ERNIE-Image + Krita AI Diffusion Plugin: Professional AI Painting Workflow for Digital Artists

Jul 10, 2026

ERNIE-Image + Krita AI Diffusion Plugin: Professional AI Painting Workflow for Digital Artists

Why This Matters

From Stable Diffusion to ComfyUI, AI image generation has always been missing a critical piece: native integration with professional digital painting software. While ComfyUI is powerful, its node-based interface forces artists into a disruptive cycle: generate in ComfyUI, screenshot, paste into Krita/Photoshop for edits, then export back to ComfyUI for regeneration. This fragmented workflow severely impacts creative efficiency.

In May 2026, Krita AI Diffusion — the most popular open-source AI painting plugin on GitHub (10.3k Stars) — officially added native support for ERNIE-Image in v1.50.0. This means digital artists can now invoke ERNIE-Image's generation capabilities directly within the Krita canvas: selection fill, inpainting, outpainting, style transfer — all completed without ever leaving the creative environment.

Krita + ERNIE-Image: A Perfect Match

Krita: The Gold Standard of Open-Source Digital Painting

Krita is a free, open-source digital painting application widely used by professional illustrators, concept artists, and comic artists worldwide. It features a complete layer system, color management, brush engines, and vector layers — the most powerful open-source alternative to Adobe Photoshop for digital painting.

Krita AI Diffusion Plugin: Bringing AI Into the Canvas

The Krita AI Diffusion plugin (Acly/krita-ai-diffusion) acts as a bridge for using generative AI inside Krita. It doesn't replace ComfyUI — it seamlessly embeds AI image generation into Krita's painting workflow:

  • Selection Fill: Select an area on the canvas, and AI automatically generates content
  • Inpainting: Mask + prompt for precise local edits
  • Outpainting: Extend canvas boundaries with AI-generated content
  • Live Painting: Every brush stroke is interpreted by AI in real time
  • ControlNet Layers: Scribble, line art, depth maps, pose guides, and more
  • IP-Adapter References: Style transfer, composition transfer, face swap
  • Regions: Different prompts for different layers, fine-grained control over each area

ERNIE-Image's Strengths Fully Realized in Krita

ERNIE-Image's three core capabilities — complex instruction following, precise text rendering, and structured image generation — shine in Krita's selection fill scenarios:

  • In poster design, select a text area, enter "Logo position: center, brand name" — ERNIE-Image's LongTextBench score of 0.8625 ensures crisp, readable text
  • In comic creation, use Krita's multi-panel layout with ERNIE-Image's structured generation — generate storyboard thumbnails in one click
  • In product retouching, use selection mask + Inpainting to fix flaws — ERNIE-Image's instruction-following capability makes repairs precise and controllable

Installation & Configuration

Requirements

Component Requirement
Krita 5.2.0 or newer
Plugin Version v1.50.0+ (recommended v1.52.1)
GPU NVIDIA 6GB+ VRAM / AMD ROCm / Apple Silicon
Backend ComfyUI (auto-managed by plugin)

Step 1: Install Krita and the Plugin

  1. Download and install Krita 5.2.0+ from krita.org
  2. Download the latest plugin ZIP from GitHub Releases
  3. Open Krita → Tools ▸ Scripts ▸ Import Python Plugin from File... → select the ZIP
  4. Restart Krita, create or open a document
  5. Enable the panel: Settings ▸ Panels ▸ AI Image Generation
  6. Click "Configure" to install the local ComfyUI backend

Step 2: Download ERNIE-Image Model Files

Place the following files into the plugin-managed models/ directory:

📂 models/
├── 📂 diffusion_models/
│   ├── ernie-image.safetensors          (Base model, 50 steps)
│   └── ernie-image-turbo.safetensors    (Turbo model, 8 steps)
├── 📂 text_encoders/
│   └── ministral-3-3b.safetensors       (3B MiniStral text encoder)
└── 📂 vae/
    └── flux2-vae.safetensors            (FLUX.2 VAE)

Download links (from HuggingFace Comfy-Org/ERNIE-Image):

⚠️ Note: The current Krita plugin version does not support the Prompt Enhancer (PE) model. MiniStral 3B serves as the primary text encoder.

Step 3: Configure a Style Preset

In the Krita AI panel, click Manage Presets to create a new Style Preset:

Parameter Recommended Value
Base Model ERNIE-Image
Sampler Euler
Scheduler Simple
CFG Scale 4.0 (Base) / 1.0 (Turbo)
Steps 50 (Base) / 8 (Turbo)

Practical Workflows

Workflow 1: Selection Fill — From Blank to Finished

This is the core functionality of Krita AI Diffusion. Create a selection on the canvas, and the plugin understands your intent:

  1. Generate New Content: Draw a rectangle in an empty area → enter a prompt → click Generate → ERNIE-Image fills the selection
  2. Replace Content: Select the area to replace → describe what you want → click Fill
  3. Remove Objects: Select unwanted objects → describe background continuation → ERNIE-Image intelligently fills

Pro tip: Use ERNIE-Image Turbo (8 steps) for rapid iterative composition, then switch to Base (50 steps) for the final output. Turbo mode delivers ~2-3 seconds per image (RTX 4090), Base mode ~8-12 seconds.

Workflow 2: Inpainting — Precision Edits

Krita AI Diffusion's Inpainting is more intuitive than ComfyUI's:

  1. Select the layer to modify
  2. Use Krita's selection tools (lasso, rectangle, circle) to mark the area
  3. In the AI panel, select Inpaint mode
  4. Enter a modification prompt (e.g., "change the red dress to blue")
  5. Adjust Strength: 0.3-0.7 for subtle edits, 0.7-1.0 for major changes

ERNIE-Image's Inpainting capability, powered by its robust img2img pipeline, combined with Krita's precise selection tools, delivers Photoshop-quality local edits. This is especially useful for e-commerce product retouching, portrait post-processing, and architectural rendering adjustments.

Four core modes:

  • Fill: General-purpose mode, works for most scenarios
  • Add: Insert new objects within the selection, preserving original content
  • Remove: Remove objects and fill the background
  • Replace: Replace entirely with new content

Workflow 3: ControlNet + ERNIE-Image = Precise Composition

Krita AI Diffusion supports the full ControlNet ecosystem as control layers:

  1. Scribble Layer: Draw rough lines that ERNIE-Image transforms into refined images
  2. Line Art Layer: Import or draw line art for AI coloring and refinement
  3. Depth Map: Control depth relationships from 3D scenes or photos
  4. Pose Map: OpenPose skeletons for precise human pose control

Typical workflow:

  1. Quickly sketch a composition in Krita
  2. Set the sketch layer as a Scribble control layer
  3. Enter a detailed prompt in the AI panel
  4. ERNIE-Image generates a complete image based on the sketch structure and prompt
  5. Continue painting on the generated result → refine details with Inpainting

Workflow 4: IP-Adapter Style Transfer

ERNIE-Image with IP-Adapter enables zero-training style transfer in Krita:

  1. Import a reference image as a layer
  2. In the AI panel's Reference Image, select that layer
  3. Choose Style Transfer mode
  4. Enter new content prompts
  5. ERNIE-Image generates new imagery while preserving the reference style

This is a powerful tool for designers maintaining brand visual consistency: train a brand LoRA once, then repeatedly apply the brand style to different designs using IP-Adapter in Krita.

Workflow 5: Live Painting

Live Painting mode is one of the Krita plugin's most innovative features. As you paint on the canvas, AI interprets each stroke in real time:

  1. Set a layer to Live Painting mode
  2. Scribble freely on the canvas — AI generates corresponding imagery in real time
  3. Continue adding strokes as AI continuously updates the composition
  4. When satisfied, lock the image and refine further with traditional painting tools

This is effectively "AI-assisted brainstorming" — exploring creative directions quickly without needing precise prompt engineering.

Best Practices for ERNIE-Image in Krita

Base vs Turbo Selection Strategy

Use Case Recommended Version Reason
Quick composition exploration Turbo (8 steps) 3-4x faster, sufficient for viability checks
Selection Fill Turbo (8 steps) Fill edges need natural blending; Turbo quality is adequate
Final output Base (50 steps) Higher instruction fidelity, more reliable for complex scenes
Text rendering Base (50 steps) LongTextBench 0.8625 vs Turbo 0.7950
Inpainting refinement Base (50 steps) Stronger semantic understanding, more natural results
Live Painting Turbo (8 steps) Real-time feedback requires minimal latency

Key Parameter Tuning

  • CFG Scale: 4.0 for Base; must be 1.0 for Turbo (Turbo is a distilled model, insensitive to CFG)
  • Strength (img2img): 0.5-0.8 is optimal for Inpainting; too low retains too much original content, too high deviates from the source
  • Prompt Guidance: Use "worst quality, low quality, blurry, distorted" as negative prompts for significantly improved output

Coordinating with ComfyUI

Krita AI Diffusion uses ComfyUI as its inference backend. If you have an existing custom ComfyUI installation, you can point the plugin to it instead of managing a new instance:

  1. In plugin config, select "Use Custom ComfyUI Server"
  2. Enter your ComfyUI address (e.g., http://localhost:8188)
  3. Ensure comfyui-tooling-nodes is up to date

This allows you to run complex workflow pipelines in ComfyUI simultaneously while doing touch-ups and refinements in Krita — both sharing the same inference backend and model cache.

Troubleshooting

ERNIE-Image Not Detected in the Plugin

Ensure Krita AI Diffusion plugin >= 1.50.0 and comfyui-tooling-nodes is updated. Some safetensors files downloaded from HuggingFace may lack metadata, but detection does not rely on metadata — updating tooling-nodes resolves the issue.

Why PE (Prompt Enhancer) Is Not Available

The current Krita plugin's text encoder pipeline doesn't support ERNIE-Image's 3B PE model. The plugin uses MiniStral 3B as the text encoder. PE is only available in native ComfyUI or Diffusers workflows. In Krita, manually write detailed prompts to compensate.

Insufficient GPU VRAM

  • Use Turbo variant (8-step inference, lower VRAM footprint)
  • Lower output resolution in plugin settings (1024×1024 → 768×768)
  • Enable enable_model_cpu_offload() mode (auto-managed by plugin)
  • Apple Silicon Mac: use MPS backend; 8GB unified memory is sufficient for Turbo

Summary

Krita AI Diffusion's native ERNIE-Image support fills the critical gap between professional digital painting tools and AI image generation. For digital artists, concept designers, comic creators, and graphic designers, this means:

  • No more window switching — complete the full workflow from concept to final output in Krita
  • Control meets creativity — use Krita's layer system to precisely define AI's working area
  • Professional-grade output — ERNIE-Image's text rendering and structured image capabilities ensure commercial-grade results
  • Zero-barrier AI workflow — presets and selection modes make AI accessible to artists who don't know ComfyUI

Krita's free, open-source philosophy + ERNIE-Image's Apache 2.0 license = a completely free open-source digital creation platform with no feature limits or monthly quotas. For independent creators and small design teams on tight budgets who demand professional quality, this is an unprecedented opportunity.

ERNIE-Image Team