ERNIE-Image ComfyUI Official Support and Custom Nodes Complete Guide

Jun 8, 2026

ERNIE-Image ComfyUI Official Support and Custom Nodes Complete Guide

Summary: ComfyUI now officially supports ERNIE-Image starting from v0.19.1, with ready-to-load official workflows available in the template library. This article provides an in-depth guide to ERNIE-Image deployment in ComfyUI, practical applications of custom nodes like Pixaroma Note, ControlNet combination workflows, and advanced techniques from single-image generation to batch production. Whether you're a ComfyUI beginner or an experienced workflow developer, you'll find practical solutions to boost your productivity.


A Milestone: Official ComfyUI Support

In early 2026, ComfyUI officially announced Day-0 support for ERNIE-Image, which means:

  1. Native Integration: ComfyUI v0.19.1+ includes ERNIE-Image workflow templates
  2. One-Click Loading: Access complete workflows via Templates → ERNIE-Image
  3. Automatic Model Download: Required model files are fetched from HuggingFace automatically
  4. Dual Model Support: ERNIE-Image Base (50 steps, high quality) and ERNIE-Image-Turbo (8 steps, fast iteration)

This marks ERNIE-Image's transition from "community manual adaptation" to "official ecosystem support."


Core ERNIE-Image Architecture in ComfyUI

The ERNIE-Image workflow in ComfyUI consists of these core nodes:

[CLIP Text Encode] → [ERNIE-Image Diffusion Model] → [VAE Decode] → [Save Image]
        ↑                       ↑                           ↑
   Prompt Text          ernie-image.safetensors      flux2_ae.safetensors
   (or PE output)       (or turbo.safetensors)

Key component details:

Component File Purpose Source
Diffusion Model ernie-image.safetensors Main generation model (8B DiT) Comfy-Org/ERNIE-Image
Turbo Model ernie-image-turbo.safetensors Distilled acceleration (8 steps) Comfy-Org/ERNIE-Image
Prompt Enhancer ministral-3-3b.safetensors 3B prompt enhancement baidu/ERNIE-Image
Text Encoder CLIP/CLIPViT Text condition encoding Built-in
VAE flux2_ae.safetensors FLUX.2 variational autoencoder FLUX.2

Important: ERNIE-Image uses FLUX.2's VAE (flux2_ae.safetensors), NOT the traditional SD VAE. This is one of the most common configuration errors in ComfyUI.


Quick Start: From Installation to First Image

Step 1: Update ComfyUI

cd ComfyUI
git pull origin master
# Ensure version ≥ v0.19.1

Step 2: Load ERNIE-Image Workflow

  1. Open the ComfyUI web interface
  2. Click the Templates tab
  3. Search for ERNIE-Image
  4. Select ERNIE-Image or ERNIE-Image-Turbo workflow
  5. Click Load to load

Step 3: Download Model Files

ComfyUI automatically detects missing models and prompts download. Files are located at:

  • ComfyUI/models/diffusion_models/ernie-image.safetensors (Base)
  • ComfyUI/models/diffusion_models/ernie-image-turbo.safetensors (Turbo)
  • ComfyUI/models/vae/flux2_ae.safetensors (VAE)
  • ComfyUI/models/text_encoders/ministral-3-3b.safetensors (PE)

Step 4: Generate Your First Image

  1. Enter your prompt in the CLIP Text Encode node
  2. Set sampler (Turbo: Euler, 8 steps; Base: DPM++ 2M, 20-50 steps)
  3. Set CFG Scale (Turbo: 1.0, Base: 5.0-7.0)
  4. Click Queue Prompt to start generation

Custom Nodes: Pixaroma Note and Resolution Nodes

Pixaroma Note: Rich Text Prompt Editor

Source: Custom node from YouTube Ep14 tutorial

Problem Solved: ERNIE-Image's text rendering is powerful, but ComfyUI's default text input box has limited support for multi-line, multilingual, and complex structured prompts.

Features:

  • Rich Text Editing: Multi-paragraph, line breaks, indentation
  • Color-Coded Text: Different colors for different prompt sections (subject, background, style, text)
  • Highlight Marking: Quick identification of key parameters
  • Lists/Code Blocks: Structured prompt organization

Installation:

cd ComfyUI/custom_nodes
git clone https://github.com/Pixaroma/comfyui-pixaroma-nodes
pip install -r requirements.txt

Typical workflow:

[Pixaroma Note (prompt editing)] → [CLIP Text Encode] → [KSampler] → [VAE Decode] → [Save Image]

Resolution Node: Smart Resolution Calculation

Problem Solved: ERNIE-Image's optimal output resolution varies by use case — posters need aspect ratios, social media needs specific dimensions, and product shots need square format.

Features: Automatically calculates optimal output resolution based on prompt content or preset templates.

Typical presets:

  • Social media cover: 1280×640 (2:1)
  • Product image: 1024×1024 (1:1)
  • Poster design: 1024×1536 (2:3)
  • YouTube thumbnail: 1280×720 (16:9)

Advanced Combination Workflows

Workflow 1: ControlNet + Super Resolution

Use Case: Convert line art into high-quality illustrations or 3D-style renders

[Load Image (line art)] → [Canny Edge Detect] → [ControlNet] → [KSampler (ERNIE-Image)] → [ESRGAN Upscale] → [Save Image]

Key configuration:

  • ControlNet strength: 0.5-0.8
  • Sampler: DPM++ 2M Karras
  • Steps: 30-50
  • Upscale ratio: 2x-4x

Workflow 2: Image-to-Video (WanVideo + GIMM-VFI)

Use Case: Convert ERNIE-Image static images into short videos

[ERNIE-Image Generate] → [WanVideo I2V] → [GIMM-VFI (frame interpolation)] → [Save Video]

Note: This is a ComfyUI officially recommended combination workflow, using WanVideo for image-to-video conversion and GIMM-VFI for frame rate enhancement.

Workflow 3: FP8 Hybrid Precision Upscaling

Use Case: Achieve high-quality high-resolution output with limited VRAM

[KSampler (FP8 Base, 512×512)] → [KSampler (FP8 Upscale, 1024×1024)] → [Save Image]

Key configuration:

  • Stage 1: 512×512, FP8 model, 20 steps
  • Stage 2: 1024×1024, FP8 model, 15 steps, low noise
  • Total VRAM usage: ~60% less than BF16 1024×1024 direct generation

Workflow 4: ICEdit + Flux + ESRGAN Local Inpainting

Use Case: Refined modification of specific regions in ERNIE-Image output

[Load Image] → [ICEdit Inpaint] → [Flux Style Transfer] → [ESRGAN Upscale] → [Save Image]

Batch Production Workflow

For e-commerce product shots and social media batch images, ComfyUI supports batch prompt input:

# Use Wildcard or text file for batch input
# prompts.txt - one prompt per line
prompt1: Blue sneakers on white background, product photography style
prompt2: Red headphones on white background, product photography style
prompt3: Black watch on white background, product photography style

ComfyUI batch configuration:

  1. Use Load Text node to read prompts.txt
  2. Pair with Batch Prompt node to split prompts
  3. Each prompt generates one independent image
  4. Save Image node automatically saves with sequential numbering

Cloud Option: Floyo H100 Cloud Workflow

If you don't have a local GPU, Floyo provides cloud-based ERNIE-Image Turbo:

Comparison with local deployment:

Dimension Floyo Cloud Local ComfyUI
Initial Cost $0 GPU hardware cost
Per-Image Cost ~$0.05-0.10 Electricity (negligible)
Latency Network + queue Zero local latency
Privacy Data uploaded to cloud Fully local
Best For Occasional use, no GPU High-frequency use, has GPU

FAQ

Q: What's the difference between ERNIE-Image Base and Turbo in ComfyUI?

Feature Base Turbo
Parameters 8B DiT 8B DiT (distilled)
Steps 20-50 8
Quality Best Close to Base
CFG Scale 5.0-7.0 1.0
Speed Slower 3-6× faster
Recommended For Final output, print Quick iteration, preview

Q: How to enable Prompt Enhancer in ComfyUI?

Add an ERNIE PE node before the CLIP Text Encode node, connected to ministral-3-3b.safetensors. Short prompts are first expanded by PE, then fed to the main model.

Q: What to do when VRAM is insufficient?

  1. Use Turbo model (lower VRAM requirements)
  2. Enable --lowvram startup parameter
  3. Use FP8 quantized models
  4. Lower output resolution and upscale

Conclusion

ComfyUI's official support for ERNIE-Image marks the maturation of the open-source text-to-image ecosystem. From basic workflows to custom nodes, from single-image generation to batch production, from local deployment to cloud options — ERNIE-Image's integration in the ComfyUI ecosystem is rapidly evolving.

Key recommendations:

  1. Beginners: Use official templates to get started, first run the Base model
  2. Intermediate Users: Pair with Pixaroma Note for complex prompt writing
  3. Production Users: Batch workflow + ControlNet combination for automated production
  4. No GPU Users: Floyo cloud solution for zero-threshold experience

Keywords: ernie-image comfyui ernie-image comfyui workflow comfyui ernie-image official support ernie-image pixaroma nodes comfyui ernie-image controlnet ernie-image batch generation comfyui

ERNIE-Image Team

ERNIE-Image ComfyUI Official Support and Custom Nodes Complete Guide | Blog