ERNIE-Image ComfyUI Official Support and Custom Nodes Complete Guide
Summary: ComfyUI now officially supports ERNIE-Image starting from v0.19.1, with ready-to-load official workflows available in the template library. This article provides an in-depth guide to ERNIE-Image deployment in ComfyUI, practical applications of custom nodes like Pixaroma Note, ControlNet combination workflows, and advanced techniques from single-image generation to batch production. Whether you're a ComfyUI beginner or an experienced workflow developer, you'll find practical solutions to boost your productivity.
A Milestone: Official ComfyUI Support
In early 2026, ComfyUI officially announced Day-0 support for ERNIE-Image, which means:
- Native Integration: ComfyUI v0.19.1+ includes ERNIE-Image workflow templates
- One-Click Loading: Access complete workflows via
Templates→ERNIE-Image - Automatic Model Download: Required model files are fetched from HuggingFace automatically
- Dual Model Support: ERNIE-Image Base (50 steps, high quality) and ERNIE-Image-Turbo (8 steps, fast iteration)
This marks ERNIE-Image's transition from "community manual adaptation" to "official ecosystem support."
Core ERNIE-Image Architecture in ComfyUI
The ERNIE-Image workflow in ComfyUI consists of these core nodes:
[CLIP Text Encode] → [ERNIE-Image Diffusion Model] → [VAE Decode] → [Save Image]
↑ ↑ ↑
Prompt Text ernie-image.safetensors flux2_ae.safetensors
(or PE output) (or turbo.safetensors)
Key component details:
| Component | File | Purpose | Source |
|---|---|---|---|
| Diffusion Model | ernie-image.safetensors |
Main generation model (8B DiT) | Comfy-Org/ERNIE-Image |
| Turbo Model | ernie-image-turbo.safetensors |
Distilled acceleration (8 steps) | Comfy-Org/ERNIE-Image |
| Prompt Enhancer | ministral-3-3b.safetensors |
3B prompt enhancement | baidu/ERNIE-Image |
| Text Encoder | CLIP/CLIPViT | Text condition encoding | Built-in |
| VAE | flux2_ae.safetensors |
FLUX.2 variational autoencoder | FLUX.2 |
Important: ERNIE-Image uses FLUX.2's VAE (
flux2_ae.safetensors), NOT the traditional SD VAE. This is one of the most common configuration errors in ComfyUI.
Quick Start: From Installation to First Image
Step 1: Update ComfyUI
cd ComfyUI
git pull origin master
# Ensure version ≥ v0.19.1
Step 2: Load ERNIE-Image Workflow
- Open the ComfyUI web interface
- Click the
Templatestab - Search for
ERNIE-Image - Select
ERNIE-ImageorERNIE-Image-Turboworkflow - Click
Loadto load
Step 3: Download Model Files
ComfyUI automatically detects missing models and prompts download. Files are located at:
ComfyUI/models/diffusion_models/ernie-image.safetensors(Base)ComfyUI/models/diffusion_models/ernie-image-turbo.safetensors(Turbo)ComfyUI/models/vae/flux2_ae.safetensors(VAE)ComfyUI/models/text_encoders/ministral-3-3b.safetensors(PE)
Step 4: Generate Your First Image
- Enter your prompt in the
CLIP Text Encodenode - Set sampler (Turbo:
Euler, 8 steps; Base:DPM++ 2M, 20-50 steps) - Set CFG Scale (Turbo: 1.0, Base: 5.0-7.0)
- Click
Queue Promptto start generation
Custom Nodes: Pixaroma Note and Resolution Nodes
Pixaroma Note: Rich Text Prompt Editor
Source: Custom node from YouTube Ep14 tutorial
Problem Solved: ERNIE-Image's text rendering is powerful, but ComfyUI's default text input box has limited support for multi-line, multilingual, and complex structured prompts.
Features:
- Rich Text Editing: Multi-paragraph, line breaks, indentation
- Color-Coded Text: Different colors for different prompt sections (subject, background, style, text)
- Highlight Marking: Quick identification of key parameters
- Lists/Code Blocks: Structured prompt organization
Installation:
cd ComfyUI/custom_nodes
git clone https://github.com/Pixaroma/comfyui-pixaroma-nodes
pip install -r requirements.txt
Typical workflow:
[Pixaroma Note (prompt editing)] → [CLIP Text Encode] → [KSampler] → [VAE Decode] → [Save Image]
Resolution Node: Smart Resolution Calculation
Problem Solved: ERNIE-Image's optimal output resolution varies by use case — posters need aspect ratios, social media needs specific dimensions, and product shots need square format.
Features: Automatically calculates optimal output resolution based on prompt content or preset templates.
Typical presets:
- Social media cover: 1280×640 (2:1)
- Product image: 1024×1024 (1:1)
- Poster design: 1024×1536 (2:3)
- YouTube thumbnail: 1280×720 (16:9)
Advanced Combination Workflows
Workflow 1: ControlNet + Super Resolution
Use Case: Convert line art into high-quality illustrations or 3D-style renders
[Load Image (line art)] → [Canny Edge Detect] → [ControlNet] → [KSampler (ERNIE-Image)] → [ESRGAN Upscale] → [Save Image]
Key configuration:
- ControlNet strength: 0.5-0.8
- Sampler: DPM++ 2M Karras
- Steps: 30-50
- Upscale ratio: 2x-4x
Workflow 2: Image-to-Video (WanVideo + GIMM-VFI)
Use Case: Convert ERNIE-Image static images into short videos
[ERNIE-Image Generate] → [WanVideo I2V] → [GIMM-VFI (frame interpolation)] → [Save Video]
Note: This is a ComfyUI officially recommended combination workflow, using WanVideo for image-to-video conversion and GIMM-VFI for frame rate enhancement.
Workflow 3: FP8 Hybrid Precision Upscaling
Use Case: Achieve high-quality high-resolution output with limited VRAM
[KSampler (FP8 Base, 512×512)] → [KSampler (FP8 Upscale, 1024×1024)] → [Save Image]
Key configuration:
- Stage 1: 512×512, FP8 model, 20 steps
- Stage 2: 1024×1024, FP8 model, 15 steps, low noise
- Total VRAM usage: ~60% less than BF16 1024×1024 direct generation
Workflow 4: ICEdit + Flux + ESRGAN Local Inpainting
Use Case: Refined modification of specific regions in ERNIE-Image output
[Load Image] → [ICEdit Inpaint] → [Flux Style Transfer] → [ESRGAN Upscale] → [Save Image]
Batch Production Workflow
For e-commerce product shots and social media batch images, ComfyUI supports batch prompt input:
# Use Wildcard or text file for batch input
# prompts.txt - one prompt per line
prompt1: Blue sneakers on white background, product photography style
prompt2: Red headphones on white background, product photography style
prompt3: Black watch on white background, product photography style
ComfyUI batch configuration:
- Use
Load Textnode to readprompts.txt - Pair with
Batch Promptnode to split prompts - Each prompt generates one independent image
Save Imagenode automatically saves with sequential numbering
Cloud Option: Floyo H100 Cloud Workflow
If you don't have a local GPU, Floyo provides cloud-based ERNIE-Image Turbo:
- Link: https://www.floyo.ai/workflows/ernie-turbo-text-to-image-workflow
- Hardware: H100 cloud GPU
- Advantage: No local installation, browser-based
- Cost: Pay-per-use, ideal for occasional users
Comparison with local deployment:
| Dimension | Floyo Cloud | Local ComfyUI |
|---|---|---|
| Initial Cost | $0 | GPU hardware cost |
| Per-Image Cost | ~$0.05-0.10 | Electricity (negligible) |
| Latency | Network + queue | Zero local latency |
| Privacy | Data uploaded to cloud | Fully local |
| Best For | Occasional use, no GPU | High-frequency use, has GPU |
FAQ
Q: What's the difference between ERNIE-Image Base and Turbo in ComfyUI?
| Feature | Base | Turbo |
|---|---|---|
| Parameters | 8B DiT | 8B DiT (distilled) |
| Steps | 20-50 | 8 |
| Quality | Best | Close to Base |
| CFG Scale | 5.0-7.0 | 1.0 |
| Speed | Slower | 3-6× faster |
| Recommended For | Final output, print | Quick iteration, preview |
Q: How to enable Prompt Enhancer in ComfyUI?
Add an ERNIE PE node before the CLIP Text Encode node, connected to ministral-3-3b.safetensors. Short prompts are first expanded by PE, then fed to the main model.
Q: What to do when VRAM is insufficient?
- Use Turbo model (lower VRAM requirements)
- Enable
--lowvramstartup parameter - Use FP8 quantized models
- Lower output resolution and upscale
Conclusion
ComfyUI's official support for ERNIE-Image marks the maturation of the open-source text-to-image ecosystem. From basic workflows to custom nodes, from single-image generation to batch production, from local deployment to cloud options — ERNIE-Image's integration in the ComfyUI ecosystem is rapidly evolving.
Key recommendations:
- Beginners: Use official templates to get started, first run the Base model
- Intermediate Users: Pair with Pixaroma Note for complex prompt writing
- Production Users: Batch workflow + ControlNet combination for automated production
- No GPU Users: Floyo cloud solution for zero-threshold experience
Keywords: ernie-image comfyui ernie-image comfyui workflow comfyui ernie-image official support ernie-image pixaroma nodes comfyui ernie-image controlnet ernie-image batch generation comfyui