ERNIE-Image Multi-Model Pipeline in ComfyUI: Production Workflows with ControlNet, IP-Adapter, and Upscaling
From single-model generation to production-grade pipelines, ERNIE-Image's true power emerges when it works in concert with other models. This deep dive explores how to build a 4-stage multi-model pipeline in ComfyUI, combining ERNIE-Image's text rendering and structured generation capabilities with ControlNet's structural precision, IP-Adapter's style transfer, and professional-grade upscalers — creating an AI image production system ready for commercial projects.
Why Build a Multi-Model Pipeline?
ERNIE-Image, as an 8B-parameter DiT model, excels in text rendering (LongTextBench 0.9733) and complex instruction following. But like any single tool, it has boundaries. ControlNet provides pixel-level structural control, IP-Adapter enables style transfer and character consistency, and upscalers bring output to print-ready resolution.
Used individually, each model solves one dimension of the problem. When fused into a pipeline, you get 1+1>2 synergy.
From a practical workflow perspective, a typical commercial project needs:
- Structural control (ControlNet): Ensure composition, perspective, and pose match requirements
- Style consistency (IP-Adapter): Brand visuals and character appearance remain consistent across images
- High-quality output (Upscaler): From 1024×1024 to 4K+ print-ready resolution
- Text precision (ERNIE-Image's core strength): Zero-error text in posters and infographics
Pipeline Architecture Overview
The complete 4-stage pipeline:
Stage 1: ERNIE-Image Base Generation
→ CheckpointLoader (ernie-image.safetensors)
→ CLIPTextEncode (prompt encoding)
→ KSampler (50 steps, DPM++ 2M)
→ VAEDecode
Stage 2: ControlNet Structural Control
→ ControlNetLoader (Canny/Depth/Pose)
→ ControlNetApply
→ Fused with Stage 1
Stage 3: IP-Adapter Style Transfer
→ CLIPVisionLoader
→ IPAdapterModelLoader (ip-adapter-ernie.bin)
→ IPAdapterApply (scale: 0.6-0.8)
Stage 4: Upscaling & Post-Processing
→ LatentUpscale (2x)
→ KSampler (refine, 20 steps)
→ ImageUpscaleWithModel (4x-UltraSharp)
→ Final Output
Stage 1: ERNIE-Image Base Generation
This is the pipeline's foundation. ERNIE-Image Base delivers the strongest instruction understanding and text rendering for complex prompts. ERNIE-Image Turbo achieves faster generation in just 8 inference steps, ideal for batch production.
# ComfyUI Node Configuration
CheckpointLoaderSimple
→ ckpt_name: "ernie-image.safetensors" (Base) or "ernie-image-turbo.safetensors" (Turbo)
CLIPTextEncode
→ clip: CLIP (from CheckpointLoader)
→ text: "detailed product photo of..."
KSampler
→ model: model (from CheckpointLoader)
→ positive: positive_conditioning
→ negative: negative_conditioning
→ steps: 50 (Base) / 8 (Turbo)
→ cfg: 4.0 (Base) / 1.0 (Turbo)
→ sampler_name: "dpmpp_2m"
→ scheduler: "karras"
Key Parameter Notes:
- Base model: 50 steps + CFG 4.0 for optimal quality
- Turbo model: 8 steps + CFG 1.0, optimized via DMD distillation
- PE (Prompt Enhancer) can be toggled on/off as needed
Stage 2: ControlNet Structural Control
ControlNet forms the structural backbone of the pipeline. ERNIE-Image supports multiple ControlNet types, each addressing different structural control needs.
Canny Edge Control
For scenarios requiring precise contour and edge control, such as product photography and architectural rendering.
ControlNetLoader
→ control_net_name: "controlnet-canny-v3.safetensors"
ControlNetApply
→ control_net: ControlNet
→ conditioning: positive_conditioning
→ image: canny_edge_map
→ strength: 0.8
Depth Control
For scenarios requiring precise spatial relationship control, such as scene reconstruction and 3D-assisted generation.
Pose Control
For scenarios requiring precise character pose control, such as character design and fashion photography.
VRAM Impact: Each ControlNet model consumes approximately 2-4 GB VRAM. When using multiple ControlNets simultaneously, memory management becomes critical.
Stage 3: IP-Adapter Style Transfer
IP-Adapter (Image Prompt Adapter) is the core tool for style transfer and character consistency. ERNIE-Image has a dedicated IP-Adapter model ip-adapter-ernie.bin.
CLIPVisionLoader
→ clip_name: "clip-vit-large-patch14.safetensors"
IPAdapterModelLoader
→ ipadapter_file: "ip-adapter-ernie.bin"
IPAdapterApply
→ ipadapter: IPAdapterModel
→ clip_vision: CLIPVision
→ image: style_reference_image
→ scale: 0.7 (recommended range 0.6-0.8)
→ start_at: 0.0
→ end_at: 1.0
Scale Parameter Tuning:
- 0.4-0.6: Light style influence, prompt content remains dominant
- 0.6-0.8: Moderate style influence, balanced between style and content
- 0.8+: Strong style influence, may override prompt content
Stage 4: Upscaling & Post-Processing
From 1024×1024 to print-ready resolution, upscaling is the pipeline's final step. A two-stage upscaling strategy is recommended.
Phase 1: Latent Upscaling
2x upscaling in latent space, followed by refinement sampling.
LatentUpscale
→ samples: latent_from_generation
→ upscale_method: "nearest-exact"
→ width: 2048
→ height: 2048
KSampler (refine)
→ steps: 20
→ cfg: 4.0
→ sampler_name: "dpmpp_2m"
→ seed: same as original (maintain consistency)
Phase 2: Image-Level Upscaling
Dedicated upscaling models for final sharpening.
UpscaleModelLoader
→ model_name: "4x-UltraSharp.pth"
ImageUpscaleWithModel
→ upscale_model: UpscaleModel
→ image: image_from_latent_upscale
Upscaling Model Selection:
- 4x-UltraSharp: General-purpose, suitable for most scenarios
- RealESRGAN: Anime and illustration styles
- SUPIR: High-fidelity upscaling, but computationally expensive
Complete Workflow JSON
The complete ComfyUI workflow JSON file is available in the GitHub repository. Below is the key node connection structure:
CheckpointLoader → CLIPTextEncode → KSampler → VAEDecode → LatentUpscale
↓
ControlNetApply → KSampler (refine)
↓
IPAdapterApply → KSampler (refine)
↓
ImageUpscaleWithModel → Final Output
Performance Optimization & VRAM Management
VRAM Requirement Estimates
| Component | VRAM Usage |
|---|---|
| ERNIE-Image Base (BF16) | ~16 GB |
| ERNIE-Image Turbo (BF16) | ~16 GB |
| ControlNet (single) | ~2-4 GB |
| IP-Adapter | ~1-2 GB |
| CLIP Vision | ~1 GB |
| Upscaling Model | ~1-2 GB |
| Total (Full Pipeline) | ~25-30 GB |
VRAM Optimization Strategies
- FP8 Quantization: Convert models to FP8 format, halving memory usage
- Model Offloading: Use
--lowvramflag, inactive models automatically offload to system memory - Staged Execution: Run generation first, free VRAM before running upscaling
- GGUF Quantization: Use Unsloth's GGUF quantized models
Recommended Hardware Configuration
| GPU | Supported Pipeline |
|---|---|
| RTX 3060 (12GB) | Base + single ControlNet or IP-Adapter |
| RTX 4070 (12GB) | Turbo + single ControlNet |
| RTX 3090 (24GB) | Full pipeline (with model offloading) |
| RTX 4090 (24GB) | Full pipeline (smooth operation) |
Application Scenarios
Scenario 1: E-commerce Product Photo Batch Production
- ControlNet (Depth): Control product angle and perspective
- IP-Adapter: Maintain brand visual style consistency
- ERNIE-Image Turbo: 8-step fast generation
- Upscaler: Output print-ready product images
Scenario 2: Character Design Consistency
- ControlNet (Pose): Control character pose
- IP-Adapter: Maintain character appearance consistency
- Generate multiple scenes with the same character
Scenario 3: Brand Visual Design
- IP-Adapter: Brand style reference
- ERNIE-Image Base: Precise text rendering (brand names, slogans)
- Upscaler: Output high-definition brand assets
Common Pitfalls & Solutions
Pitfall 1: ControlNet-Prompt Conflict
When ControlNet's structural information conflicts with prompt description, output becomes distorted.
Solution: Lower ControlNet strength (0.5-0.7), ensure prompt matches reference image.
Pitfall 2: IP-Adapter Overrides Prompt Content
When scale is too high, IP-Adapter's style information overrides prompt content.
Solution: Reduce scale to 0.6-0.7, or use start_at/end_at to control influence range.
Pitfall 3: Text Blur After Upscaling
Text edges may blur after Latent Upscale.
Solution: Generate high-res text images before upscaling, or use high-fidelity models like SUPIR.
Summary
The multi-model pipeline combines ERNIE-Image's core strengths (text rendering, structured generation) with ControlNet's structural control, IP-Adapter's style transfer, and professional-grade upscaling into a complete commercial-grade AI image production system.
Key takeaways:
- 4-stage pipeline: Generation → Structural Control → Style Transfer → Upscaling
- VRAM is the primary bottleneck; FP8 quantization and model offloading are essential
- Parameter tuning requires experimentation; no universal configuration exists
- Start with simple pipelines, gradually increase complexity
ERNIE-Image's role in multi-model pipelines is the "content generation engine" — it understands complex prompts and generates high-quality content foundations, while other models provide fine-grained control and enhancement on top. This division of labor is the right direction for modern AI image generation.
Reference Resources:
- ERNIE-Image ComfyUI Official Workflow: https://docs.comfy.org/tutorials/image/ernie-image/ernie-image
- ERNIE-Image GitHub: https://github.com/baidu/ernie-image
- ComfyUI ControlNet + IP-Adapter Workflow: https://comfyui.org/en/image-style-transfer-controlnet-ipadapter-workflow
- ComfyUI VRAM Optimization Guide: https://www.synpixcloud.com/blog/comfyui-complex-workflow-gpu-guide