ERNIE-Image Multi-Model Pipeline in ComfyUI: Production Workflows with ControlNet, IP-Adapter, and Upscaling

Jul 3, 2026

ERNIE-Image Multi-Model Pipeline in ComfyUI: Production Workflows with ControlNet, IP-Adapter, and Upscaling

From single-model generation to production-grade pipelines, ERNIE-Image's true power emerges when it works in concert with other models. This deep dive explores how to build a 4-stage multi-model pipeline in ComfyUI, combining ERNIE-Image's text rendering and structured generation capabilities with ControlNet's structural precision, IP-Adapter's style transfer, and professional-grade upscalers — creating an AI image production system ready for commercial projects.

Why Build a Multi-Model Pipeline?

ERNIE-Image, as an 8B-parameter DiT model, excels in text rendering (LongTextBench 0.9733) and complex instruction following. But like any single tool, it has boundaries. ControlNet provides pixel-level structural control, IP-Adapter enables style transfer and character consistency, and upscalers bring output to print-ready resolution.

Used individually, each model solves one dimension of the problem. When fused into a pipeline, you get 1+1>2 synergy.

From a practical workflow perspective, a typical commercial project needs:

  • Structural control (ControlNet): Ensure composition, perspective, and pose match requirements
  • Style consistency (IP-Adapter): Brand visuals and character appearance remain consistent across images
  • High-quality output (Upscaler): From 1024×1024 to 4K+ print-ready resolution
  • Text precision (ERNIE-Image's core strength): Zero-error text in posters and infographics

Pipeline Architecture Overview

The complete 4-stage pipeline:

Stage 1: ERNIE-Image Base Generation
  → CheckpointLoader (ernie-image.safetensors)
  → CLIPTextEncode (prompt encoding)
  → KSampler (50 steps, DPM++ 2M)
  → VAEDecode

Stage 2: ControlNet Structural Control
→ ControlNetLoader (Canny/Depth/Pose)
→ ControlNetApply
→ Fused with Stage 1

Stage 3: IP-Adapter Style Transfer
→ CLIPVisionLoader
→ IPAdapterModelLoader (ip-adapter-ernie.bin)
→ IPAdapterApply (scale: 0.6-0.8)

Stage 4: Upscaling & Post-Processing
→ LatentUpscale (2x)
→ KSampler (refine, 20 steps)
→ ImageUpscaleWithModel (4x-UltraSharp)
→ Final Output

Stage 1: ERNIE-Image Base Generation

This is the pipeline's foundation. ERNIE-Image Base delivers the strongest instruction understanding and text rendering for complex prompts. ERNIE-Image Turbo achieves faster generation in just 8 inference steps, ideal for batch production.

# ComfyUI Node Configuration
CheckpointLoaderSimple
  → ckpt_name: "ernie-image.safetensors" (Base) or "ernie-image-turbo.safetensors" (Turbo)

CLIPTextEncode
→ clip: CLIP (from CheckpointLoader)
→ text: "detailed product photo of..."

KSampler
→ model: model (from CheckpointLoader)
→ positive: positive_conditioning
→ negative: negative_conditioning
→ steps: 50 (Base) / 8 (Turbo)
→ cfg: 4.0 (Base) / 1.0 (Turbo)
→ sampler_name: "dpmpp_2m"
→ scheduler: "karras"

Key Parameter Notes:

  • Base model: 50 steps + CFG 4.0 for optimal quality
  • Turbo model: 8 steps + CFG 1.0, optimized via DMD distillation
  • PE (Prompt Enhancer) can be toggled on/off as needed

Stage 2: ControlNet Structural Control

ControlNet forms the structural backbone of the pipeline. ERNIE-Image supports multiple ControlNet types, each addressing different structural control needs.

Canny Edge Control

For scenarios requiring precise contour and edge control, such as product photography and architectural rendering.

ControlNetLoader
  → control_net_name: "controlnet-canny-v3.safetensors"

ControlNetApply
→ control_net: ControlNet
→ conditioning: positive_conditioning
→ image: canny_edge_map
→ strength: 0.8

Depth Control

For scenarios requiring precise spatial relationship control, such as scene reconstruction and 3D-assisted generation.

Pose Control

For scenarios requiring precise character pose control, such as character design and fashion photography.

VRAM Impact: Each ControlNet model consumes approximately 2-4 GB VRAM. When using multiple ControlNets simultaneously, memory management becomes critical.

Stage 3: IP-Adapter Style Transfer

IP-Adapter (Image Prompt Adapter) is the core tool for style transfer and character consistency. ERNIE-Image has a dedicated IP-Adapter model ip-adapter-ernie.bin.

CLIPVisionLoader
  → clip_name: "clip-vit-large-patch14.safetensors"

IPAdapterModelLoader
→ ipadapter_file: "ip-adapter-ernie.bin"

IPAdapterApply
→ ipadapter: IPAdapterModel
→ clip_vision: CLIPVision
→ image: style_reference_image
→ scale: 0.7 (recommended range 0.6-0.8)
→ start_at: 0.0
→ end_at: 1.0

Scale Parameter Tuning:

  • 0.4-0.6: Light style influence, prompt content remains dominant
  • 0.6-0.8: Moderate style influence, balanced between style and content
  • 0.8+: Strong style influence, may override prompt content

Stage 4: Upscaling & Post-Processing

From 1024×1024 to print-ready resolution, upscaling is the pipeline's final step. A two-stage upscaling strategy is recommended.

Phase 1: Latent Upscaling

2x upscaling in latent space, followed by refinement sampling.

LatentUpscale
  → samples: latent_from_generation
  → upscale_method: "nearest-exact"
  → width: 2048
  → height: 2048

KSampler (refine)
→ steps: 20
→ cfg: 4.0
→ sampler_name: "dpmpp_2m"
→ seed: same as original (maintain consistency)

Phase 2: Image-Level Upscaling

Dedicated upscaling models for final sharpening.

UpscaleModelLoader
  → model_name: "4x-UltraSharp.pth"

ImageUpscaleWithModel
→ upscale_model: UpscaleModel
→ image: image_from_latent_upscale

Upscaling Model Selection:

  • 4x-UltraSharp: General-purpose, suitable for most scenarios
  • RealESRGAN: Anime and illustration styles
  • SUPIR: High-fidelity upscaling, but computationally expensive

Complete Workflow JSON

The complete ComfyUI workflow JSON file is available in the GitHub repository. Below is the key node connection structure:

CheckpointLoader → CLIPTextEncode → KSampler → VAEDecode → LatentUpscale
                                                    ↓
                                            ControlNetApply → KSampler (refine)
                                                    ↓
                                            IPAdapterApply → KSampler (refine)
                                                    ↓
                                            ImageUpscaleWithModel → Final Output

Performance Optimization & VRAM Management

VRAM Requirement Estimates

Component VRAM Usage
ERNIE-Image Base (BF16) ~16 GB
ERNIE-Image Turbo (BF16) ~16 GB
ControlNet (single) ~2-4 GB
IP-Adapter ~1-2 GB
CLIP Vision ~1 GB
Upscaling Model ~1-2 GB
Total (Full Pipeline) ~25-30 GB

VRAM Optimization Strategies

  1. FP8 Quantization: Convert models to FP8 format, halving memory usage
  2. Model Offloading: Use --lowvram flag, inactive models automatically offload to system memory
  3. Staged Execution: Run generation first, free VRAM before running upscaling
  4. GGUF Quantization: Use Unsloth's GGUF quantized models

Recommended Hardware Configuration

GPU Supported Pipeline
RTX 3060 (12GB) Base + single ControlNet or IP-Adapter
RTX 4070 (12GB) Turbo + single ControlNet
RTX 3090 (24GB) Full pipeline (with model offloading)
RTX 4090 (24GB) Full pipeline (smooth operation)

Application Scenarios

Scenario 1: E-commerce Product Photo Batch Production

  • ControlNet (Depth): Control product angle and perspective
  • IP-Adapter: Maintain brand visual style consistency
  • ERNIE-Image Turbo: 8-step fast generation
  • Upscaler: Output print-ready product images

Scenario 2: Character Design Consistency

  • ControlNet (Pose): Control character pose
  • IP-Adapter: Maintain character appearance consistency
  • Generate multiple scenes with the same character

Scenario 3: Brand Visual Design

  • IP-Adapter: Brand style reference
  • ERNIE-Image Base: Precise text rendering (brand names, slogans)
  • Upscaler: Output high-definition brand assets

Common Pitfalls & Solutions

Pitfall 1: ControlNet-Prompt Conflict

When ControlNet's structural information conflicts with prompt description, output becomes distorted.
Solution: Lower ControlNet strength (0.5-0.7), ensure prompt matches reference image.

Pitfall 2: IP-Adapter Overrides Prompt Content

When scale is too high, IP-Adapter's style information overrides prompt content.
Solution: Reduce scale to 0.6-0.7, or use start_at/end_at to control influence range.

Pitfall 3: Text Blur After Upscaling

Text edges may blur after Latent Upscale.
Solution: Generate high-res text images before upscaling, or use high-fidelity models like SUPIR.

Summary

The multi-model pipeline combines ERNIE-Image's core strengths (text rendering, structured generation) with ControlNet's structural control, IP-Adapter's style transfer, and professional-grade upscaling into a complete commercial-grade AI image production system.

Key takeaways:

  • 4-stage pipeline: Generation → Structural Control → Style Transfer → Upscaling
  • VRAM is the primary bottleneck; FP8 quantization and model offloading are essential
  • Parameter tuning requires experimentation; no universal configuration exists
  • Start with simple pipelines, gradually increase complexity

ERNIE-Image's role in multi-model pipelines is the "content generation engine" — it understands complex prompts and generates high-quality content foundations, while other models provide fine-grained control and enhancement on top. This division of labor is the right direction for modern AI image generation.


Reference Resources:

ERNIE-Image Team