Multi-Modal Agent with ERNIE-Image: Building a Fully Automated Text-to-Image-to-Video Pipeline

Jul 6, 2026

Multi-Modal Agent with ERNIE-Image: Building a Fully Automated Text-to-Image-to-Video Pipeline

Abstract: In the AI Agent era, a single image generation model is no longer enough. This article introduces how to build a multi-modal Agent workflow based on ERNIE-Image — from user natural language input, automatically through prompt enhancement, image generation, aesthetic scoring, image-to-video conversion, and final professional-grade output. The entire pipeline uses open-source models and can be fully deployed locally.

Background: Why Do We Need Multi-Modal Agents?

AI image generation in 2026 is no longer a simple "input prompt, output image" process. Professional users need a complete automated pipeline:

  1. User describes needs in natural language → 2. AI automatically enhances the prompt → 3. Image generation → 4. Aesthetic scoring → 5. Image-to-video conversion → 6. Final output

The ERNIE-Image ecosystem is naturally suited for building such a pipeline — it already includes a Prompt Enhancer (3B), an aesthetic evaluation model (ERNIE-Image-Aes, 8B), and the main generation model (8B). Combined with ComfyUI's visual orchestration capabilities, you can build a complete multi-modal Agent system.

System Architecture: Four-Stage Pipeline

Stage 1: Prompt Enhancement (Prompt Enhancer 3B)

User input: "A dog running on the moon"

Prompt Enhancer automatically expands to: "A golden retriever running across the lunar surface, dust kicking up, Earth visible in the background, cinematic lighting, 8K ultra HD, HDR, motion blur, sci-fi film style"

This is a unique element in the ERNIE-Image ecosystem — most image generation models either require users to manually write enhanced prompts or rely on external LLMs. ERNIE-Image includes a built-in Prompt Enhancer fine-tuned from Ministral 3B, specifically optimized for image generation tasks.

Stage 2: Image Generation (ERNIE-Image 8B / Turbo)

Using the enhanced prompt, generate images through ERNIE-Image Base (50 steps, high quality) or Turbo (8 steps, fast).

Key parameter configuration:

  • Sampler: dpmpp_2s_ancestral (community-verified best sampler)
  • CFG Scale: 5.0-7.0 (medium guidance strength)
  • Steps: 50 (Base) or 8 (Turbo)
  • Resolution: 1024×1024 or custom aspect ratio based on needs

Stage 3: Aesthetic Scoring (ERNIE-Image-Aes 8B)

Generated images are scored by the ERNIE-Image-Aes aesthetic evaluation model. This model achieves SRCC 0.7445 and PLCC 0.7598 on the ERIA-1K benchmark, far surpassing previous open-source aesthetic evaluation models.

Scoring strategy:

  • Score ≥ 7.0: Keep for downstream processing
  • Score 5.0-7.0: Backup, may need prompt fine-tuning and regeneration
  • Score < 5.0: Discard, adjust prompt and retry

Stage 4: Image-to-Video Conversion (Wan 2.7 / LTX 2.3)

Convert static images to dynamic video through ComfyUI's image-to-video nodes. Currently supported options:

  • Wan 2.7: High-quality image-to-video, supports 1080P output
  • LTX 2.3 Sulphur: Fast video generation, suitable for real-time preview
  • Hunyuan3D 2.0: If 3D textured models are needed

Practical ComfyUI Workflow Setup

Environment Setup

# 1. Install ComfyUI
git clone https://github.com/comfyanonymous/ComfyUI
cd ComfyUI
pip install -r requirements.txt

2. Download ERNIE-Image models

Base model: https://huggingface.co/baidu/ERNIE-Image

Turbo model: https://huggingface.co/baidu/ERNIE-Image-Turbo

3. Install necessary custom nodes

pip install comfyui-easy-install

Or manually install needed nodes

Node Connection Logic

[User Text Input]
      ↓
[CLIP Text Encode (PE Enhancer)] → Enhanced prompt
      ↓
[ERNIE-Image DiT Sampler] → Image output
      ↓
[ERNIE-Image-Aes Scorer] → Aesthetic score
      ↓
[Conditional: Score ≥ 7.0?]
      ↓ Yes
[Wan 2.7 Image-to-Video] → Video output
      ↓
[Final Result Save]

Key Configuration Points

  1. VRAM Management: Loading multiple models simultaneously (PE 3B + DiT 8B + Aes 8B + Video model) requires 48GB+ VRAM. Consider using SGLang for independent PE deployment, or use FP8/GGUF quantized versions.

  2. Batch Processing: ComfyUI's Batch mode supports generating multiple images at once, with aesthetic scoring automatically filtering the best results.

  3. Error Handling: When aesthetic scores fall below the threshold, automatically loop back to modify the prompt and regenerate.

Practical Application Cases

Case 1: Automated E-commerce Product Video Generation

Need: E-commerce sellers need promotional videos for new products

Workflow:

  1. Input product name and selling points: "Smart watch, black dial, sporty style"
  2. PE auto-enhances the prompt with lighting and scene descriptions
  3. ERNIE-Image generates product photos (multiple angles: front, side, wearing effect)
  4. ERNIE-Image-Aes selects the best angles
  5. Wan 2.7 converts product photos to 10-second product showcase videos
  6. Output ready-to-use promotional materials for e-commerce platforms

Cost Comparison:

  • Traditional (designer + photographer + video editor): ¥5,000-20,000/product
  • ERNIE-Image Agent: ¥0 (local deployment) or ≈ ¥50/product (cloud API)

Case 2: Social Media Content Batch Production

Need: Content creators need customized visual content for different platforms

Workflow:

  1. Input core theme: "AI technology changes lives"
  2. PE generates different-style prompts for different platforms
  3. ERNIE-Image batch-generates multi-size images (1:1 Instagram, 16:9 YouTube, 9:16 TikTok)
  4. Each image includes platform-specific text (titles, tags)
  5. Best results converted to short video via Wan 2.7
  6. Auto-outputs platform-ready formats

Efficiency Gain: From 2 days of work down to 30 minutes

Case 3: Academic Research Visualization

Need: Researchers need to convert complex data into intuitive charts

Workflow:

  1. Input data description and visualization requirements
  2. PE enhances the prompt with specific chart types and color schemes
  3. ERNIE-Image generates publication-quality charts
  4. ERNIE-Image-Aes ensures visual quality
  5. Output charts ready for academic papers

Advantage: ERNIE-Image's text rendering capability makes it particularly suitable for generating labeled scientific charts.

Deployment Option Comparison

Option VRAM Required Cost Use Case
Local A100 80GB 80GB Owned GPU Professional batch production
Local RTX 5090 32GB (FP8) Owned GPU Individual creators
Google Colab Free 16GB (T4) ¥0 Learning and small-scale use
RunPod A100 40-80GB ≈ $0.5-2/hour Medium-scale production
SiliconFlow API No GPU ≈ ¥0.11/image Light usage
Civitai API No GPU Buzz billing On-demand use

Common Issues and Solutions

Issue 1: Insufficient VRAM

Solutions:

  • Use FP8 quantization (~50% VRAM reduction)
  • Use GGUF format (supports Q4/Q8 quantization)
  • Deploy PE independently via SGLang, freeing main GPU memory
  • Use ComfyUI's "Force Dispatch" node for dynamic model loading

Issue 2: Aesthetic Scoring Too High/Low

Solutions:

  • ERNIE-Image-Aes scores "photography style" images higher, "illustration/anime" lower
  • Adjust thresholds by style category:
    • Photorealistic: ≥ 7.0
    • Illustration/Anime: ≥ 5.5
    • Abstract Art: ≥ 5.0

Issue 3: Poor Video Quality

Solutions:

  • Input image resolution affects video output quality — ensure ≥ 1024×1024
  • Wan 2.7 works better with "simple scenes"; complex multi-element scenes may produce artifacts
  • Use LTX 2.3 for quick preview, confirm direction, then use Wan 2.7 for final rendering

Conclusion

The ERNIE-Image multi-modal Agent workflow represents the future direction of AI content generation: from a single model to a complete pipeline, from manual operation to automated production.

Core advantages:

  • Fully open-source: All components can be deployed locally, no API restrictions
  • Cost-effective: Near-zero cost for local deployment (with existing GPU)
  • Native Chinese support: Full Chinese support from prompt enhancement to text rendering
  • Flexible composition: Each stage can be independently replaced or upgraded

With the official release of the ERNIE-Image editing model and more community nodes coming, this pipeline will continue to improve. For teams and individuals who need to batch-produce professional-grade visual content, now is the best time to start building.


Sources: ERNIE-Image official documentation, ComfyUI official templates, Pixaroma tutorial videos, community testing feedback.

ERNIE-Image Team