Multi-Modal Agent with ERNIE-Image: Building a Fully Automated Text-to-Image-to-Video Pipeline
Abstract: In the AI Agent era, a single image generation model is no longer enough. This article introduces how to build a multi-modal Agent workflow based on ERNIE-Image — from user natural language input, automatically through prompt enhancement, image generation, aesthetic scoring, image-to-video conversion, and final professional-grade output. The entire pipeline uses open-source models and can be fully deployed locally.
Background: Why Do We Need Multi-Modal Agents?
AI image generation in 2026 is no longer a simple "input prompt, output image" process. Professional users need a complete automated pipeline:
- User describes needs in natural language → 2. AI automatically enhances the prompt → 3. Image generation → 4. Aesthetic scoring → 5. Image-to-video conversion → 6. Final output
The ERNIE-Image ecosystem is naturally suited for building such a pipeline — it already includes a Prompt Enhancer (3B), an aesthetic evaluation model (ERNIE-Image-Aes, 8B), and the main generation model (8B). Combined with ComfyUI's visual orchestration capabilities, you can build a complete multi-modal Agent system.
System Architecture: Four-Stage Pipeline
Stage 1: Prompt Enhancement (Prompt Enhancer 3B)
User input: "A dog running on the moon"
Prompt Enhancer automatically expands to: "A golden retriever running across the lunar surface, dust kicking up, Earth visible in the background, cinematic lighting, 8K ultra HD, HDR, motion blur, sci-fi film style"
This is a unique element in the ERNIE-Image ecosystem — most image generation models either require users to manually write enhanced prompts or rely on external LLMs. ERNIE-Image includes a built-in Prompt Enhancer fine-tuned from Ministral 3B, specifically optimized for image generation tasks.
Stage 2: Image Generation (ERNIE-Image 8B / Turbo)
Using the enhanced prompt, generate images through ERNIE-Image Base (50 steps, high quality) or Turbo (8 steps, fast).
Key parameter configuration:
- Sampler: dpmpp_2s_ancestral (community-verified best sampler)
- CFG Scale: 5.0-7.0 (medium guidance strength)
- Steps: 50 (Base) or 8 (Turbo)
- Resolution: 1024×1024 or custom aspect ratio based on needs
Stage 3: Aesthetic Scoring (ERNIE-Image-Aes 8B)
Generated images are scored by the ERNIE-Image-Aes aesthetic evaluation model. This model achieves SRCC 0.7445 and PLCC 0.7598 on the ERIA-1K benchmark, far surpassing previous open-source aesthetic evaluation models.
Scoring strategy:
- Score ≥ 7.0: Keep for downstream processing
- Score 5.0-7.0: Backup, may need prompt fine-tuning and regeneration
- Score < 5.0: Discard, adjust prompt and retry
Stage 4: Image-to-Video Conversion (Wan 2.7 / LTX 2.3)
Convert static images to dynamic video through ComfyUI's image-to-video nodes. Currently supported options:
- Wan 2.7: High-quality image-to-video, supports 1080P output
- LTX 2.3 Sulphur: Fast video generation, suitable for real-time preview
- Hunyuan3D 2.0: If 3D textured models are needed
Practical ComfyUI Workflow Setup
Environment Setup
# 1. Install ComfyUI
git clone https://github.com/comfyanonymous/ComfyUI
cd ComfyUI
pip install -r requirements.txt
2. Download ERNIE-Image models
Base model: https://huggingface.co/baidu/ERNIE-Image
Turbo model: https://huggingface.co/baidu/ERNIE-Image-Turbo
3. Install necessary custom nodes
pip install comfyui-easy-install
Or manually install needed nodes
Node Connection Logic
[User Text Input]
↓
[CLIP Text Encode (PE Enhancer)] → Enhanced prompt
↓
[ERNIE-Image DiT Sampler] → Image output
↓
[ERNIE-Image-Aes Scorer] → Aesthetic score
↓
[Conditional: Score ≥ 7.0?]
↓ Yes
[Wan 2.7 Image-to-Video] → Video output
↓
[Final Result Save]
Key Configuration Points
VRAM Management: Loading multiple models simultaneously (PE 3B + DiT 8B + Aes 8B + Video model) requires 48GB+ VRAM. Consider using SGLang for independent PE deployment, or use FP8/GGUF quantized versions.
Batch Processing: ComfyUI's Batch mode supports generating multiple images at once, with aesthetic scoring automatically filtering the best results.
Error Handling: When aesthetic scores fall below the threshold, automatically loop back to modify the prompt and regenerate.
Practical Application Cases
Case 1: Automated E-commerce Product Video Generation
Need: E-commerce sellers need promotional videos for new products
Workflow:
- Input product name and selling points: "Smart watch, black dial, sporty style"
- PE auto-enhances the prompt with lighting and scene descriptions
- ERNIE-Image generates product photos (multiple angles: front, side, wearing effect)
- ERNIE-Image-Aes selects the best angles
- Wan 2.7 converts product photos to 10-second product showcase videos
- Output ready-to-use promotional materials for e-commerce platforms
Cost Comparison:
- Traditional (designer + photographer + video editor): ¥5,000-20,000/product
- ERNIE-Image Agent: ¥0 (local deployment) or ≈ ¥50/product (cloud API)
Case 2: Social Media Content Batch Production
Need: Content creators need customized visual content for different platforms
Workflow:
- Input core theme: "AI technology changes lives"
- PE generates different-style prompts for different platforms
- ERNIE-Image batch-generates multi-size images (1:1 Instagram, 16:9 YouTube, 9:16 TikTok)
- Each image includes platform-specific text (titles, tags)
- Best results converted to short video via Wan 2.7
- Auto-outputs platform-ready formats
Efficiency Gain: From 2 days of work down to 30 minutes
Case 3: Academic Research Visualization
Need: Researchers need to convert complex data into intuitive charts
Workflow:
- Input data description and visualization requirements
- PE enhances the prompt with specific chart types and color schemes
- ERNIE-Image generates publication-quality charts
- ERNIE-Image-Aes ensures visual quality
- Output charts ready for academic papers
Advantage: ERNIE-Image's text rendering capability makes it particularly suitable for generating labeled scientific charts.
Deployment Option Comparison
| Option | VRAM Required | Cost | Use Case |
|---|---|---|---|
| Local A100 80GB | 80GB | Owned GPU | Professional batch production |
| Local RTX 5090 | 32GB (FP8) | Owned GPU | Individual creators |
| Google Colab Free | 16GB (T4) | ¥0 | Learning and small-scale use |
| RunPod A100 | 40-80GB | ≈ $0.5-2/hour | Medium-scale production |
| SiliconFlow API | No GPU | ≈ ¥0.11/image | Light usage |
| Civitai API | No GPU | Buzz billing | On-demand use |
Common Issues and Solutions
Issue 1: Insufficient VRAM
Solutions:
- Use FP8 quantization (~50% VRAM reduction)
- Use GGUF format (supports Q4/Q8 quantization)
- Deploy PE independently via SGLang, freeing main GPU memory
- Use ComfyUI's "Force Dispatch" node for dynamic model loading
Issue 2: Aesthetic Scoring Too High/Low
Solutions:
- ERNIE-Image-Aes scores "photography style" images higher, "illustration/anime" lower
- Adjust thresholds by style category:
- Photorealistic: ≥ 7.0
- Illustration/Anime: ≥ 5.5
- Abstract Art: ≥ 5.0
Issue 3: Poor Video Quality
Solutions:
- Input image resolution affects video output quality — ensure ≥ 1024×1024
- Wan 2.7 works better with "simple scenes"; complex multi-element scenes may produce artifacts
- Use LTX 2.3 for quick preview, confirm direction, then use Wan 2.7 for final rendering
Conclusion
The ERNIE-Image multi-modal Agent workflow represents the future direction of AI content generation: from a single model to a complete pipeline, from manual operation to automated production.
Core advantages:
- Fully open-source: All components can be deployed locally, no API restrictions
- Cost-effective: Near-zero cost for local deployment (with existing GPU)
- Native Chinese support: Full Chinese support from prompt enhancement to text rendering
- Flexible composition: Each stage can be independently replaced or upgraded
With the official release of the ERNIE-Image editing model and more community nodes coming, this pipeline will continue to improve. For teams and individuals who need to batch-produce professional-grade visual content, now is the best time to start building.
Sources: ERNIE-Image official documentation, ComfyUI official templates, Pixaroma tutorial videos, community testing feedback.