ERNIE-Image Multi-Resolution Training Pipeline and Aspect Ratio Curriculum Deep Dive

Jun 17, 2026

ERNIE-Image Multi-Resolution Training Pipeline and Aspect Ratio Curriculum Deep Dive

In the AI image generation space, most people focus on model parameter counts — 8B, 12B, 80B — but what truly determines output quality are the "invisible" training details. How does the model learn? From what resolution? How do aspect ratios vary? How is data filtered?

ERNIE-Image's technical report (arXiv 2605.25347) reveals Baidu's complete training pipeline design: three-stage resolution curriculum learning, 10,000-level visual classification, Qwen3 VLM auto-captioning, and ERNIE-Image-Aes aesthetic scoring-driven data filtering. Behind these design choices lies a core philosophy: precise data pipelines are more effective than blind parameter scaling.

This article dives deep into ERNIE-Image's training pipeline, revealing why an 8B parameter model can challenge competitors with several times more parameters across multiple benchmarks.

Why Resolution Curriculum Learning Matters

Imagine teaching a student to draw. Do you start with an 8K canvas, or begin with small compositions in a sketchbook?

Most open-source models adopt a "one-shot" training strategy — training directly at target resolution (e.g., 1024×1024). Simple and direct, but two problems emerge:

  1. Low optimization efficiency: Large resolution means bigger computation graphs and longer gradient propagation paths — early training converges slowly
  2. Missing semantic learning: The model wastes compute on detail-level optimization without fully learning "what to draw"

ERNIE-Image uses three-stage resolution curriculum learning, letting the model master image generation "from small to large."

Three-Stage Resolution Curriculum Learning

Stage 1: 256×256 — Rapid Semantic Learning

In the early phase, the model trains at 256×256. The goal isn't beautiful images — it's:

  • Building semantic mapping: Learning correspondences between concepts ("cat", "forest", "sunset") and visual representations
  • Fast convergence: Small resolution = fewer computations = more data seen in less time
  • Basic composition: Object positioning, foreground/background distinction

Technical details:

  • Larger batch sizes (same VRAM)
  • Higher gradient update frequency
  • Can use higher learning rates

Analogy: Like sketching — quickly capturing the rough outline and semantic content without detail precision.

Stage 2: 512×512 — Detail Refinement

After mastering basic semantics, the model upgrades to 512×512:

  • Texture and material refinement: Fur texture, metal reflection, water ripples
  • Intermediate composition: Spatial relationships in multi-object scenes, perspective
  • Color and lighting consistency: Cross-region light-shadow coordination

This is the "concept to texture" transition. The model starts understanding: two cats can look completely different — a Ragdoll vs. a British Shorthair.

Stage 3: 1024×1024 — Fine Output

The final stage trains at target resolution:

  • Pixel-level precision: Stroke details in text rendering, facial feature fine-tuning
  • Global consistency: Ensuring unified style, lighting, and color across the whole image
  • Complex scene handling: Multi-object, multi-layer, multi-lighting compositions

Analogy: Going from sketch to detailed illustration — filling in details stroke by stroke on an established composition.

Three-Stage Advantages

Dimension One-Shot (1024) Three-Stage Curriculum
Early convergence Slow (heavy compute) Fast (small resolution, rapid iteration)
Semantic learning depth Shallow (detail interference) Deep (semantics first, details later)
Optimization stability Lower (large gradient fluctuations) Higher (progressive optimization)
Final quality Data-dependent High data utilization efficiency

ERNIE-Image's report shows this curriculum approach achieves better optimization with the same data volume. 8B parameters + efficient pipeline = benchmark performance surpassing 12B+ models.

Aspect Ratio Diversity: Beyond Square

Most text-to-image models default to square (1:1) training. ERNIE-Image introduces multiple aspect ratios during training:

Aspect Ratio Shape Use Case
1:1 Square Avatars, icons, product photos
3:4 Vertical Phone wallpapers, social media vertical
4:3 Horizontal Traditional photography, PPT slides
16:9 Widescreen Landscapes, video covers
2:3 Tall Magazine covers, e-commerce details
9:16 Portrait Short video covers, Stories format

Why Aspect Ratio Diversity Matters

  1. Prevents square overfitting: Training only at 1:1 significantly degrades non-square composition ability
  2. Improves layout flexibility: Different ratios require different composition strategies — widescreen emphasizes horizontal extension, portrait emphasizes vertical layering
  3. Matches real usage: Users don't always need squares — posters, banners, and wallpapers have specific aspect ratio requirements

ERNIE-Image supports multiple aspect ratios at generation time, so you can specify target ratios directly via API or local deployment without post-cropping.

Data Pipeline: From Raw Images to Training Samples

The core of a training pipeline is transforming massive raw images into high-quality training data. ERNIE-Image uses a complete five-stage pipeline:

Stage 1: Fine-Grained Classification — 10,000 Visual Categories

Traditional methods classify images into dozens to hundreds of broad categories ("people", "landscape", "animals"). ERNIE-Image goes further:

  • 10,000 fine-grained categories: Not "cat" but "Ragdoll", "British Shorthair", "Dragon Li"
  • Long-tail concept preservation: Prevents dominant categories from overwhelming rare ones
  • Balanced category sampling: Appearance frequency in training is carefully adjusted

Stage 2: Qwen3 VLM Auto-Captioning

Manually annotating 10,000 categories isn't feasible. ERNIE-Image fine-tuned Qwen3 (a powerful multimodal language model) as an auto-captioning engine:

  • Structured description extraction: Not just "this is a cat" but "an orange Dragon Li cat sitting on a green sofa, looking out the window, afternoon sunlight on its fur"
  • Text content recognition: For images containing text, specific text content is annotated (e.g., billboard slogans)
  • Multilingual captioning: Supports Chinese, English, and Japanese description generation

Stage 3: ERNIE-Image-Aes Aesthetic Scoring

Not all images are suitable for training. ERNIE-Image-Aes (an 8B-parameter vision-language model) scores each image aesthetically:

  • SRCC: 0.7445 / PLCC: 0.7598 (on ERIA-1K benchmark)
  • Eliminates systematic bias: Won't over-favor AI-generated images, B&W photos, or casual snapshots
  • Data filtering: Images below the aesthetic threshold are excluded

This sets a "quality gate" for training data — only high-aesthetic-score images enter the training set.

Stage 4: Hierarchical Sampling

Inter-category sampling:

  • Balances corpus size and aggregate aesthetic quality
  • Large categories aren't oversampled, small categories aren't ignored

Intra-category sampling:

  • Within a category, higher-quality instances are sampled more frequently
  • Quality determined by ERNIE-Image-Aes score

Stage 5: SFT Domain Fine-Tuning

After pre-training, the model enters SFT (supervised fine-tuning), fine-tuning for high-demand domains:

  • Target domains: Poster design, game art, portrait photography, product photography
  • K2.5 VLM caption rewriting: Technical descriptions rewritten into natural user styles
    • Keyword style: "cyberpunk, neon, rain night, future city"
    • Natural language: "A future city in the rain, neon lights reflecting colorful light in puddles"
    • Instruction style: "Please generate a cyberpunk-style night scene with neon lights and rain effects"

DPO Alignment Training: Teaching the Model "Good vs Bad"

DPO (Direct Preference Optimization) is a critical step in ERNIE-Image's pipeline. But ERNIE-Image's DPO isn't a simple copy from LLM methods — it's redesigned for the Flow Matching paradigm.

Flow Matching DPO Core Formula

Traditional DPO compares reward differences between two outputs. ERNIE-Image reformulates it as velocity-field L2 reconstruction error:

Diff_policy = ||v_pol(x_win) - v_win||² - ||v_pol(x_lose) - v_lose||²

Where:

  • v_pol is the policy model's predicted velocity field
  • x_win / x_lose are preferred positive and negative samples
  • v_win / v_lose are ground truth velocity fields

Anchor Loss: Preventing Reward Hacking

Pure DPO training can lead to "reward hacking" — the model learns to optimize reward scores while ignoring true quality. ERNIE-Image introduces anchor loss:

L_total = L_DPO + λ_win * E[L_policy_win] + λ_lose * E[L_policy_lose]

Parameters: β=0.05, λ_win=0.35, λ_lose=0.15

This ensures the model maintains good reconstruction capability while optimizing preferences — it won't sacrifice detail for the sake of looking "good."

MT-DMD Multi-Teacher Distillation: Turbo's Secret

ERNIE-Image Turbo (8-step generation) relies on MT-DMD (Multi-Teacher Distribution Matching Distillation).

The Single-Teacher Problem

Traditional knowledge distillation uses one teacher model to guide one student model. But in image generation, different domains need different expertise — a "text rendering expert" and a "landscape rendering expert" have completely different focuses.

MT-DMD Solution

MT-DMD uses a committee of domain experts {E_1, ..., E_K}:

  • Text expert: Optimizes text rendering and OCR accuracy
  • Art expert: Optimizes color harmony and compositional beauty
  • Layout expert: Optimizes spatial relationships in multi-element layouts
  • Portrait expert: Optimizes facial features and expression naturalness

These experts are weighted through a dynamic routing manifold W — automatically selecting the most relevant expert based on the current latent state, noise scale, and semantic condition.

Practical Results

MT-DMD enables ERNIE-Image Turbo to maintain quality close to the 50-step Standard with only 8 inference steps:

  • 6x+ speed improvement: From ~15 seconds to ~2-3 seconds
  • Minimal quality loss: Visually indistinguishable in most scenarios
  • Cost reduction: Same API price but much faster

Lessons for Developers and Researchers

ERNIE-Image's pipeline design reveals key engineering principles:

1. Data Quality > Data Quantity

  • 10,000 fine-grained categories + aesthetic filtering = high-quality training set
  • More effective than simply accumulating data volume

2. Curriculum Learning > One-Shot

  • Three-stage resolution progression = more efficient optimization
  • Analogy: Learn outlines first, then details

3. Multi-Teacher > Single-Teacher

  • Different domains need different expertise
  • Dynamic routing enables seamless handoff

4. Anchor Constraints > Pure Reward Optimization

  • Prevents reward hacking
  • Balances reconstruction quality with preference optimization

How to Apply These Techniques Locally?

While you can't fully replicate Baidu's pipeline, these techniques are applicable:

Resolution Curriculum (Simplified)

# Train at small resolution first, then upgrade
pipe = ErnieImagePipeline.from_pretrained("baidu/ERNIE-Image")

Step 1: 512x512 rapid iteration

for i in range(100):
image = pipe(prompt, height=512, width=512).images[0]

Step 2: 1024x1024 fine output

image = pipe(prompt, height=1024, width=1024).images[0]

Multi-Aspect-Ratio Generation

# ERNIE-Image natively supports multiple aspect ratios
sizes = [(1024, 1024), (768, 1024), (1024, 768), (1280, 720)]
for h, w in sizes:
    image = pipe(prompt, height=h, width=w).images[0]
    image.save(f"output_{h}x{w}.png")

Aesthetic Score Filtering

# Use ERNIE-Image-Aes to score generated results
from ernie_image_aes import AesModel

aes = AesModel.from_pretrained("baidu/ERNIE-Image-Aes")
score = aes.score(image)
if score > 0.7: # Keep only high-quality results
image.save("high_quality.png")

Summary

ERNIE-Image's training pipeline design reflects Baidu's engineering depth in AI image generation. From three-stage resolution curriculum to 10,000 fine-grained categories, from Qwen3 VLM auto-captioning to MT-DMD multi-teacher distillation — every design choice serves one goal: achieve better quality with fewer parameters.

Key takeaways:

  1. Resolution curriculum learning lets the model learn semantics before details
  2. Aspect ratio diversity prevents square overfitting and improves layout flexibility
  3. Aesthetic score filtering ensures high-quality training data
  4. DPO anchor loss prevents reward hacking while maintaining reconstruction quality
  5. MT-DMD multi-teacher distillation enables Turbo to reach near-50-step quality in 8 steps

For researchers and developers, ERNIE-Image's pipeline provides a valuable reference — with limited resources, precise data pipelines and training strategies are more effective than blind parameter scaling.


Further Reading:

ERNIE-Image Team