ERNIE-Image Training Data Pipeline Deep Dive: Qwen3 VLM Annotation, Swiss-Tournament Aesthetics, and MT-DMD Distillation

Jun 9, 2026

ERNIE-Image Training Data Pipeline Deep Dive: Qwen3 VLM Annotation, Swiss-Tournament Aesthetics, and MT-DMD Distillation

Summary: The ERNIE-Image technical report reveals its complete training data pipeline: from massive raw corpus to Qwen3 VLM automated annotation, from Swiss-tournament aesthetic evaluation to Flow Matching DPO, to multi-teacher distillation (MT-DMD). This article deconstructs the technical details of each stage, explaining how these design choices enable an 8B-parameter model to reach SOTA performance in open-source text-to-image generation.


1. Introduction: Data Pipeline as the Great Divider

In April 2026, Baidu published the ERNIE-Image technical report (arXiv: 2605.25347), revealing for the first time its complete training data pipeline and training strategy. Unlike teams that only release final models, the ERNIE-Image team detailed every stage from raw data collection to final alignment training.

Key finding: ERNIE-Image's success stems not from a single breakthrough, but from a carefully engineered pipeline — Qwen3 VLM automated annotation + Swiss-tournament aesthetic scoring + curriculum training + DPO alignment + MT-DMD distillation — working together.


2. Pre-training Pipeline: Bottom-Up Strategy

2.1 Raw Data Collection and Classification

ERNIE-Image's pre-training data starts from a massive raw corpus, processed through:

  1. Fine-grained visual classification: Raw data split into 10,000 visual categories
  2. VLM Captioning: Qwen3 VLM extracts structural descriptions and in-image text
  3. Aesthetic scoring: ERNIE-Image-Aes model scores each image

2.2 Qwen3 VLM: Automated Annotation Engine

Key technical choice: ERNIE-Image uses Qwen3 VLM as its caption model, not traditional CLIP or BLIP.

Why Qwen3 VLM?

Feature CLIP/BLIP Qwen3 VLM
Text recognition ❌ Weak ✅ Strong (OCR)
Structural description Basic Detailed (position/relationship/attribute)
Long text generation ❌ Limited ✅ Complex descriptions
Multilingual English-first Chinese + English

Annotation flow:

Raw image → Qwen3 VLM → Structured caption
                              ↓
        "An Asian man in a blue suit standing in a modern office,
         floor-to-ceiling windows and city skyline in the background,
         natural light from the left-front, smiling expression, 45-degree angle"

Key insight: ERNIE-Image's SOTA text rendering performance (LongTextBench 0.964) largely comes from Qwen3 VLM's precise in-image text annotation.


3. Aesthetic Scoring: Swiss-Tournament Annotation

3.1 Limitations of Traditional Methods

Method Problem
Likert scoring (1-5) Annotator inconsistency, score drift
ELO pairwise comparison High comparison overhead, expensive

3.2 Swiss-Tournament Method

ERNIE-Image introduces a Swiss-system tournament, inspired by chess competitions:

Round 1: Image A vs B, C vs D, E vs F...
Round 2: Winners vs winners, losers vs losers
Round 3: Continue grouping by "win count"
...
Final: Each image gets an aesthetic score based on relative ranking

Advantages:

  • ✅ Avoids Likert score drift
  • ✅ 50%+ fewer comparisons than ELO
  • ✅ Annotators pass aesthetic calibration tests

3.3 ERNIE-Image-Aes Model Performance

Model SRCC PLCC
LAION AES 0.2944 0.3138
ArtiMuse 0.4277 0.4704
UniPercept 0.4533 0.4748
ERNIE-Image-Aes 0.7445 0.7598

Bias correction: ERNIE-Image-Aes fixes known biases — LAION-Aes over-scores AI/anime/casual shots; ArtiMuse/UniPercept over-scores B&W/casual photos.

3.4 ERIA-1K Benchmark Distribution

Category Share
Photography 49.28%
Illustration/Anime 23.16%
Graphic Design/Posters 11.14%
Mixed Web 10.44%
Film Photography 5.42%
Product/Collectible 0.56%

4. Hierarchical Sampling and Curriculum Training

4.1 Hierarchical Sampling

Level Sampling Basis Purpose
Inter-category Corpus size + aggregate aesthetic quality Ensure all categories well-represented
Intra-category Individual image aesthetic scores Prioritize high-quality data

4.2 Three-Stage Curriculum Training

Stage 1: 256×256 → Learn basic features and composition
         ↓
Stage 2: 512×512 → Refine textures and details
         ↓
Stage 3: 1024×1024 → Final high-resolution output

Key design: Each stage uses diverse aspect ratios (not just squares), improving layout flexibility.


5. Post-training Pipeline: Top-Down Strategy

5.1 Supervised Fine-Tuning (SFT)

Focus: High-demand domains (posters, game screenshots, portraits, anime, etc.)

Innovation: K2.5 VLM Caption Rewriting

SFT stage uses K2.5 VLM to rewrite captions into diverse user-style prompts:

Style Example
Keywords "woman, blue, smile, studio lighting"
Natural language "A smiling woman portrait, studio lighting"
Instructional "Generate a woman portrait with studio lighting"
Compositional "4K, woman portrait, blue tones, studio lighting, shallow depth"

Purpose: Simulate real-world diverse prompt patterns for robustness.

5.2 DPO Alignment: Flow Matching + Anchor Losses

Flow Matching DPO: Combines Direct Preference Optimization with Flow Matching for aesthetic alignment.

Key innovation: Anchor Losses

Traditional DPO suffers from reward hacking — the model learns to "cheat" for high scores rather than truly improving aesthetics. ERNIE-Image introduces Anchor Losses:

Total loss = DPO preference loss + λ × Anchor loss

Anchor Losses constrain the model to stay close to high-quality reference data during preference optimization, preventing over-deviation.

5.3 Multi-Teacher Distillation (MT-DMD)

Background: Traditional DMD/DMDR uses single-teacher distillation, potentially causing capability drift.

MT-DMD design:

Teacher Specialty Routing
Text rendering expert Precise text generation Text-related prompts
Digital art expert Artistic styles Style-related prompts
Spatial layout expert Multi-object layout Complex composition prompts

Dynamic routing: Weights shift between teachers based on noise level and optimization objectives.

High noise → More layout teacher (learn macro structure)
Low noise → More text teacher (learn fine details)

6. Complete Pipeline Summary

Raw data
  ↓
[10,000 visual categories]
  ↓
[Qwen3 VLM Caption] → Structural description + text annotation
  ↓
[ERNIE-Image-Aes scoring] → Swiss-tournament annotation
  ↓
[Hierarchical sampling] → Inter + intra category weighting
  ↓
[Curriculum training] → 256→512→1024
  ↓
Pre-training complete (ERNIE-Image Base)
  ↓
[SFT] → K2.5 VLM caption rewriting
  ↓
[DPO + Anchor Losses] → Aesthetic alignment
  ↓
[MT-DMD] → Multi-teacher distillation
  ↓
ERNIE-Image-Turbo (8 steps)

7. Real-World Impact of Pipeline Design

7.1 Closing the Gap with Closed Models

Model Total HP Spatial Knowledge Aesthetic
Nano Banana 2.0 (closed) 5.39 95.54 99.40 91.37
ERNIE-Image (open 8B) 5.07 89.88 95.24 83.04
Seedream 5.0 (closed) 5.03 90.48 97.02 80.65

Key conclusion: ERNIE-Image achieves performance close to closed flagship models with only 8B parameters — a gap of just ~0.32 HP.

7.2 Text Rendering SOTA

Benchmark ERNIE-Image (w/o PE) ERNIE-Image (w/ PE)
LongTextBench (EN) 0.968 0.980
LongTextBench (ZH) 0.959 0.966

Both Chinese and English text rendering achieve 96%+ accuracy, far exceeding other open-source models.


8. Takeaways for the Community

8.1 Data Quality > Data Quantity

ERNIE-Image proves that carefully curated and annotated data matters more than massive raw data.

8.2 VLM Annotation is the Future

Using Qwen3 VLM for automated annotation ensures quality while controlling costs. Community users can apply the same approach — use open-source VLMs to generate high-quality captions for custom datasets.

8.3 Aesthetic Evaluation Needs De-biasing

ERNIE-Image-Aes' bias correction reminds us: general-purpose aesthetic models have systematic biases. Domain-specific fine-tuning matters.

8.4 Multi-Teacher > Single-Teacher

MT-DMD's design — different experts for different capabilities — may become the standard paradigm for future distillation training.


9. Conclusion

ERNIE-Image's training pipeline is systems engineering: from data collection to annotation, from aesthetic evaluation to alignment training, from curriculum learning to multi-teacher distillation — every stage is carefully designed.

Core formula:

Quality data × VLM annotation × Aesthetic filtering × Curriculum training × DPO alignment × Multi-teacher distillation = 8B SOTA

For researchers wanting to reproduce or improve this pipeline, the technical report provides sufficient detail. For end users, understanding this pipeline helps use ERNIE-Image better — understanding what the model excels at means writing better prompts.

Further reading: ERNIE-Image Technical Report, ERNIE-Image-Aes Aesthetic Evaluation, DPO Alignment Training


This article is based on the ERNIE-Image technical report (arXiv: 2605.25347v1). All data from officially published materials.

ERNIE-Image Team