ERNIE-Image Training Data Pipeline Deep Dive: Qwen3 VLM Annotation, Swiss-Tournament Aesthetics, and MT-DMD Distillation
Summary: The ERNIE-Image technical report reveals its complete training data pipeline: from massive raw corpus to Qwen3 VLM automated annotation, from Swiss-tournament aesthetic evaluation to Flow Matching DPO, to multi-teacher distillation (MT-DMD). This article deconstructs the technical details of each stage, explaining how these design choices enable an 8B-parameter model to reach SOTA performance in open-source text-to-image generation.
1. Introduction: Data Pipeline as the Great Divider
In April 2026, Baidu published the ERNIE-Image technical report (arXiv: 2605.25347), revealing for the first time its complete training data pipeline and training strategy. Unlike teams that only release final models, the ERNIE-Image team detailed every stage from raw data collection to final alignment training.
Key finding: ERNIE-Image's success stems not from a single breakthrough, but from a carefully engineered pipeline — Qwen3 VLM automated annotation + Swiss-tournament aesthetic scoring + curriculum training + DPO alignment + MT-DMD distillation — working together.
2. Pre-training Pipeline: Bottom-Up Strategy
2.1 Raw Data Collection and Classification
ERNIE-Image's pre-training data starts from a massive raw corpus, processed through:
- Fine-grained visual classification: Raw data split into 10,000 visual categories
- VLM Captioning: Qwen3 VLM extracts structural descriptions and in-image text
- Aesthetic scoring: ERNIE-Image-Aes model scores each image
2.2 Qwen3 VLM: Automated Annotation Engine
Key technical choice: ERNIE-Image uses Qwen3 VLM as its caption model, not traditional CLIP or BLIP.
Why Qwen3 VLM?
| Feature | CLIP/BLIP | Qwen3 VLM |
|---|---|---|
| Text recognition | ❌ Weak | ✅ Strong (OCR) |
| Structural description | Basic | Detailed (position/relationship/attribute) |
| Long text generation | ❌ Limited | ✅ Complex descriptions |
| Multilingual | English-first | Chinese + English |
Annotation flow:
Raw image → Qwen3 VLM → Structured caption
↓
"An Asian man in a blue suit standing in a modern office,
floor-to-ceiling windows and city skyline in the background,
natural light from the left-front, smiling expression, 45-degree angle"
Key insight: ERNIE-Image's SOTA text rendering performance (LongTextBench 0.964) largely comes from Qwen3 VLM's precise in-image text annotation.
3. Aesthetic Scoring: Swiss-Tournament Annotation
3.1 Limitations of Traditional Methods
| Method | Problem |
|---|---|
| Likert scoring (1-5) | Annotator inconsistency, score drift |
| ELO pairwise comparison | High comparison overhead, expensive |
3.2 Swiss-Tournament Method
ERNIE-Image introduces a Swiss-system tournament, inspired by chess competitions:
Round 1: Image A vs B, C vs D, E vs F...
Round 2: Winners vs winners, losers vs losers
Round 3: Continue grouping by "win count"
...
Final: Each image gets an aesthetic score based on relative ranking
Advantages:
- ✅ Avoids Likert score drift
- ✅ 50%+ fewer comparisons than ELO
- ✅ Annotators pass aesthetic calibration tests
3.3 ERNIE-Image-Aes Model Performance
| Model | SRCC | PLCC |
|---|---|---|
| LAION AES | 0.2944 | 0.3138 |
| ArtiMuse | 0.4277 | 0.4704 |
| UniPercept | 0.4533 | 0.4748 |
| ERNIE-Image-Aes | 0.7445 | 0.7598 |
Bias correction: ERNIE-Image-Aes fixes known biases — LAION-Aes over-scores AI/anime/casual shots; ArtiMuse/UniPercept over-scores B&W/casual photos.
3.4 ERIA-1K Benchmark Distribution
| Category | Share |
|---|---|
| Photography | 49.28% |
| Illustration/Anime | 23.16% |
| Graphic Design/Posters | 11.14% |
| Mixed Web | 10.44% |
| Film Photography | 5.42% |
| Product/Collectible | 0.56% |
4. Hierarchical Sampling and Curriculum Training
4.1 Hierarchical Sampling
| Level | Sampling Basis | Purpose |
|---|---|---|
| Inter-category | Corpus size + aggregate aesthetic quality | Ensure all categories well-represented |
| Intra-category | Individual image aesthetic scores | Prioritize high-quality data |
4.2 Three-Stage Curriculum Training
Stage 1: 256×256 → Learn basic features and composition
↓
Stage 2: 512×512 → Refine textures and details
↓
Stage 3: 1024×1024 → Final high-resolution output
Key design: Each stage uses diverse aspect ratios (not just squares), improving layout flexibility.
5. Post-training Pipeline: Top-Down Strategy
5.1 Supervised Fine-Tuning (SFT)
Focus: High-demand domains (posters, game screenshots, portraits, anime, etc.)
Innovation: K2.5 VLM Caption Rewriting
SFT stage uses K2.5 VLM to rewrite captions into diverse user-style prompts:
| Style | Example |
|---|---|
| Keywords | "woman, blue, smile, studio lighting" |
| Natural language | "A smiling woman portrait, studio lighting" |
| Instructional | "Generate a woman portrait with studio lighting" |
| Compositional | "4K, woman portrait, blue tones, studio lighting, shallow depth" |
Purpose: Simulate real-world diverse prompt patterns for robustness.
5.2 DPO Alignment: Flow Matching + Anchor Losses
Flow Matching DPO: Combines Direct Preference Optimization with Flow Matching for aesthetic alignment.
Key innovation: Anchor Losses
Traditional DPO suffers from reward hacking — the model learns to "cheat" for high scores rather than truly improving aesthetics. ERNIE-Image introduces Anchor Losses:
Total loss = DPO preference loss + λ × Anchor loss
Anchor Losses constrain the model to stay close to high-quality reference data during preference optimization, preventing over-deviation.
5.3 Multi-Teacher Distillation (MT-DMD)
Background: Traditional DMD/DMDR uses single-teacher distillation, potentially causing capability drift.
MT-DMD design:
| Teacher | Specialty | Routing |
|---|---|---|
| Text rendering expert | Precise text generation | Text-related prompts |
| Digital art expert | Artistic styles | Style-related prompts |
| Spatial layout expert | Multi-object layout | Complex composition prompts |
Dynamic routing: Weights shift between teachers based on noise level and optimization objectives.
High noise → More layout teacher (learn macro structure)
Low noise → More text teacher (learn fine details)
6. Complete Pipeline Summary
Raw data
↓
[10,000 visual categories]
↓
[Qwen3 VLM Caption] → Structural description + text annotation
↓
[ERNIE-Image-Aes scoring] → Swiss-tournament annotation
↓
[Hierarchical sampling] → Inter + intra category weighting
↓
[Curriculum training] → 256→512→1024
↓
Pre-training complete (ERNIE-Image Base)
↓
[SFT] → K2.5 VLM caption rewriting
↓
[DPO + Anchor Losses] → Aesthetic alignment
↓
[MT-DMD] → Multi-teacher distillation
↓
ERNIE-Image-Turbo (8 steps)
7. Real-World Impact of Pipeline Design
7.1 Closing the Gap with Closed Models
| Model | Total HP | Spatial | Knowledge | Aesthetic |
|---|---|---|---|---|
| Nano Banana 2.0 (closed) | 5.39 | 95.54 | 99.40 | 91.37 |
| ERNIE-Image (open 8B) | 5.07 | 89.88 | 95.24 | 83.04 |
| Seedream 5.0 (closed) | 5.03 | 90.48 | 97.02 | 80.65 |
Key conclusion: ERNIE-Image achieves performance close to closed flagship models with only 8B parameters — a gap of just ~0.32 HP.
7.2 Text Rendering SOTA
| Benchmark | ERNIE-Image (w/o PE) | ERNIE-Image (w/ PE) |
|---|---|---|
| LongTextBench (EN) | 0.968 | 0.980 |
| LongTextBench (ZH) | 0.959 | 0.966 |
Both Chinese and English text rendering achieve 96%+ accuracy, far exceeding other open-source models.
8. Takeaways for the Community
8.1 Data Quality > Data Quantity
ERNIE-Image proves that carefully curated and annotated data matters more than massive raw data.
8.2 VLM Annotation is the Future
Using Qwen3 VLM for automated annotation ensures quality while controlling costs. Community users can apply the same approach — use open-source VLMs to generate high-quality captions for custom datasets.
8.3 Aesthetic Evaluation Needs De-biasing
ERNIE-Image-Aes' bias correction reminds us: general-purpose aesthetic models have systematic biases. Domain-specific fine-tuning matters.
8.4 Multi-Teacher > Single-Teacher
MT-DMD's design — different experts for different capabilities — may become the standard paradigm for future distillation training.
9. Conclusion
ERNIE-Image's training pipeline is systems engineering: from data collection to annotation, from aesthetic evaluation to alignment training, from curriculum learning to multi-teacher distillation — every stage is carefully designed.
Core formula:
Quality data × VLM annotation × Aesthetic filtering × Curriculum training × DPO alignment × Multi-teacher distillation = 8B SOTA
For researchers wanting to reproduce or improve this pipeline, the technical report provides sufficient detail. For end users, understanding this pipeline helps use ERNIE-Image better — understanding what the model excels at means writing better prompts.
Further reading: ERNIE-Image Technical Report, ERNIE-Image-Aes Aesthetic Evaluation, DPO Alignment Training
This article is based on the ERNIE-Image technical report (arXiv: 2605.25347v1). All data from officially published materials.