Baidu Open-Sources ERNIE-Image: Why Is This 8B Text-to-Image Model Ranked #1 in Open Source?
To be blunt, AI image generation stopped feeling novel a long time ago.
Midjourney makes beautiful images, Stable Diffusion is endlessly hackable, and FLUX is both fast and stable—it seems like everything you could want is already here. But there’s one problem almost everyone in design has run into: text inside AI-generated images is almost always garbled.
A poster is supposed to say “New Product Launch,” but what you get is a blob of alien symbols. A comic speech bubble should contain dialogue, but it comes out as unreadable scribbles. Mixed Chinese and English text? Forget it—it breaks immediately.
That changed yesterday (April 15), when Baidu officially open-sourced ERNIE-Image, an 8B-parameter text-to-image model that ranked first among open-source models across multiple public benchmarks. More importantly, its text rendering finally looks usable.
1. What Exactly Is ERNIE-Image?

ERNIE-Image is an open-source text-to-image model developed by Baidu’s Wenxin team. It is built on a single-stream Diffusion Transformer (DiT) architecture, with a core parameter count of just 8B (8 billion).
Put simply, it works like most AI image generators: you enter a text description, and it creates an image for you.
But it does have a few important differences.
💡 Core concept: ERNIE-Image = single-stream DiT (handles image generation) + Prompt Enhancer (interprets what you mean) + 8B parameters (compact, but highly capable)
Two versions:
| Version | Inference Steps | CFG | Features |
|---|---|---|---|
| ERNIE-Image | 50 steps | 4.0 | SFT version with higher quality and more precise instruction following |
| ERNIE-Image-Turbo | 8 steps | 1.0 | Distilled acceleration version, 6x faster, with stronger aesthetics |
The Turbo version uses both DMD (Distribution Matching Distillation) and RL (Reinforcement Learning) for acceleration, compressing the process from 50 steps down to 8 and improving speed by nearly 6x.
2. Why Is It Worth Paying Attention To?
Right now, the open-source text-to-image landscape looks roughly like this:
- FLUX.2-klein (9B): Strong overall capability, but weak at rendering Chinese text
- Z-Image (6B): Released by Alibaba Tongyi, strong in photorealistic style, but text details still have flaws
- Stable Diffusion 3.5 (8B): The most mature ecosystem, but text has always been a weak spot
- Qwen-Image (7B): Also from Alibaba, with strong semantic understanding, but only average text rendering
What makes ERNIE-Image stand out is this: it treats “controllability” as being just as important as “visual quality.”
💡 One-line summary: While other models focus on “making it look good,” ERNIE-Image is also focused on “making it correct.”
Its 6 core strengths:
- Precise text rendering: Handles mixed Chinese-English text, dense long-form text, and layout-sensitive text well
- Complex instruction following: Accurately executes multi-object, multi-relation, and knowledge-dense descriptions
- Structured image generation: Especially strong at posters, comic layouts, and multi-panel compositions
- Broad style coverage: Can handle realistic photography, graphic design, and cinematic looks
- Compact but powerful: Only 8B parameters, yet stronger than many larger models
- Consumer-grade deployment: Runs on 24GB VRAM; an RTX 4090 or RTX 3090 is enough
3. How Does It Pull This Off?
Let’s start with an analogy.
Imagine asking an artist to draw a poster. Most AI models are like an “impatient artist”—before you’ve even finished explaining your request, they’ve already started drawing, and the result doesn’t quite match what you had in mind.
ERNIE-Image is more like “an artist with a translator.”
Your Chinese instruction first goes through a Prompt Enhancer (PE), which acts like a translator, expanding it into a more detailed and structured description. Only then does the artist start working from that more complete “translated brief.”
Naturally, the result is much more accurate.
📖 Technical explanation: ERNIE-Image uses a single-stream DiT architecture that fuses text tokens, visual semantic tokens, and image VAE tokens into a unified sequence for processing. Combined with a lightweight Prompt Enhancer, it expands short user inputs into richer, more structured descriptions, improving generation quality.
About the Prompt Enhancer (PE):
This is one of ERNIE-Image’s standout design choices. It is a lightweight model responsible for automatically expanding something casual like “make a poster” into a professional-grade prompt that includes composition, color, text content, layout style, and other details.
Benchmark data shows that with PE enabled, the model improves noticeably on both text rendering (LongTextBench) and instruction following (GENEval).
4. How Good Are the Results?
The numbers speak for themselves. On four mainstream benchmark tests, ERNIE-Image performs as follows:
- GENEval (compositional generation capability): ERNIE-Image ranks first with an overall score of 0.8856, surpassing Qwen-Image (0.8683), FLUX.2-klein-9B (0.8481), and Z-Image (0.8400).
- LongTextBench (long-text rendering): ERNIE-Image (w/ PE) ranks second with 0.9733 (behind only the closed-source commercial model Seedream 4.5 at 0.9882), far ahead of FLUX.2-klein-9B (0.5413).
- OneIG-EN / OneIG-ZH (open-domain image generation): In both English and Chinese settings, ERNIE-Image remains firmly among the top open-source models.
⚠️ Note: FLUX.2-klein-9B scores only 0.2183 on the Chinese portion of LongTextBench, which means it is almost unusable for generating images containing long Chinese text. This is ERNIE-Image’s biggest advantage over FLUX.
5. How Does It Perform in Practice?
Based on official demos and third-party evaluations, ERNIE-Image performs especially well in the following scenarios:

- Posters and design layout: It doesn’t just generate “something that looks like a poster”—it can accurately place titles, subtitles, and body text while maintaining a sensible hierarchy and layout. That is quite rare among open-source models.
- Comics and storyboards: It doesn’t just draw characters; it can generate a complete comic page—with panel separation, coherent actions, and scene transitions—so it reads like a real comic rather than a pile of disconnected illustrations.
- Multi-panel visual storytelling: It can generate four-panel comics, comparison sets, and emotionally consistent sequences, with decent consistency between frames.
- Multilingual text rendering: It can accurately render both Chinese and English, including dense paragraphs, title layouts, annotation text, and comic dialogue. Mixed Chinese-English text also works.
- Cinematic texture and artistic style: It does more than produce the typical high-saturation “AI look”; it can also create muted cinematic tones, film grain, and mood-driven lighting with a stronger visual identity.
6. How Do You Use It?
Hardware requirements:
- A consumer GPU with 24GB VRAM (RTX 3090 / 4090 is enough)
- Support for bfloat16 precision
Supported resolutions: 1024×1024, 848×1264, 1264×848, 768×1376, 896×1200, 1376×768, 1200×896
Option 1: Diffusers (recommended for beginners)
import torch
from diffusers import ErnieImagePipeline
# Use ERNIE-Image (higher quality, 50 steps)
pipe = ErnieImagePipeline.from_pretrained(
"baidu/ERNIE-Image",
torch_dtype=torch.bfloat16,
).to("cuda")
image = pipe(
prompt="A black-and-white Chinese rural dog running across the grass",
height=1024,
width=1024,
num_inference_steps=50,
guidance_scale=4.0,
use_pe=True # Enable prompt enhancement
).images[0]
image.save("output.png")
Want it faster? Switch to the Turbo version—just change the model name and number of steps:
pipe = ErnieImagePipeline.from_pretrained(
"baidu/ERNIE-Image-Turbo",
torch_dtype=torch.bfloat16,
).to("cuda")
image = pipe(
prompt="A black-and-white Chinese rural dog running across the grass",
height=1024,
width=1024,
num_inference_steps=8, # Only 8 steps needed
guidance_scale=1.0,
use_pe=True
).images[0]
Option 2: SGLang (best for service deployment)
git clone https://github.com/sgl-project/sglang.git
sglang serve --model-path baidu/ERNIE-Image-Turbo
curl -X POST http://localhost:30000/generate \
-H "Content-Type: application/json" \
-d '{
"prompt": "A black-and-white Chinese rural dog running across the grass",
"height": 1024,
"width": 1024,
"num_inference_steps": 8,
"guidance_scale": 1.0,
"use_pe": true
}' \
--output output.png
Option 3: ComfyUI
The latest version of ComfyUI already supports ERNIE-Image, and workflow templates can be found on the official GitHub repository.
Option 4: Try it online
Don’t want to deploy it locally? You can try it online directly:
- HuggingFace Demo:https://huggingface.co/spaces/baidu/ERNIE-Image-Turbo
- Baidu AI Studio:https://aistudio.baidu.com/ernieimage
7. Compared with Other Models, Which One Should You Choose?
Here’s my view. If you’re making a model selection decision, this quick guide may help:
| Your Need | Recommendation |
|---|---|
| Chinese posters, design layout | ERNIE-Image (best text rendering) |
| Maximum image quality and style diversity | FLUX.2 or Midjourney |
| Extensive LoRA ecosystem and plugins | Stable Diffusion 3.5 |
| Real-time generation, speed first | ERNIE-Image-Turbo (8-step generation) |
| Custom modification and fine-tuning | ERNIE-Image (supports Unsloth GGUF + AI-Toolkit fine-tuning) |
💡 My take: ERNIE-Image isn’t trying to crush every competitor on every dimension. What it has done is make a major leap in practical usability for Chinese-language scenarios. For Chinese users working in e-commerce design, social media operations, and content creation, “getting the text right inside the image” matters far more than “making the image look just a little better.” From that perspective, the significance of ERNIE-Image being open-sourced is on par with the breakthroughs domestic large models have made in language.
8. A Technical Trend Worth Watching
ERNIE-Image’s release actually reflects an important trend: text-to-image models are shifting from “making images look good” to “making images controllable.”
In 2024, everyone was still competing on realism. In 2025, the competition shifted toward faster inference and fewer parameters. By 2026, controllability has become the new battleground.
A few important signals worth watching:
- Text rendering is becoming a standard capability: ERNIE-Image, Z-Image, and LongCat-Image are all focusing heavily on this direction, showing that the industry now recognizes it as a key bottleneck for real-world adoption of AI image generation.
- Distillation-based acceleration is becoming standard: ERNIE-Image-Turbo (DMD+RL) and Z-Image-Turbo (Decoupled-DMD) show that distilling large models into smaller, faster ones is already the mainstream route.
- Single-stream architectures are becoming the trend: The field is moving from dual-stream designs (where text and image are processed separately) toward single-stream DiT, and ERNIE-Image, Z-Image, and Qwen-Image have all chosen this path.
- Prompt Enhancers are becoming standard components: Automatically expanding user prompts has worked very well in ERNIE-Image and may well be adopted by more models going forward.
Summary
- ERNIE-Image is Baidu’s open-source 8B text-to-image model, released under Apache-2.0 and friendly to commercial use.
- Among open-source models, it has the strongest text rendering capability, leading in mixed Chinese-English text and layout accuracy.
- It can be deployed with just 24GB of VRAM, making it friendly to consumer GPUs.
- The Turbo version needs only 8 inference steps, giving it a speed advantage.
- It is well suited for Chinese-language scenarios such as poster design, comic creation, and social media assets.