ERNIE-Image 8B Turbo: A Deep Dive from Benchmarks to LoRA Fine-Tuning
00 Introduction
Baidu's ERNIE team has open-sourced ERNIE-Image, an 8B-parameter text-to-image model built on a single-stream DiT architecture. It runs on consumer-grade GPUs with just 24GB of VRAM, leading open-source models across instruction following, text rendering, and other mainstream benchmarks. It excels at structured scenes such as posters, comic panels, and multi-panel layouts. The team also released ERNIE-Image Turbo, which generates high-fidelity images in just 8 inference steps.
Try it: ModelScope Studio
Open-source: ERNIE-Image / ERNIE-Image-Turbo
01 Model Overview
ERNIE-Image is built on a DiT architecture with 8 billion parameters, requiring only 24GB of VRAM to generate complex images that rival top-tier commercial models. It leads open-source models across GenEval, OneIG, and LongTextBench benchmarks, with overall performance approaching state-of-the-art models like NanoBanana and Seedream 4.5.
Key Features
| Feature | Description |
|---|---|
| Small Model, Strong Performance | 8B parameter model ranks #1 among open-source models across mainstream benchmarks, approaching top closed-source models like NanoBanana 2.0 and Seedream 4.5 |
| Precise Text Rendering | Stable performance on high-density text, long text, and layout-sensitive tasks; ideal for posters, infographics, and UI-like images |
| Complex Instruction Following | Strong comprehension and precise execution on multi-subject relationships, detail constraints, and knowledge-intensive prompts |
| Structured Generation | Maintains layout logic and visual organization in posters, comics, storyboards, and panel-based tasks |
| Multi-Style Coverage | Supports realistic photography, anime/2D, film, surrealism, silhouettes, vintage photos, and cinematic soft-light styles |
| Consumer Hardware Friendly | Deployable on 24GB VRAM, significantly lowering the barrier for research and production environments |
Prompt Enhancer
ERNIE-Image performs best with detailed, structured long prompts, but users often provide short, casual descriptions. The team built a lightweight 3B-parameter Prompt Enhancer that automatically expands brief inputs into richer, more structured prompts without changing the original intent. The effect is especially noticeable in structured visual tasks like posters, anime, web layouts, and game screenshots.
02 Benchmark Results
ERNIE-Image was evaluated across four mainstream text-to-image benchmarks: GenEval (compositional generation), OneIG-EN / OneIG-ZH (English/Chinese open-domain image generation), and LongTextBench (long text rendering fidelity).
| Benchmark | Rank | Score |
|---|---|---|
| GenEval (Compositional Generation) | #1 | 0.8856 |
| LongTextBench (Long Text Rendering) | #2 | 0.9733 |
| OneIG-ZH (Chinese Open-Domain) | #2 | 0.5543 |
| OneIG-EN (English Open-Domain) | #3 | 0.5750 |
Key Insight: Ranked #2 on LongTextBench with excellent performance in both Chinese and English long text rendering; highly competitive on the Text dimension of OneIG, demonstrating strong multilingual text generation. These results come from just an 8B-parameter DiT architecture—one of the most parameter-efficient models at this performance level.
03 Model Positioning and Comparison
In the open-source text-to-image landscape, ERNIE-Image-Turbo 8B sits in the mid-range for parameter count, but its controllability, instruction following, and text rendering make it highly competitive.
| Model | Parameters | Architecture / Key Traits | Use Cases & Hardware Requirements |
|---|---|---|---|
| ERNIE-Image-Turbo 8B | 8B | Single-stream DiT + LDM + built-in 3B Prompt Enhancer | Complex instruction tracking, precise text rendering, structured image generation |
| HunyuanImage-3.0 (Tencent) | 80B (MoE, ~13B active) | Native multimodal autoregressive | Complex prompts / knowledge reasoning / CN-EN rendering, requires datacenter-grade hardware |
| FLUX.2 [dev] (Black Forest) | 32B | Rectified Flow Transformer | Extremely strong prompt following / detail / coherence; quantized version runs on consumer GPUs |
| FLUX.1 [dev/schnell] | ~12B | Classic DiT / Flow Matching | Top-tier text rendering, richest community ecosystem (ComfyUI, etc.) |
| SD 3.5 Large (Stability) | 8.1B (MMDiT) | Latest SD flagship, supports 1MP+ | Significant improvements in prompt following / layout; most mature LoRA / fine-tuning ecosystem |
| Qwen-Image 2.0 (Alibaba) | ~7B | Lightweight & efficient, unified generation + editing | Strong Chinese / multilingual rendering, native 2K resolution, great for infographics |
| Z-Image-Turbo | ~6B | Efficient real-time / edge deployment | Low-resource environments, fast speed |
Core Insight: Parameter count ≠ absolute capability. ERNIE-Image achieves SOTA-level controllability with just 8B parameters, significantly outperforming most open-source models in complex instruction tracking, precise text rendering, and structured generation.
04 Inference and Deployment Guide
1. Diffusers Inference
pip install git+https://github.com/huggingface/diffusers
import torch
from diffusers import ErnieImagePipeline
pipe = ErnieImagePipeline.from_pretrained(
"Baidu/ERNIE-Image-Turbo",
torch_dtype=torch.bfloat16
).to("cuda")
image = pipe(
prompt="A photography shot of an urban street scene",
height=1264,
width=848,
num_inference_steps=8,
guidance_scale=1.0,
use_pe=True
).images[0]
image.save("output.png")
2. SGLang Inference (Server Deployment)
git clone https://github.com/sgl-project/sglang.git
sglang serve --model-path baidu/ERNIE-Image-Turbo
3. Diffsynth Inference (Low VRAM Optimization)
pip install -U diffsynth==2.0.8
from diffsynth.pipelines.ernie_image import ErnieImagePipeline, ModelConfig
import torch
vram_config = {
"offload_dtype": torch.bfloat16, "offload_device": "cpu",
"onload_dtype": torch.bfloat16, "onload_device": "cpu",
"preparing_dtype": torch.bfloat16, "preparing_device": "cuda",
"computation_dtype": torch.bfloat16, "computation_device": "cuda"
}
pipe = ErnieImagePipeline.from_pretrained(
torch_dtype=torch.bfloat16, device="cuda",
model_configs=[
ModelConfig(model_id="PaddlePaddle/ERNIE-Image", origin_file_pattern="transformer/diffusion_pytorch_model*.safetensors", **vram_config),
ModelConfig(model_id="PaddlePaddle/ERNIE-Image", origin_file_pattern="text_encoder/model.safetensors", **vram_config),
ModelConfig(model_id="PaddlePaddle/ERNIE-Image", origin_file_pattern="vae/diffusion_pytorch_model.safetensors", **vram_config)
],
tokenizer_config=ModelConfig(model_id="PaddlePaddle/ERNIE-Image", origin_file_pattern="tokenizer/"),
vram_limit=torch.cuda.mem_get_info("cuda")[1] / (1024 ** 3) - 0.5
)
image = pipe(
prompt="A black-and-white Chinese rural dog",
height=1024, width=1024, seed=42,
num_inference_steps=50, cfg_scale=4.0
)
image.save("output.jpg")
05 Prompt Writing Rules and Best Practices
ERNIE-Image heavily relies on long, detailed, structured prompts. The model does not "hallucinate" or fill in gaps—it requires explicit description of everything.
✅ Core Rules
- Long prompts far outperform short ones: Short prompts → literal interpretation, messy layout, text errors, lack of narrative
- Explicit description required: Text content, layout position, logical relationships, visual hierarchy, narrative flow
- Leverage the enhancer: Brief idea → 3B PE auto-expansion → manual refinement
- Language choice: Chinese preferred (strongest semantic control), Chinese-English mix or pure English also works
🏗️ Structured Writing Framework
- Subject (Subject): Objects, scenes, characters, actions
- Details & Relations (Details & Relations): Position, interaction, lighting, material, quantity, perspective
- Composition (Composition): Panels, poster layout, multi-panel, text position/font
- Style (Style): Art style, era, mood, lighting
- Quality Boosters (Quality Boosters): High detail, sharp, cinematic lighting, no watermark, commercial quality
- Negative Prompt (Negative Prompt): e.g., --no blurry, low resolution, distorted
📥 Complex Scene Prompt Example
A studio macro photography shot showcasing a handmade polymer clay miniature diorama. At the center is a tiny Oreo-themed shop, presented in a vertical composition. The shop's roof is constructed from several giant Oreo sandwich cookies interlocked together—the cookies are deep black with a thick white cream layer in between, featuring classic embossed textures and clear "OREO" lettering.
06 Model LoRA Training
DiffSynth-Studio supports image-to-image LoRA training for ERNIE-Image with automatic VRAM management.
Training Command
accelerate launch examples/ernie_image/model_training/train.py \
--dataset_base_path data/diffsynth_example_dataset/ernie_image/ \
--dataset_metadata_path data/diffsynth_example_dataset/ernie_image/metadata.csv \
--max_pixels 1048576 \
--dataset_repeat 50 \
--learning_rate 1e-4 \
--num_epochs 5 \
--output_path "./models/train/Ernie-Image-T2I_lora" \
--lora_rank 32 \
--use_gradient_checkpointing
07 Core Insights and Recommendations
- Rule-driven model: Not following its prompt rules will drastically reduce quality
- King of structured scenes: Posters, comic panels, UI mockups, game screenshots, and infographics are its absolute strength
- Text rendering advantage: Ideal for generating images with extensive Chinese/English text
- Recommended workflow: Brief concept → 3B Prompt Enhancer → Manual layout/text review → Generate
- Parameter count ≠ absolute capability: 8B parameters achieve SOTA-level controllability
- Consumer hardware friendly: 24GB VRAM for deployment; Diffsynth framework supports as low as 3GB VRAM
Source: ModelScope Community — ERNIE-Image Team Open-Source Release