ERNIE-Image 8B Turbo: A Deep Dive from Benchmarks to LoRA Fine-Tuning

Apr 27, 2026

ERNIE-Image 8B Turbo: A Deep Dive from Benchmarks to LoRA Fine-Tuning

00 Introduction

Baidu's ERNIE team has open-sourced ERNIE-Image, an 8B-parameter text-to-image model built on a single-stream DiT architecture. It runs on consumer-grade GPUs with just 24GB of VRAM, leading open-source models across instruction following, text rendering, and other mainstream benchmarks. It excels at structured scenes such as posters, comic panels, and multi-panel layouts. The team also released ERNIE-Image Turbo, which generates high-fidelity images in just 8 inference steps.

Try it: ModelScope Studio
Open-source: ERNIE-Image / ERNIE-Image-Turbo


01 Model Overview

ERNIE-Image is built on a DiT architecture with 8 billion parameters, requiring only 24GB of VRAM to generate complex images that rival top-tier commercial models. It leads open-source models across GenEval, OneIG, and LongTextBench benchmarks, with overall performance approaching state-of-the-art models like NanoBanana and Seedream 4.5.

Key Features

Feature Description
Small Model, Strong Performance 8B parameter model ranks #1 among open-source models across mainstream benchmarks, approaching top closed-source models like NanoBanana 2.0 and Seedream 4.5
Precise Text Rendering Stable performance on high-density text, long text, and layout-sensitive tasks; ideal for posters, infographics, and UI-like images
Complex Instruction Following Strong comprehension and precise execution on multi-subject relationships, detail constraints, and knowledge-intensive prompts
Structured Generation Maintains layout logic and visual organization in posters, comics, storyboards, and panel-based tasks
Multi-Style Coverage Supports realistic photography, anime/2D, film, surrealism, silhouettes, vintage photos, and cinematic soft-light styles
Consumer Hardware Friendly Deployable on 24GB VRAM, significantly lowering the barrier for research and production environments

Prompt Enhancer

ERNIE-Image performs best with detailed, structured long prompts, but users often provide short, casual descriptions. The team built a lightweight 3B-parameter Prompt Enhancer that automatically expands brief inputs into richer, more structured prompts without changing the original intent. The effect is especially noticeable in structured visual tasks like posters, anime, web layouts, and game screenshots.


02 Benchmark Results

ERNIE-Image was evaluated across four mainstream text-to-image benchmarks: GenEval (compositional generation), OneIG-EN / OneIG-ZH (English/Chinese open-domain image generation), and LongTextBench (long text rendering fidelity).

Benchmark Rank Score
GenEval (Compositional Generation) #1 0.8856
LongTextBench (Long Text Rendering) #2 0.9733
OneIG-ZH (Chinese Open-Domain) #2 0.5543
OneIG-EN (English Open-Domain) #3 0.5750

Key Insight: Ranked #2 on LongTextBench with excellent performance in both Chinese and English long text rendering; highly competitive on the Text dimension of OneIG, demonstrating strong multilingual text generation. These results come from just an 8B-parameter DiT architecture—one of the most parameter-efficient models at this performance level.


03 Model Positioning and Comparison

In the open-source text-to-image landscape, ERNIE-Image-Turbo 8B sits in the mid-range for parameter count, but its controllability, instruction following, and text rendering make it highly competitive.

Model Parameters Architecture / Key Traits Use Cases & Hardware Requirements
ERNIE-Image-Turbo 8B 8B Single-stream DiT + LDM + built-in 3B Prompt Enhancer Complex instruction tracking, precise text rendering, structured image generation
HunyuanImage-3.0 (Tencent) 80B (MoE, ~13B active) Native multimodal autoregressive Complex prompts / knowledge reasoning / CN-EN rendering, requires datacenter-grade hardware
FLUX.2 [dev] (Black Forest) 32B Rectified Flow Transformer Extremely strong prompt following / detail / coherence; quantized version runs on consumer GPUs
FLUX.1 [dev/schnell] ~12B Classic DiT / Flow Matching Top-tier text rendering, richest community ecosystem (ComfyUI, etc.)
SD 3.5 Large (Stability) 8.1B (MMDiT) Latest SD flagship, supports 1MP+ Significant improvements in prompt following / layout; most mature LoRA / fine-tuning ecosystem
Qwen-Image 2.0 (Alibaba) ~7B Lightweight & efficient, unified generation + editing Strong Chinese / multilingual rendering, native 2K resolution, great for infographics
Z-Image-Turbo ~6B Efficient real-time / edge deployment Low-resource environments, fast speed

Core Insight: Parameter count ≠ absolute capability. ERNIE-Image achieves SOTA-level controllability with just 8B parameters, significantly outperforming most open-source models in complex instruction tracking, precise text rendering, and structured generation.


04 Inference and Deployment Guide

1. Diffusers Inference

pip install git+https://github.com/huggingface/diffusers
import torch
from diffusers import ErnieImagePipeline

pipe = ErnieImagePipeline.from_pretrained(
"Baidu/ERNIE-Image-Turbo",
torch_dtype=torch.bfloat16
).to("cuda")

image = pipe(
prompt="A photography shot of an urban street scene",
height=1264,
width=848,
num_inference_steps=8,
guidance_scale=1.0,
use_pe=True
).images[0]

image.save("output.png")

2. SGLang Inference (Server Deployment)

git clone https://github.com/sgl-project/sglang.git
sglang serve --model-path baidu/ERNIE-Image-Turbo

3. Diffsynth Inference (Low VRAM Optimization)

pip install -U diffsynth==2.0.8
from diffsynth.pipelines.ernie_image import ErnieImagePipeline, ModelConfig
import torch

vram_config = {
"offload_dtype": torch.bfloat16, "offload_device": "cpu",
"onload_dtype": torch.bfloat16, "onload_device": "cpu",
"preparing_dtype": torch.bfloat16, "preparing_device": "cuda",
"computation_dtype": torch.bfloat16, "computation_device": "cuda"
}

pipe = ErnieImagePipeline.from_pretrained(
torch_dtype=torch.bfloat16, device="cuda",
model_configs=[
ModelConfig(model_id="PaddlePaddle/ERNIE-Image", origin_file_pattern="transformer/diffusion_pytorch_model*.safetensors", **vram_config),
ModelConfig(model_id="PaddlePaddle/ERNIE-Image", origin_file_pattern="text_encoder/model.safetensors", **vram_config),
ModelConfig(model_id="PaddlePaddle/ERNIE-Image", origin_file_pattern="vae/diffusion_pytorch_model.safetensors", **vram_config)
],
tokenizer_config=ModelConfig(model_id="PaddlePaddle/ERNIE-Image", origin_file_pattern="tokenizer/"),
vram_limit=torch.cuda.mem_get_info("cuda")[1] / (1024 ** 3) - 0.5
)

image = pipe(
prompt="A black-and-white Chinese rural dog",
height=1024, width=1024, seed=42,
num_inference_steps=50, cfg_scale=4.0
)
image.save("output.jpg")


05 Prompt Writing Rules and Best Practices

ERNIE-Image heavily relies on long, detailed, structured prompts. The model does not "hallucinate" or fill in gaps—it requires explicit description of everything.

✅ Core Rules

  • Long prompts far outperform short ones: Short prompts → literal interpretation, messy layout, text errors, lack of narrative
  • Explicit description required: Text content, layout position, logical relationships, visual hierarchy, narrative flow
  • Leverage the enhancer: Brief idea → 3B PE auto-expansion → manual refinement
  • Language choice: Chinese preferred (strongest semantic control), Chinese-English mix or pure English also works

🏗️ Structured Writing Framework

  1. Subject (Subject): Objects, scenes, characters, actions
  2. Details & Relations (Details & Relations): Position, interaction, lighting, material, quantity, perspective
  3. Composition (Composition): Panels, poster layout, multi-panel, text position/font
  4. Style (Style): Art style, era, mood, lighting
  5. Quality Boosters (Quality Boosters): High detail, sharp, cinematic lighting, no watermark, commercial quality
  6. Negative Prompt (Negative Prompt): e.g., --no blurry, low resolution, distorted

📥 Complex Scene Prompt Example

A studio macro photography shot showcasing a handmade polymer clay miniature diorama. At the center is a tiny Oreo-themed shop, presented in a vertical composition. The shop's roof is constructed from several giant Oreo sandwich cookies interlocked together—the cookies are deep black with a thick white cream layer in between, featuring classic embossed textures and clear "OREO" lettering.

06 Model LoRA Training

DiffSynth-Studio supports image-to-image LoRA training for ERNIE-Image with automatic VRAM management.

Training Command

accelerate launch examples/ernie_image/model_training/train.py \
  --dataset_base_path data/diffsynth_example_dataset/ernie_image/ \
  --dataset_metadata_path data/diffsynth_example_dataset/ernie_image/metadata.csv \
  --max_pixels 1048576 \
  --dataset_repeat 50 \
  --learning_rate 1e-4 \
  --num_epochs 5 \
  --output_path "./models/train/Ernie-Image-T2I_lora" \
  --lora_rank 32 \
  --use_gradient_checkpointing

07 Core Insights and Recommendations

  • Rule-driven model: Not following its prompt rules will drastically reduce quality
  • King of structured scenes: Posters, comic panels, UI mockups, game screenshots, and infographics are its absolute strength
  • Text rendering advantage: Ideal for generating images with extensive Chinese/English text
  • Recommended workflow: Brief concept → 3B Prompt Enhancer → Manual layout/text review → Generate
  • Parameter count ≠ absolute capability: 8B parameters achieve SOTA-level controllability
  • Consumer hardware friendly: 24GB VRAM for deployment; Diffsynth framework supports as low as 3GB VRAM

Source: ModelScope Community — ERNIE-Image Team Open-Source Release

ERNIE-Image Team

ERNIE-Image 8B Turbo: A Deep Dive from Benchmarks to LoRA Fine-Tuning | Blog