ERNIE-Image Prompt Enhancer Toggle Strategy and Best Practices: When to Turn It On, When to Turn It Off
Published: 2026-06-15
Author: ERNIE-Image Blog
Keywords: ernie-image prompt enhancer, ernie-image PE toggle, ernie-image use_pe, ernie-image PE best practices
Introduction
ERNIE-Image's Prompt Enhancer (PE) is a 3B-parameter model fine-tuned on Ministral 3B that automatically expands your brief prompt input into rich, structured descriptions. Sounds great — but community feedback has been mixed.
On Reddit, users complain: "i did some test previously with the enhancer and it makes the result less coherent (i got consistently missing people in a long prompt)." Others say: "PE dramatically improves results from simple prompts."
The issue isn't whether PE is good or bad — it's when to turn it on, and when to turn it off. Based on official benchmark data and community testing, this article provides clear PE toggle strategies.
How PE Works
PE is essentially a prompt expander. When you input "a cat sitting on a sofa," PE expands it to something like:
"A photorealistic image of a fluffy orange tabby cat sitting comfortably on a modern grey fabric sofa. Natural window light illuminates the scene from the left, casting soft shadows on the wooden floor. The cat's eyes are focused on something off-frame, creating a sense of curiosity. Shot on 50mm lens, f/2.8, warm color temperature."
This process is handled by Ministral 3B — a lightweight but powerful language model specifically trained for image prompt expansion.
Key Data: How PE Affects Performance
Official benchmark data from ERNIE-Image's HuggingFace card reveals the double-edged nature of PE:
Instruction Following (GENEval)
| Configuration | Overall Score |
|---|---|
| ERNIE-Image (w/o PE) | 0.8856 |
| ERNIE-Image (w/ PE) | 0.8728 |
Conclusion: PE slightly reduces instruction following scores (-1.5%). For scenarios requiring precise instruction execution (multi-object positioning, complex relationship descriptions), disabling PE is more reliable.
English Image Generation (OneIG-EN)
| Capability | w/o PE | w/ PE | Change |
|---|---|---|---|
| Overall Score | — | 0.5750 | — |
| Reasoning | — | 0.3566 | Improvement |
| Alignment | — | 0.8678 | Improvement |
| Text Rendering | — | 0.9788 | Significant improvement |
Conclusion: PE significantly improves text rendering (0.9788) and alignment. For text generation scenarios (posters, infographics, comics), PE must be enabled.
Long Prompt Handling (LongTextBench)
| Model | Score |
|---|---|
| Seedream 4.5 | 0.9882 |
| ERNIE-Image (w/ PE) | 0.9733 |
| FLUX.2-klein-9B | 0.5413 |
Conclusion: ERNIE-Image with PE ranks near the top for long prompt handling. For prompts requiring complex scene descriptions, PE is a key advantage.
PE Toggle Decision Tree
Based on the data and community experience above, here are clear PE toggle strategies:
Enable PE ✅
- Short prompts (< 30 words): PE dramatically expands information, improving generation quality
- Text rendering needed (posters, infographics, comic bubbles): PE pushes text rendering to 0.9788
- Creative expansion needed (simple descriptions with unclear style): PE adds lighting, composition, color details
- Long, complex scene prompts: LongTextBench score of 0.9733 shows PE excels at complex descriptions
Disable PE ❌
- Already detailed prompts (> 100 words): PE may add redundant or conflicting descriptions
- Multi-object precise positioning: GENEval shows PE slightly reduces instruction following
- Precise control needed (specific composition, specific color schemes): PE's "creative expansion" may deviate from intent
- Missing people in long prompts: Known issue reported on Reddit community
Practical Code Examples
Using PE with Diffusers
import torch
from diffusers import ErnieImagePipeline
pipe = ErnieImagePipeline.from_pretrained(
"Baidu/ERNIE-Image",
torch_dtype=torch.bfloat16,
).to("cuda")
Scenario 1: Short prompt → Enable PE
image = pipe(
prompt="a cat on a sofa",
height=1264,
width=848,
num_inference_steps=50,
guidance_scale=4.0,
use_pe=True # ✅ Short prompt, enable PE for quality boost
).images[0]
Scenario 2: Detailed prompt → Disable PE
image = pipe(
prompt="A photorealistic image of a fluffy orange tabby cat sitting on a modern grey fabric sofa. Natural window light from the left, soft shadows on wooden floor. Shot on 50mm lens, f/2.8.",
height=1264,
width=848,
num_inference_steps=50,
guidance_scale=4.0,
use_pe=False # ❌ Already detailed, disable PE to avoid conflicts
).images[0]
Scenario 3: Text rendering → Enable PE
image = pipe(
prompt="A poster that says 'COFFEE' with beans and steam",
height=1264,
width=848,
num_inference_steps=50,
guidance_scale=4.0,
use_pe=True # ✅ Text rendering, must enable PE
).images[0]
Controlling PE in ComfyUI
In ComfyUI, PE control is managed through text encoder loading:
- PE On: Load both
ministral-3-3b.safetensorsANDernie-image-prompt-enhancer.safetensors - PE Off: Load only
ministral-3-3b.safetensors
Controlling PE via SGLang API
# Enable PE
curl -X POST http://localhost:30000/v1/images/generations \
-H "Content-Type: application/json" \
-d '{
"prompt": "a cat on a sofa",
"use_pe": true,
"height": 1264,
"width": 848
}'
Disable PE
curl -X POST http://localhost:30000/v1/images/generations
-H "Content-Type: application/json"
-d '{
"prompt": "a detailed description...",
"use_pe": false,
"height": 1264,
"width": 848
}'
Community Test Cases
Case 1: PE On vs Off — Short Prompt
Prompt: "a cyberpunk city street"
- PE On: Generates a complete cyberpunk scene with neon lights, rain, holographic billboards, steam pipes
- PE Off: Generates a basic city street lacking atmospheric detail
Conclusion: For short prompts, PE On > Off.
Case 2: PE On vs Off — Multi-Object Scene
Prompt: "a woman in red coat standing next to a man in blue suit, with a dog between them, in a park"
- PE On: Reddit users report "consistently missing people" — only 2 of 3 characters appear
- PE Off: More precise instruction following, all three characters appear
Conclusion: For multi-object precise positioning, PE Off > On.
Case 3: PE On vs Off — Text Rendering
Prompt: "A movie poster with the text 'THE MATRIX' in green digital rain style"
- PE On: Clear text rendering, "THE MATRIX" correctly spelled, green digital rain effect
- PE Off: Text may be misspelled or blurry
Conclusion: For text rendering, PE On >> Off.
PE Behavior in Turbo Mode
ERNIE-Image-Turbo also supports PE, but with slightly different behavior:
| Benchmark | Turbo w/o PE | Turbo w/ PE |
|---|---|---|
| GENEval | — | 0.8510 |
| OneIG-EN | — | 0.8375 (text: 0.8351) |
Turbo mode PE text rendering score (0.8351) is lower than SFT version (0.9788), because Turbo prioritizes speed and aesthetics. For maximum text rendering quality, use SFT version + PE.
PE Strategy in Batch Generation
In batch generation scenarios (see EI-093), PE strategy should be grouped by prompt type:
# Group strategy
short_prompts = ["a cat", "city street", "mountain landscape"] # → use_pe=True
detailed_prompts = ["A photorealistic..."] # → use_pe=False
text_prompts = ["poster with text 'SALE'"] # → use_pe=True
Batch processing
for group, use_pe in [(short_prompts, True), (detailed_prompts, False), (text_prompts, True)]:
for prompt in group:
image = pipe(prompt=prompt, use_pe=use_pe, ...)
Summary: PE Toggle Quick Reference
| Scenario | PE | Reason |
|---|---|---|
| Short prompt (< 30 words) | ✅ On | Expand information |
| Detailed prompt (> 100 words) | ❌ Off | Avoid redundancy/conflicts |
| Text rendering (posters/infographics) | ✅ On | 0.9788 rendering score |
| Multi-object precise positioning | ❌ Off | Maintain instruction following |
| Precise composition control | ❌ Off | Avoid PE creative deviation |
| Creative exploration / flexible style | ✅ On | PE adds details |
| Turbo mode fast generation | ⚠️ Depends | Turbo PE text score lower |
| Long prompt complex scenes | ✅ On | LongTextBench advantage |
Core principle: PE is not "the more the better" — it's "use as needed." Enable PE for short prompts and text rendering; disable it for precise control and multi-object scenes.