ERNIE-Image Prompt Enhancer Toggle Strategy and Best Practices: When to Turn It On, When to Turn It Off

Jun 15, 2026

ERNIE-Image Prompt Enhancer Toggle Strategy and Best Practices: When to Turn It On, When to Turn It Off

Published: 2026-06-15
Author: ERNIE-Image Blog
Keywords: ernie-image prompt enhancer, ernie-image PE toggle, ernie-image use_pe, ernie-image PE best practices


Introduction

ERNIE-Image's Prompt Enhancer (PE) is a 3B-parameter model fine-tuned on Ministral 3B that automatically expands your brief prompt input into rich, structured descriptions. Sounds great — but community feedback has been mixed.

On Reddit, users complain: "i did some test previously with the enhancer and it makes the result less coherent (i got consistently missing people in a long prompt)." Others say: "PE dramatically improves results from simple prompts."

The issue isn't whether PE is good or bad — it's when to turn it on, and when to turn it off. Based on official benchmark data and community testing, this article provides clear PE toggle strategies.

How PE Works

PE is essentially a prompt expander. When you input "a cat sitting on a sofa," PE expands it to something like:

"A photorealistic image of a fluffy orange tabby cat sitting comfortably on a modern grey fabric sofa. Natural window light illuminates the scene from the left, casting soft shadows on the wooden floor. The cat's eyes are focused on something off-frame, creating a sense of curiosity. Shot on 50mm lens, f/2.8, warm color temperature."

This process is handled by Ministral 3B — a lightweight but powerful language model specifically trained for image prompt expansion.

Key Data: How PE Affects Performance

Official benchmark data from ERNIE-Image's HuggingFace card reveals the double-edged nature of PE:

Instruction Following (GENEval)

Configuration Overall Score
ERNIE-Image (w/o PE) 0.8856
ERNIE-Image (w/ PE) 0.8728

Conclusion: PE slightly reduces instruction following scores (-1.5%). For scenarios requiring precise instruction execution (multi-object positioning, complex relationship descriptions), disabling PE is more reliable.

English Image Generation (OneIG-EN)

Capability w/o PE w/ PE Change
Overall Score — 0.5750 —
Reasoning — 0.3566 Improvement
Alignment — 0.8678 Improvement
Text Rendering — 0.9788 Significant improvement

Conclusion: PE significantly improves text rendering (0.9788) and alignment. For text generation scenarios (posters, infographics, comics), PE must be enabled.

Long Prompt Handling (LongTextBench)

Model Score
Seedream 4.5 0.9882
ERNIE-Image (w/ PE) 0.9733
FLUX.2-klein-9B 0.5413

Conclusion: ERNIE-Image with PE ranks near the top for long prompt handling. For prompts requiring complex scene descriptions, PE is a key advantage.

PE Toggle Decision Tree

Based on the data and community experience above, here are clear PE toggle strategies:

Enable PE ✅

  1. Short prompts (< 30 words): PE dramatically expands information, improving generation quality
  2. Text rendering needed (posters, infographics, comic bubbles): PE pushes text rendering to 0.9788
  3. Creative expansion needed (simple descriptions with unclear style): PE adds lighting, composition, color details
  4. Long, complex scene prompts: LongTextBench score of 0.9733 shows PE excels at complex descriptions

Disable PE ❌

  1. Already detailed prompts (> 100 words): PE may add redundant or conflicting descriptions
  2. Multi-object precise positioning: GENEval shows PE slightly reduces instruction following
  3. Precise control needed (specific composition, specific color schemes): PE's "creative expansion" may deviate from intent
  4. Missing people in long prompts: Known issue reported on Reddit community

Practical Code Examples

Using PE with Diffusers

import torch
from diffusers import ErnieImagePipeline

pipe = ErnieImagePipeline.from_pretrained(
"Baidu/ERNIE-Image",
torch_dtype=torch.bfloat16,
).to("cuda")

Scenario 1: Short prompt → Enable PE

image = pipe(
prompt="a cat on a sofa",
height=1264,
width=848,
num_inference_steps=50,
guidance_scale=4.0,
use_pe=True # ✅ Short prompt, enable PE for quality boost
).images[0]

Scenario 2: Detailed prompt → Disable PE

image = pipe(
prompt="A photorealistic image of a fluffy orange tabby cat sitting on a modern grey fabric sofa. Natural window light from the left, soft shadows on wooden floor. Shot on 50mm lens, f/2.8.",
height=1264,
width=848,
num_inference_steps=50,
guidance_scale=4.0,
use_pe=False # ❌ Already detailed, disable PE to avoid conflicts
).images[0]

Scenario 3: Text rendering → Enable PE

image = pipe(
prompt="A poster that says 'COFFEE' with beans and steam",
height=1264,
width=848,
num_inference_steps=50,
guidance_scale=4.0,
use_pe=True # ✅ Text rendering, must enable PE
).images[0]

Controlling PE in ComfyUI

In ComfyUI, PE control is managed through text encoder loading:

  • PE On: Load both ministral-3-3b.safetensors AND ernie-image-prompt-enhancer.safetensors
  • PE Off: Load only ministral-3-3b.safetensors

Controlling PE via SGLang API

# Enable PE
curl -X POST http://localhost:30000/v1/images/generations \
  -H "Content-Type: application/json" \
  -d '{
    "prompt": "a cat on a sofa",
    "use_pe": true,
    "height": 1264,
    "width": 848
  }'

Disable PE

curl -X POST http://localhost:30000/v1/images/generations
-H "Content-Type: application/json"
-d '{
"prompt": "a detailed description...",
"use_pe": false,
"height": 1264,
"width": 848
}'

Community Test Cases

Case 1: PE On vs Off — Short Prompt

Prompt: "a cyberpunk city street"

  • PE On: Generates a complete cyberpunk scene with neon lights, rain, holographic billboards, steam pipes
  • PE Off: Generates a basic city street lacking atmospheric detail

Conclusion: For short prompts, PE On > Off.

Case 2: PE On vs Off — Multi-Object Scene

Prompt: "a woman in red coat standing next to a man in blue suit, with a dog between them, in a park"

  • PE On: Reddit users report "consistently missing people" — only 2 of 3 characters appear
  • PE Off: More precise instruction following, all three characters appear

Conclusion: For multi-object precise positioning, PE Off > On.

Case 3: PE On vs Off — Text Rendering

Prompt: "A movie poster with the text 'THE MATRIX' in green digital rain style"

  • PE On: Clear text rendering, "THE MATRIX" correctly spelled, green digital rain effect
  • PE Off: Text may be misspelled or blurry

Conclusion: For text rendering, PE On >> Off.

PE Behavior in Turbo Mode

ERNIE-Image-Turbo also supports PE, but with slightly different behavior:

Benchmark Turbo w/o PE Turbo w/ PE
GENEval — 0.8510
OneIG-EN — 0.8375 (text: 0.8351)

Turbo mode PE text rendering score (0.8351) is lower than SFT version (0.9788), because Turbo prioritizes speed and aesthetics. For maximum text rendering quality, use SFT version + PE.

PE Strategy in Batch Generation

In batch generation scenarios (see EI-093), PE strategy should be grouped by prompt type:

# Group strategy
short_prompts = ["a cat", "city street", "mountain landscape"]  # → use_pe=True
detailed_prompts = ["A photorealistic..."]  # → use_pe=False
text_prompts = ["poster with text 'SALE'"]  # → use_pe=True

Batch processing

for group, use_pe in [(short_prompts, True), (detailed_prompts, False), (text_prompts, True)]:
for prompt in group:
image = pipe(prompt=prompt, use_pe=use_pe, ...)

Summary: PE Toggle Quick Reference

Scenario PE Reason
Short prompt (< 30 words) ✅ On Expand information
Detailed prompt (> 100 words) ❌ Off Avoid redundancy/conflicts
Text rendering (posters/infographics) ✅ On 0.9788 rendering score
Multi-object precise positioning ❌ Off Maintain instruction following
Precise composition control ❌ Off Avoid PE creative deviation
Creative exploration / flexible style ✅ On PE adds details
Turbo mode fast generation ⚠️ Depends Turbo PE text score lower
Long prompt complex scenes ✅ On LongTextBench advantage

Core principle: PE is not "the more the better" — it's "use as needed." Enable PE for short prompts and text rendering; disable it for precise control and multi-object scenes.

ERNIE-Image Team