What Exactly is ERNIE-Image's PE? Uncovering the 3B Prompt Enhancer's Mechanics and Toggle Tips

Apr 30, 2026

What Exactly is ERNIE-Image's PE? Uncovering the 3B Prompt Enhancer's Mechanics and Toggle Tips

If you've used ERNIE-Image, you've probably noticed a module called "Prompt Enhancer" (PE) in the workflow.

Enabled by default in ComfyUI, enabled by default in the Diffusers API. You type a simple description, it automatically expands it into a professional prompt with lighting, composition, and style — and then generates a noticeably better image.

But at the same time, the community has plenty of complaints:

"I clearly wrote 'Hello World', but the image shows Chinese text."
"PE completely ruined my carefully crafted prompt."
"When should I turn it on, and when should I turn it off?"

What exactly is PE? Why does it improve image quality? In what cases does it actually hurt you? This article covers it all.


1. Where PE Sits in the ERNIE-Image Inference Pipeline

The complete ERNIE-Image inference pipeline consists of four core components:

User prompt input
    ↓
[PE Module]  3B language model, auto-expands prompt (optional)
    ↓
[Text Encoder]  Ministral-3B, encodes prompt into text features
    ↓
[8B DiT Diffusion Model]  Generates image from text features
    ↓
[VAE Decode]  Outputs final image

PE is the first preprocessing module. It runs before image generation, rewrites the user's original prompt into a more detailed structured description, then passes it to the Text Encoder.

PE is not part of the diffusion model — it's an independent language model. Think of it as ERNIE-Image's "pre-prompt translator" — converting casual user expressions into professional descriptions that the diffusion model understands better.

ComfyUI file structure:

ComfyUI/models/
├── diffusion_models/ernie-image.safetensors      # 8B DiT main model
├── text_encoders/
│   ├── ministral-3-3b.safetensors                # Text Encoder
│   └── ernie-image-prompt-enhancer.safetensors   # PE Module
└── vae/flux2-vae.safetensors                     # VAE Decoder

The PE file ernie-image-prompt-enhancer.safetensors is a complete 3B-parameter language model.


2. PE's Base Model — Why Ministral-3B?

PE's base model is Ministral-3B-Instruct (by Mistral AI). Baidu fine-tuned it with instructions specifically for text-to-image prompt expansion tasks.

Ministral-3B is the smallest model in the Mistral 3 series, released at the end of 2025, using a MoE (Mixture of Experts) architecture.

Four reasons for choosing it as the PE base:

1. Inference Speed is a Hard Requirement

PE is a preprocessing step — users don't want to wait. The 3B model runs extremely fast on consumer GPUs — a single prompt expansion typically completes in seconds. Using a 7B or larger model would make PE the bottleneck itself, negating the speed advantage of ERNIE-Image Turbo's 8-step inference.

2. MoE Architecture is Naturally Efficient

The Ministral 3 series uses a Mixture of Experts structure, activating only a subset of parameters per inference. The actual compute is far less than the nominal 3B. For a "plug-and-play" preprocessing module like PE, efficiency determines user experience.

3. Strong Instruction-Tuning Capability

Ministral-3B-Instruct is already an instruction-tuned model, excelling at "receive text, rewrite as instructed." This is exactly PE's core task.

4. Open-Source License Compatibility

Ministral uses an open-source-friendly license, consistent with ERNIE-Image's overall Apache 2.0 open-source strategy.


3. What Does PE Actually Do?

PE's core capability is expanding brief descriptions into structured prompts covering 5 key elements:

Element What PE Adds Example
Subject Details Material, color, form "matte ceramic texture", "weathered oak surface"
Environment/Scene Background, atmosphere "morning kitchen setting", "soft diffused daylight"
Lighting Light source direction, quality "soft window light from the left", "volumetric fog"
Composition Lens, depth of field, perspective "shallow depth of field", "centered composition"
Style Photography type, artistic style "commercial product photography", "cinematic lighting"

Real Comparison

Original Input (3 words):

a ceramic coffee mug

After PE Expansion (similar output):

Close-up product photograph of a matte ceramic coffee mug on a weathered oak table in a morning kitchen setting, soft window light from the left casting warm tones, shallow depth of field, commercial product photography style, 8K detail, centered composition

The original prompt had only 3 words. PE expands it to include product photography, lighting direction, depth of field, style, and more. The diffusion model receives a professional-level prompt instead of a keyword — the quality naturally improves.

PE Parameter Settings in ComfyUI

In ComfyUI workflows, the PE node is loaded via CLIPLoader (though it's actually an LLM — ComfyUI reuses the CLIPLoader interface). Key parameters:

Parameter Recommended Description
max_length 1536–2048 Maximum output prompt length
temperature 0.6 Sampling temperature, controls randomness
top_p 0.8 Nucleus sampling threshold
thinking mode Disabled Turn off reasoning mode

4. PE's Real Contribution to Benchmarks

The official HuggingFace model card provides GenEval benchmark data with/without PE:

Configuration GenEval Score
ERNIE-Image (w/ PE) 0.8625
ERNIE-Image (w/o PE) ~0.73
ERNIE-Image-Turbo (w/ PE) 0.8510

PE contributes roughly 13% score improvement on the GenEval benchmark. The gain is especially significant for short-prompt scenarios — the shorter the input, the larger the benefit from PE.

But here's a key contradiction: for text rendering scenarios, PE should actually be turned off.

LongTextBench shows ERNIE-Image (w/ PE) scoring 0.9788 — looks perfect. But in practice, if your prompt contains text that needs precise rendering, PE may rewrite that text content, causing the image to display text you didn't want.

This is why ERNIE-Image is strong at text rendering (high LongTextBench score), but PE can actually undermine text precision — PE rewrites the prompt text, and the diffusion model faithfully executes the rewritten text.


5. PE's Biggest Pitfall: It Translates Your Prompt to Chinese

This is a key behavior discovered by Reddit community users through testing (r/StableDiffusion, 2026-04):

"The default workflow from the Comfy templates has a 'Prompt Enhancer' section that among other things, translates your prompt to Chinese."

During expansion, PE tends to convert prompts to Chinese for processing. This means:

  • English prompts may be translated to Chinese, so the diffusion model actually receives Chinese instructions
  • If the prompt contains precise English text to render (e.g., a title "SUMMER BEATS 2026"), PE may rewrite it to Chinese, and Chinese characters appear on the image

Scenario.com's official documentation also explicitly recommends: for Chinese text rendering scenarios, turn off Prompt Enhancer.


6. Quick Toggle Reference

Scenario PE Recommendation Reason
One-line simple description ✅ On PE auto-fills lighting, composition, style details
Don't know prompt engineering ✅ On Built-in prompt assistant, lowers the barrier
Turbo mode fast generation ✅ On PE + 8 steps = fastest workflow
Need precise English text ❌ Off PE may translate to Chinese
Need precise Chinese text ❌ Off PE may rewrite text content
Already have detailed structured prompt ❌ Off Avoid conflicts and redundant expansion
Fixed seed iteration ❌ Off Isolate variables, judge the effect of changes
GGUF quantized workflow ❌ N/A Unsloth doesn't provide GGUF version of PE

7. PE Alternatives — Stronger LLMs

Official recommendation: using a stronger LLM (such as GPT-4 or Claude) as a Prompt Enhancer can further improve results.

The built-in PE's advantages are free, offline, and fast. Its disadvantage is limited understanding depth from a 3B model — expansion results tend to be templated. If you pursue maximum quality, adjust your workflow:

  1. Use GPT-4/Claude to expand short prompts into professional-level detailed descriptions
  2. Turn off PE in ERNIE-Image
  3. Paste the expanded prompt and generate

This completely bypasses PE's rewriting behavior, so the diffusion model executes the carefully crafted instructions from you (or a stronger LLM).


8. Summary

PE is one of the most misunderstood modules in ERNIE-Image. At its core, it's a 3B language model whose task is to expand user briefs into structured prompts.

  • Short prompts → PE fills in details, quality improves noticeably
  • Precise text → PE rewrites text content, must be turned off
  • Already detailed prompts → PE is overkill, recommend off
  • Pursuing the best quality → Replace built-in PE with GPT-4/Claude

Understanding PE's mechanics and boundaries matters more than blindly toggling it on or off.

ERNIE-Image Team