A 3B Model Rewrites Your Prompt: The Complete Guide to ERNIE-Image PE Enhancer
ERNIE-Image's Prompt Enhancer is not a simple keyword completion tool — it's a complete 3-billion-parameter language model. Based on Mistral AI's Ministral-3B-Instruct, it is fine-tuned for one task: rewriting your prompt into a version the diffusion model can understand better.
This article fully dissects PE from a technical perspective: working principles, model architecture, parameter configuration, platform integration, and underlying behavior mechanisms.
1. Model Architecture: What is Ministral-3B?
Base Model
| Property | Value |
|---|---|
| Model Name | Ministral-3B-Instruct-2512 |
| Developer | Mistral AI |
| Parameters | ~3B (3 billion) |
| Architecture | Transformer Decoder + MoE (Mixture of Experts) |
| Context Window | 32K tokens |
| Release | December 2025 |
| License | Apache 2.0 / Open-source friendly |
Core Advantages of MoE Architecture
The Ministral 3 series uses a Mixture of Experts (MoE) architecture. Unlike traditional Dense Transformers that activate all parameters per inference, MoE structures activate only a subset of "expert" modules each time.
This means:
- Nominal parameters: 3B
- Active parameters: Significantly lower than 3B
- Inference speed: Much faster than same-parameter Dense models
- Resource consumption: Lower VRAM usage
For a "fast-response" preprocessing module like PE, MoE architecture is a natural advantage.
What Did Baidu Do?
Baidu performed instruction fine-tuning on top of Ministral-3B-Instruct, using high-quality paired data of "short prompt → structured long prompt".
The fine-tuned model is saved as ernie-image-prompt-enhancer.safetensors, approximately the compressed size of the base model (depending on quantization format).
2. PE's Inference Flow: Every Step from Input to Output
Step 1: Prompt Reception and Preprocessing
The user inputs an original prompt, PE receives and tokenizes it:
Raw input: "a ceramic coffee mug"
Tokenize: [a] [ceramic] [coffee] [mug]
Token count: ~4 tokens
Step 2: System Instruction Injection
PE injects a system instruction (like a Chat model's system prompt) before the user prompt, instructing the model to perform prompt expansion. Approximate system instruction:
You are a prompt enhancer for image generation. Expand the user's brief description into a detailed, structured prompt that includes subject details, environment, lighting, composition, and style. Maintain the user's original intent. Output the enhanced prompt.
Step 3: MoE Inference
Ministral-3B's MoE routing mechanism selects which expert modules to activate. For prompt expansion tasks, experts related to "text generation" and "descriptive language" are primarily activated.
Inference parameters:
- max_length: 1536–2048 (output token cap)
- temperature: 0.6 (balance creativity and stability)
- top_p: 0.8 (nucleus sampling)
- thinking mode: Disabled (skip chain of thought, output directly)
Step 4: Output Truncation and Post-processing
PE generates the expanded prompt, truncates to the specified length, then passes it to the Text Encoder (Ministral-3B Text Encoder — note this is a separate model instance, another Ministral-3B used for text encoding).
3. ComfyUI Integration Details
File Structure
ComfyUI/models/text_encoders/
├── ministral-3-3b.safetensors ← Text Encoder (text encoding)
└── ernie-image-prompt-enhancer.safetensors ← PE module (prompt expansion)
Note: These two models are both based on Ministral-3B but serve different purposes:
- Text Encoder: Encodes prompts into text feature vectors for the diffusion model
- PE module: Rewrites prompt text, output is still text
Node Configuration
PE is loaded via CLIPLoader node in ComfyUI, even though it's an LLM, not a CLIP model. This is ComfyUI's interface reuse design.
Key PE settings (via subgraph):
{
"model": "ernie-image-prompt-enhancer.safetensors",
"max_length": 2048,
"temperature": 0.6,
"top_p": 0.8,
"thinking_enabled": false,
"enabled": true
}
Position in Workflow
In ComfyUI's standard ERNIE-Image workflow, the PE node sits before the Text Encoder:
[CustomText (user input)] → [PE Node] → [Text Encoder (Ministral-3B)] → [DiT Diffusion Model]
If PE is disabled (enabled: false), the workflow becomes:
[CustomText (user input)] → [Text Encoder (Ministral-3B)] → [DiT Diffusion Model]
4. Diffusers API Integration
Python Code Example
from diffusers import ErnieImagePipeline
import torch
pipe = ErnieImagePipeline.from_pretrained(
"Baidu/ERNIE-Image-Turbo",
torch_dtype=torch.bfloat16
).to("cuda")
PE on (default)
image_on = pipe(
prompt="a ceramic coffee mug",
use_pe=True, # PE enabled
height=1264,
width=848,
num_inference_steps=8,
guidance_scale=1.0
).images[0]
PE off
image_off = pipe(
prompt="a ceramic coffee mug",
use_pe=False, # PE disabled
height=1264,
width=848,
num_inference_steps=8,
guidance_scale=1.0
).images[0]
PE's Inference Overhead
Enabling PE adds one LLM inference overhead:
| Configuration | Extra Time | Extra VRAM |
|---|---|---|
| PE on + 8B DiT | ~2-5 sec (PE inference) | ~6-8 GB (PE model) |
| PE off + 8B DiT | 0 | 0 |
In VRAM-constrained environments (12-16GB GPUs), turning off PE frees ~6-8GB for the diffusion model.
5. PE's Language Behavior Analysis
Roots of Chinese Translation Tendency
PE's tendency to translate prompts to Chinese comes from training data distribution:
- Baidu's training data: Chinese prompts dominate Baidu's prompt expansion training data
- Ministral base: Ministral-3B-Instruct's multilingual ability is balanced, but fine-tuning shifts toward the dominant training language
- ERNIE-Image positioning: ERNIE-Image's primary target market is China, so Chinese-first is a reasonable training strategy
Impact on Multilingual Users
| User Language | PE Behavior | Effect |
|---|---|---|
| English input | May translate to Chinese or mixed | English text rendering may become Chinese |
| Chinese input | Stays Chinese or mixed | Minimal impact |
| Other languages | May translate to Chinese | May alter original semantics |
Solutions
- Precise text rendering: Turn off PE
- English users: Turn off PE or use external LLM
- Chinese users: PE-friendly, but still turn off for text rendering
6. GGUF Quantization: PE's Blind Spot
Unsloth provides GGUF quantized versions of ERNIE-Image DiT and Ministral-3B Text Encoder, but no GGUF version exists for PE.
Likely reasons:
- PE is already a 3B model — quantization benefit is limited
- PE loads via CLIPLoader interface — GGUF requires a specialized Loader
- Lower priority — core diffusion model and Text Encoder quantization save more VRAM
Result: GGUF workflows cannot use PE. Low-VRAM users must write detailed prompts themselves.
7. PE Parameter Tuning Guide
temperature Effects
| temperature | Effect | Use Case |
|---|---|---|
| 0.3 | Conservative, stable but templated | Consistency needed in batch generation |
| 0.6 | Balanced (default) | General purpose |
| 0.8+ | Creative but may drift from original intent | Exploratory creation |
max_length Effects
| max_length | Effect | Use Case |
|---|---|---|
| 512 | Brief expansion | Quick preview |
| 1536 | Standard expansion (recommended) | General purpose |
| 2048 | Detailed expansion | Complex scenarios |
| 4096+ | Too long, wasteful | Not recommended |
top_p Effects
| top_p | Effect |
|---|---|
| 0.5 | Very conservative, only most likely words |
| 0.8 | Balanced (recommended) |
| 0.95 | More open, allows long-tail vocabulary |
8. External LLM Alternative Implementation
Approach 1: Manual Replacement
1. Expand prompt with GPT-4/Claude
2. Turn off PE in ERNIE-Image
3. Paste expanded prompt
Approach 2: Automated Pipeline
# Pseudocode
def generate_with_custom_pe(user_prompt, llm_client):
enhanced = llm_client.chat(
system="Expand this image generation prompt into detailed description...",
user=user_prompt
)
image = ernie_image_pipe(
prompt=enhanced,
use_pe=False
)
return image
Approach 3: Custom PE Fine-tuning
If you have extensive domain-specific prompt data, you can fine-tune PE based on Ministral-3B:
1. Download ministral/Ministral-3b-instruct
2. Prepare [short_prompt, long_prompt] paired data
3. Fine-tune with LoRA/QLoRA
4. Export as safetensors
5. Replace PE module in ComfyUI
9. Summary
PE is an instruction-fine-tuned language model based on Ministral-3B MoE architecture, whose core task is expanding short prompts into structured long prompts.
- Architecture: MoE → fast inference, low VRAM
- Behavior: Rewrites rather than enhances, leans toward Chinese translation
- Parameters: temperature 0.6, max_length 2048, top_p 0.8 (recommended)
- Limitations: Not available in GGUF workflows
- Alternatives: External LLM or custom fine-tuning
Understanding PE's technical underpinnings helps you make more accurate toggle decisions and parameter optimizations.