A 3B Model Rewrites Your Prompt: The Complete Guide to ERNIE-Image PE Enhancer

May 1, 2026

A 3B Model Rewrites Your Prompt: The Complete Guide to ERNIE-Image PE Enhancer

ERNIE-Image's Prompt Enhancer is not a simple keyword completion tool — it's a complete 3-billion-parameter language model. Based on Mistral AI's Ministral-3B-Instruct, it is fine-tuned for one task: rewriting your prompt into a version the diffusion model can understand better.

This article fully dissects PE from a technical perspective: working principles, model architecture, parameter configuration, platform integration, and underlying behavior mechanisms.


1. Model Architecture: What is Ministral-3B?

Base Model

Property Value
Model Name Ministral-3B-Instruct-2512
Developer Mistral AI
Parameters ~3B (3 billion)
Architecture Transformer Decoder + MoE (Mixture of Experts)
Context Window 32K tokens
Release December 2025
License Apache 2.0 / Open-source friendly

Core Advantages of MoE Architecture

The Ministral 3 series uses a Mixture of Experts (MoE) architecture. Unlike traditional Dense Transformers that activate all parameters per inference, MoE structures activate only a subset of "expert" modules each time.

This means:

  • Nominal parameters: 3B
  • Active parameters: Significantly lower than 3B
  • Inference speed: Much faster than same-parameter Dense models
  • Resource consumption: Lower VRAM usage

For a "fast-response" preprocessing module like PE, MoE architecture is a natural advantage.

What Did Baidu Do?

Baidu performed instruction fine-tuning on top of Ministral-3B-Instruct, using high-quality paired data of "short prompt → structured long prompt".

The fine-tuned model is saved as ernie-image-prompt-enhancer.safetensors, approximately the compressed size of the base model (depending on quantization format).


2. PE's Inference Flow: Every Step from Input to Output

Step 1: Prompt Reception and Preprocessing

The user inputs an original prompt, PE receives and tokenizes it:

Raw input: "a ceramic coffee mug"
Tokenize: [a] [ceramic] [coffee] [mug]
Token count: ~4 tokens

Step 2: System Instruction Injection

PE injects a system instruction (like a Chat model's system prompt) before the user prompt, instructing the model to perform prompt expansion. Approximate system instruction:

You are a prompt enhancer for image generation. Expand the user's brief description into a detailed, structured prompt that includes subject details, environment, lighting, composition, and style. Maintain the user's original intent. Output the enhanced prompt.

Step 3: MoE Inference

Ministral-3B's MoE routing mechanism selects which expert modules to activate. For prompt expansion tasks, experts related to "text generation" and "descriptive language" are primarily activated.

Inference parameters:

  • max_length: 1536–2048 (output token cap)
  • temperature: 0.6 (balance creativity and stability)
  • top_p: 0.8 (nucleus sampling)
  • thinking mode: Disabled (skip chain of thought, output directly)

Step 4: Output Truncation and Post-processing

PE generates the expanded prompt, truncates to the specified length, then passes it to the Text Encoder (Ministral-3B Text Encoder — note this is a separate model instance, another Ministral-3B used for text encoding).


3. ComfyUI Integration Details

File Structure

ComfyUI/models/text_encoders/
├── ministral-3-3b.safetensors          ← Text Encoder (text encoding)
└── ernie-image-prompt-enhancer.safetensors  ← PE module (prompt expansion)

Note: These two models are both based on Ministral-3B but serve different purposes:

  • Text Encoder: Encodes prompts into text feature vectors for the diffusion model
  • PE module: Rewrites prompt text, output is still text

Node Configuration

PE is loaded via CLIPLoader node in ComfyUI, even though it's an LLM, not a CLIP model. This is ComfyUI's interface reuse design.

Key PE settings (via subgraph):

{
  "model": "ernie-image-prompt-enhancer.safetensors",
  "max_length": 2048,
  "temperature": 0.6,
  "top_p": 0.8,
  "thinking_enabled": false,
  "enabled": true
}

Position in Workflow

In ComfyUI's standard ERNIE-Image workflow, the PE node sits before the Text Encoder:

[CustomText (user input)] → [PE Node] → [Text Encoder (Ministral-3B)] → [DiT Diffusion Model]

If PE is disabled (enabled: false), the workflow becomes:

[CustomText (user input)] → [Text Encoder (Ministral-3B)] → [DiT Diffusion Model]

4. Diffusers API Integration

Python Code Example

from diffusers import ErnieImagePipeline
import torch

pipe = ErnieImagePipeline.from_pretrained(
"Baidu/ERNIE-Image-Turbo",
torch_dtype=torch.bfloat16
).to("cuda")

PE on (default)

image_on = pipe(
prompt="a ceramic coffee mug",
use_pe=True, # PE enabled
height=1264,
width=848,
num_inference_steps=8,
guidance_scale=1.0
).images[0]

PE off

image_off = pipe(
prompt="a ceramic coffee mug",
use_pe=False, # PE disabled
height=1264,
width=848,
num_inference_steps=8,
guidance_scale=1.0
).images[0]

PE's Inference Overhead

Enabling PE adds one LLM inference overhead:

Configuration Extra Time Extra VRAM
PE on + 8B DiT ~2-5 sec (PE inference) ~6-8 GB (PE model)
PE off + 8B DiT 0 0

In VRAM-constrained environments (12-16GB GPUs), turning off PE frees ~6-8GB for the diffusion model.


5. PE's Language Behavior Analysis

Roots of Chinese Translation Tendency

PE's tendency to translate prompts to Chinese comes from training data distribution:

  1. Baidu's training data: Chinese prompts dominate Baidu's prompt expansion training data
  2. Ministral base: Ministral-3B-Instruct's multilingual ability is balanced, but fine-tuning shifts toward the dominant training language
  3. ERNIE-Image positioning: ERNIE-Image's primary target market is China, so Chinese-first is a reasonable training strategy

Impact on Multilingual Users

User Language PE Behavior Effect
English input May translate to Chinese or mixed English text rendering may become Chinese
Chinese input Stays Chinese or mixed Minimal impact
Other languages May translate to Chinese May alter original semantics

Solutions

  • Precise text rendering: Turn off PE
  • English users: Turn off PE or use external LLM
  • Chinese users: PE-friendly, but still turn off for text rendering

6. GGUF Quantization: PE's Blind Spot

Unsloth provides GGUF quantized versions of ERNIE-Image DiT and Ministral-3B Text Encoder, but no GGUF version exists for PE.

Likely reasons:

  1. PE is already a 3B model — quantization benefit is limited
  2. PE loads via CLIPLoader interface — GGUF requires a specialized Loader
  3. Lower priority — core diffusion model and Text Encoder quantization save more VRAM

Result: GGUF workflows cannot use PE. Low-VRAM users must write detailed prompts themselves.


7. PE Parameter Tuning Guide

temperature Effects

temperature Effect Use Case
0.3 Conservative, stable but templated Consistency needed in batch generation
0.6 Balanced (default) General purpose
0.8+ Creative but may drift from original intent Exploratory creation

max_length Effects

max_length Effect Use Case
512 Brief expansion Quick preview
1536 Standard expansion (recommended) General purpose
2048 Detailed expansion Complex scenarios
4096+ Too long, wasteful Not recommended

top_p Effects

top_p Effect
0.5 Very conservative, only most likely words
0.8 Balanced (recommended)
0.95 More open, allows long-tail vocabulary

8. External LLM Alternative Implementation

Approach 1: Manual Replacement

1. Expand prompt with GPT-4/Claude
2. Turn off PE in ERNIE-Image
3. Paste expanded prompt

Approach 2: Automated Pipeline

# Pseudocode
def generate_with_custom_pe(user_prompt, llm_client):
    enhanced = llm_client.chat(
        system="Expand this image generation prompt into detailed description...",
        user=user_prompt
    )
    image = ernie_image_pipe(
        prompt=enhanced,
        use_pe=False
    )
    return image

Approach 3: Custom PE Fine-tuning

If you have extensive domain-specific prompt data, you can fine-tune PE based on Ministral-3B:

1. Download ministral/Ministral-3b-instruct
2. Prepare [short_prompt, long_prompt] paired data
3. Fine-tune with LoRA/QLoRA
4. Export as safetensors
5. Replace PE module in ComfyUI

9. Summary

PE is an instruction-fine-tuned language model based on Ministral-3B MoE architecture, whose core task is expanding short prompts into structured long prompts.

  • Architecture: MoE → fast inference, low VRAM
  • Behavior: Rewrites rather than enhances, leans toward Chinese translation
  • Parameters: temperature 0.6, max_length 2048, top_p 0.8 (recommended)
  • Limitations: Not available in GGUF workflows
  • Alternatives: External LLM or custom fine-tuning

Understanding PE's technical underpinnings helps you make more accurate toggle decisions and parameter optimizations.

ERNIE-Image Team