ERNIE-Image Prompt Enhancer 开关策略与最佳实践:什么时候开?什么时候关?

Jun 15, 2026

ERNIE-Image Prompt Enhancer 开关策略与最佳实践:什么时候开?什么时候关?

发布日期:2026-06-15
作者:ERNIE-Image 博客
关键词:ernie-image prompt enhancer、ernie-image PE 开关、ernie-image use_pe、ernie-image PE 最佳实践


引言

ERNIE-Image 的 Prompt Enhancer(PE)是一个基于 Ministral 3B 微调的 3B 参数模型,它可以将你输入的简短 prompt 自动扩展为丰富、结构化的详细描述。听起来很棒——但社区用户的实际反馈却两极分化。

Reddit 上有人抱怨:"i did some test previously with the enhancer and it makes the result less coherent (i got consistently missing people in a long prompt)." 另一些人则说:"PE 让简单 prompt 的效果提升巨大。"

问题不在于 PE 好不好,而在于什么时候该开,什么时候该关。本文将基于官方基准数据和社区实测经验,给出明确的 PE 开关策略。

PE 工作原理

PE 本质上是一个 prompt 扩展器。当你输入"a cat sitting on a sofa",PE 会将其扩展为类似这样的详细描述:

"A photorealistic image of a fluffy orange tabby cat sitting comfortably on a modern grey fabric sofa. Natural window light illuminates the scene from the left, casting soft shadows on the wooden floor. The cat's eyes are focused on something off-frame, creating a sense of curiosity. Shot on 50mm lens, f/2.8, warm color temperature."

这个过程由 Ministral 3B 完成——一个轻量级但强大的语言模型,专门训练用于图像 prompt 扩展。

关键数据:PE 对性能的影响

来自 ERNIE-Image 官方 HuggingFace 卡片的基准数据揭示了 PE 的双面效应:

指令遵循(GENEval)

配置 综合得分
ERNIE-Image (w/o PE) 0.8856
ERNIE-Image (w/ PE) 0.8728

结论:PE 略微降低指令遵循得分(-1.5%)。对于需要精确指令执行的场景(多对象定位、复杂关系描述),关闭 PE 更可靠。

英文图像生成(OneIG-EN)

能力 w/o PE w/ PE 变化
综合得分 — 0.5750 —
Reasoning — 0.3566 提升
Alignment — 0.8678 提升
Text Rendering — 0.9788 大幅提升

结论:PE 显著提升文字渲染能力(0.9788)和对齐能力。对于需要文字生成的场景(海报、信息图、漫画),必须开启 PE。

长 Prompt 处理(LongTextBench)

模型 得分
Seedream 4.5 0.9882
ERNIE-Image (w/ PE) 0.9733
FLUX.2-klein-9B 0.5413

结论:PE 加持下的 ERNIE-Image 在长 prompt 处理上排名前列。对于需要复杂场景描述的 prompt,PE 是关键优势。

PE 开关决策树

基于以上数据和社区经验,以下是明确的 PE 开关策略:

开启 PE 的场景 ✅

  1. 简短 prompt(< 30 词):PE 能大幅扩展信息量,提升生成质量
  2. 需要文字渲染(海报、信息图、漫画气泡):PE 将 text rendering 提升到 0.9788
  3. 需要创意扩展(风格不明确的简单描述):PE 补充光照、构图、色彩等细节
  4. 长 prompt 复杂场景:LongTextBench 得分 0.9733,PE 擅长处理复杂描述

关闭 PE 的场景 ❌

  1. 已经非常详细的 prompt(> 100 词):PE 可能添加冗余或冲突描述
  2. 多对象精确定位:GENEval 得分显示 PE 略微降低指令遵循
  3. 需要精确控制(特定构图、特定颜色方案):PE 的"创意扩展"可能偏离预期
  4. 长 prompt 中出现人物缺失:Reddit 社区反馈的已知问题

实战代码示例

Diffusers 中使用 PE

import torch
from diffusers import ErnieImagePipeline

pipe = ErnieImagePipeline.from_pretrained(
"Baidu/ERNIE-Image",
torch_dtype=torch.bfloat16,
).to("cuda")

场景 1:简短 prompt → 开 PE

image = pipe(
prompt="a cat on a sofa",
height=1264,
width=848,
num_inference_steps=50,
guidance_scale=4.0,
use_pe=True # ✅ 简短 prompt,开 PE 提升质量
).images[0]

场景 2:详细 prompt → 关 PE

image = pipe(
prompt="A photorealistic image of a fluffy orange tabby cat sitting on a modern grey fabric sofa. Natural window light from the left, soft shadows on wooden floor. Shot on 50mm lens, f/2.8.",
height=1264,
width=848,
num_inference_steps=50,
guidance_scale=4.0,
use_pe=False # ❌ 已经详细,关 PE 避免冲突
).images[0]

场景 3:需要文字渲染 → 开 PE

image = pipe(
prompt="A poster that says 'COFFEE' with beans and steam",
height=1264,
width=848,
num_inference_steps=50,
guidance_scale=4.0,
use_pe=True # ✅ 文字渲染,必须开 PE
).images[0]

ComfyUI 中控制 PE

在 ComfyUI 中,PE 控制通过加载不同的文本编码器实现:

  • 开启 PE:同时加载 ministral-3-3b.safetensors 和 ernie-image-prompt-enhancer.safetensors
  • 关闭 PE:只加载 ministral-3-3b.safetensors

SGLang API 中控制 PE

# 开启 PE
curl -X POST http://localhost:30000/v1/images/generations \
  -H "Content-Type: application/json" \
  -d '{
    "prompt": "a cat on a sofa",
    "use_pe": true,
    "height": 1264,
    "width": 848
  }'

关闭 PE

curl -X POST http://localhost:30000/v1/images/generations
-H "Content-Type: application/json"
-d '{
"prompt": "a detailed description...",
"use_pe": false,
"height": 1264,
"width": 848
}'

社区实测案例

案例 1:PE 开 vs 关 — 简短 prompt

Prompt:"a cyberpunk city street"

  • PE 开:生成包含霓虹灯、雨夜、全息广告牌、蒸汽管道的完整赛博朋克场景
  • PE 关:生成基本的城市街道,缺乏氛围细节

结论:简短 prompt 下,PE 开 > 关。

案例 2:PE 开 vs 关 — 多对象场景

Prompt:"a woman in red coat standing next to a man in blue suit, with a dog between them, in a park"

  • PE 开:Reddit 用户报告"consistently missing people"——三人中可能只出现 2 人
  • PE 关:指令遵循更精确,三人都出现在画面中

结论:多对象精确定位场景,PE 关 > 开。

案例 3:PE 开 vs 关 — 文字渲染

Prompt:"A movie poster with the text 'THE MATRIX' in green digital rain style"

  • PE 开:文字渲染清晰,"THE MATRIX" 正确拼写,绿色数字雨效果
  • PE 关:文字可能拼写错误或模糊

结论:文字渲染场景,PE 开 >> 关。

Turbo 模式下的 PE 行为

ERNIE-Image-Turbo 也支持 PE,但行为略有不同:

基准 Turbo w/o PE Turbo w/ PE
GENEval — 0.8510
OneIG-EN — 0.8375 (text: 0.8351)

Turbo 模式下的 PE 文字渲染得分(0.8351)低于 SFT 版本(0.9788),这是因为 Turbo 优先优化速度和美学。对于极致文字渲染需求,建议使用 SFT 版本 + PE。

批量生成中的 PE 策略

在批量生成场景(参见 EI-093),PE 策略应按 prompt 类型分组:

# 分组策略
short_prompts = ["a cat", "city street", "mountain landscape"]  # → use_pe=True
detailed_prompts = ["A photorealistic..."]  # → use_pe=False
text_prompts = ["poster with text 'SALE'"]  # → use_pe=True

批量处理

for group, use_pe in [(short_prompts, True), (detailed_prompts, False), (text_prompts, True)]:
for prompt in group:
image = pipe(prompt=prompt, use_pe=use_pe, ...)

总结:PE 开关决策速查表

场景 PE 原因
简短 prompt(< 30 词) ✅ 开 扩展信息量
详细 prompt(> 100 词) ❌ 关 避免冗余冲突
文字渲染(海报/信息图) ✅ 开 0.9788 渲染得分
多对象精确定位 ❌ 关 保持指令遵循
需要精确构图控制 ❌ 关 避免 PE 创意偏移
创意探索 / 风格不固定 ✅ 开 PE 补充细节
Turbo 模式快速出图 ⚠️ 看情况 Turbo PE 文字得分较低
长 prompt 复杂场景 ✅ 开 LongTextBench 优势

核心原则:PE 不是"越好越开",而是"按需开关"。简短 prompt 和文字渲染场景开 PE,精确控制和多对象场景关 PE。

ERNIE-Image Team