ERNIE-Image Prompt Enhancer 开关策略与最佳实践:什么时候开?什么时候关?
发布日期:2026-06-15
作者:ERNIE-Image 博客
关键词:ernie-image prompt enhancer、ernie-image PE 开关、ernie-image use_pe、ernie-image PE 最佳实践
引言
ERNIE-Image 的 Prompt Enhancer(PE)是一个基于 Ministral 3B 微调的 3B 参数模型,它可以将你输入的简短 prompt 自动扩展为丰富、结构化的详细描述。听起来很棒——但社区用户的实际反馈却两极分化。
Reddit 上有人抱怨:"i did some test previously with the enhancer and it makes the result less coherent (i got consistently missing people in a long prompt)." 另一些人则说:"PE 让简单 prompt 的效果提升巨大。"
问题不在于 PE 好不好,而在于什么时候该开,什么时候该关。本文将基于官方基准数据和社区实测经验,给出明确的 PE 开关策略。
PE 工作原理
PE 本质上是一个 prompt 扩展器。当你输入"a cat sitting on a sofa",PE 会将其扩展为类似这样的详细描述:
"A photorealistic image of a fluffy orange tabby cat sitting comfortably on a modern grey fabric sofa. Natural window light illuminates the scene from the left, casting soft shadows on the wooden floor. The cat's eyes are focused on something off-frame, creating a sense of curiosity. Shot on 50mm lens, f/2.8, warm color temperature."
这个过程由 Ministral 3B 完成——一个轻量级但强大的语言模型,专门训练用于图像 prompt 扩展。
关键数据:PE 对性能的影响
来自 ERNIE-Image 官方 HuggingFace 卡片的基准数据揭示了 PE 的双面效应:
指令遵循(GENEval)
| 配置 | 综合得分 |
|---|---|
| ERNIE-Image (w/o PE) | 0.8856 |
| ERNIE-Image (w/ PE) | 0.8728 |
结论:PE 略微降低指令遵循得分(-1.5%)。对于需要精确指令执行的场景(多对象定位、复杂关系描述),关闭 PE 更可靠。
英文图像生成(OneIG-EN)
| 能力 | w/o PE | w/ PE | 变化 |
|---|---|---|---|
| 综合得分 | — | 0.5750 | — |
| Reasoning | — | 0.3566 | 提升 |
| Alignment | — | 0.8678 | 提升 |
| Text Rendering | — | 0.9788 | 大幅提升 |
结论:PE 显著提升文字渲染能力(0.9788)和对齐能力。对于需要文字生成的场景(海报、信息图、漫画),必须开启 PE。
长 Prompt 处理(LongTextBench)
| 模型 | 得分 |
|---|---|
| Seedream 4.5 | 0.9882 |
| ERNIE-Image (w/ PE) | 0.9733 |
| FLUX.2-klein-9B | 0.5413 |
结论:PE 加持下的 ERNIE-Image 在长 prompt 处理上排名前列。对于需要复杂场景描述的 prompt,PE 是关键优势。
PE 开关决策树
基于以上数据和社区经验,以下是明确的 PE 开关策略:
开启 PE 的场景 ✅
- 简短 prompt(< 30 词):PE 能大幅扩展信息量,提升生成质量
- 需要文字渲染(海报、信息图、漫画气泡):PE 将 text rendering 提升到 0.9788
- 需要创意扩展(风格不明确的简单描述):PE 补充光照、构图、色彩等细节
- 长 prompt 复杂场景:LongTextBench 得分 0.9733,PE 擅长处理复杂描述
关闭 PE 的场景 ❌
- 已经非常详细的 prompt(> 100 词):PE 可能添加冗余或冲突描述
- 多对象精确定位:GENEval 得分显示 PE 略微降低指令遵循
- 需要精确控制(特定构图、特定颜色方案):PE 的"创意扩展"可能偏离预期
- 长 prompt 中出现人物缺失:Reddit 社区反馈的已知问题
实战代码示例
Diffusers 中使用 PE
import torch
from diffusers import ErnieImagePipeline
pipe = ErnieImagePipeline.from_pretrained(
"Baidu/ERNIE-Image",
torch_dtype=torch.bfloat16,
).to("cuda")
场景 1:简短 prompt → 开 PE
image = pipe(
prompt="a cat on a sofa",
height=1264,
width=848,
num_inference_steps=50,
guidance_scale=4.0,
use_pe=True # ✅ 简短 prompt,开 PE 提升质量
).images[0]
场景 2:详细 prompt → 关 PE
image = pipe(
prompt="A photorealistic image of a fluffy orange tabby cat sitting on a modern grey fabric sofa. Natural window light from the left, soft shadows on wooden floor. Shot on 50mm lens, f/2.8.",
height=1264,
width=848,
num_inference_steps=50,
guidance_scale=4.0,
use_pe=False # ❌ 已经详细,关 PE 避免冲突
).images[0]
场景 3:需要文字渲染 → 开 PE
image = pipe(
prompt="A poster that says 'COFFEE' with beans and steam",
height=1264,
width=848,
num_inference_steps=50,
guidance_scale=4.0,
use_pe=True # ✅ 文字渲染,必须开 PE
).images[0]
ComfyUI 中控制 PE
在 ComfyUI 中,PE 控制通过加载不同的文本编码器实现:
- 开启 PE:同时加载
ministral-3-3b.safetensors和ernie-image-prompt-enhancer.safetensors - 关闭 PE:只加载
ministral-3-3b.safetensors
SGLang API 中控制 PE
# 开启 PE
curl -X POST http://localhost:30000/v1/images/generations \
-H "Content-Type: application/json" \
-d '{
"prompt": "a cat on a sofa",
"use_pe": true,
"height": 1264,
"width": 848
}'
关闭 PE
curl -X POST http://localhost:30000/v1/images/generations
-H "Content-Type: application/json"
-d '{
"prompt": "a detailed description...",
"use_pe": false,
"height": 1264,
"width": 848
}'
社区实测案例
案例 1:PE 开 vs 关 — 简短 prompt
Prompt:"a cyberpunk city street"
- PE 开:生成包含霓虹灯、雨夜、全息广告牌、蒸汽管道的完整赛博朋克场景
- PE 关:生成基本的城市街道,缺乏氛围细节
结论:简短 prompt 下,PE 开 > 关。
案例 2:PE 开 vs 关 — 多对象场景
Prompt:"a woman in red coat standing next to a man in blue suit, with a dog between them, in a park"
- PE 开:Reddit 用户报告"consistently missing people"——三人中可能只出现 2 人
- PE 关:指令遵循更精确,三人都出现在画面中
结论:多对象精确定位场景,PE 关 > 开。
案例 3:PE 开 vs 关 — 文字渲染
Prompt:"A movie poster with the text 'THE MATRIX' in green digital rain style"
- PE 开:文字渲染清晰,"THE MATRIX" 正确拼写,绿色数字雨效果
- PE 关:文字可能拼写错误或模糊
结论:文字渲染场景,PE 开 >> 关。
Turbo 模式下的 PE 行为
ERNIE-Image-Turbo 也支持 PE,但行为略有不同:
| 基准 | Turbo w/o PE | Turbo w/ PE |
|---|---|---|
| GENEval | — | 0.8510 |
| OneIG-EN | — | 0.8375 (text: 0.8351) |
Turbo 模式下的 PE 文字渲染得分(0.8351)低于 SFT 版本(0.9788),这是因为 Turbo 优先优化速度和美学。对于极致文字渲染需求,建议使用 SFT 版本 + PE。
批量生成中的 PE 策略
在批量生成场景(参见 EI-093),PE 策略应按 prompt 类型分组:
# 分组策略
short_prompts = ["a cat", "city street", "mountain landscape"] # → use_pe=True
detailed_prompts = ["A photorealistic..."] # → use_pe=False
text_prompts = ["poster with text 'SALE'"] # → use_pe=True
批量处理
for group, use_pe in [(short_prompts, True), (detailed_prompts, False), (text_prompts, True)]:
for prompt in group:
image = pipe(prompt=prompt, use_pe=use_pe, ...)
总结:PE 开关决策速查表
| 场景 | PE | 原因 |
|---|---|---|
| 简短 prompt(< 30 词) | ✅ 开 | 扩展信息量 |
| 详细 prompt(> 100 词) | ❌ 关 | 避免冗余冲突 |
| 文字渲染(海报/信息图) | ✅ 开 | 0.9788 渲染得分 |
| 多对象精确定位 | ❌ 关 | 保持指令遵循 |
| 需要精确构图控制 | ❌ 关 | 避免 PE 创意偏移 |
| 创意探索 / 风格不固定 | ✅ 开 | PE 补充细节 |
| Turbo 模式快速出图 | ⚠️ 看情况 | Turbo PE 文字得分较低 |
| 长 prompt 复杂场景 | ✅ 开 | LongTextBench 优势 |
核心原则:PE 不是"越好越开",而是"按需开关"。简短 prompt 和文字渲染场景开 PE,精确控制和多对象场景关 PE。