ERNIE-Image PE Independent Deployment with SGLang Acceleration: From 1 Minute to 10 Seconds
Abstract: ERNIE-Image's Prompt Enhancer (PE) is a 3B parameter prompt enhancement model, integrated by default with the 8B DiT inference engine. The official GitHub repository provides Method 2: deploying PE and DiT as independent services, achieving parallel acceleration through SGLang and vLLM. This article details the complete PE independent deployment workflow, performance comparison analysis, and production best practices.
What Exactly is PE?
ERNIE-Image's core architecture consists of two parts:
- 8B DiT Inference Engine: Responsible for converting text prompts into high-quality images
- 3B Prompt Enhancer (PE): Responsible for expanding brief user inputs into richer, more structured descriptions
PE is fine-tuned from Ministral 3B, essentially a lightweight language model specifically designed to understand short user prompts and generate more detailed image generation instructions.
PE's Performance Impact
According to official benchmark data from the GitHub repository, PE has a significant impact on ERNIE-Image:
| Benchmark | Metric | ERNIE-Image (w/o PE) | ERNIE-Image (w/ PE) | Difference |
|---|---|---|---|---|
| GenEval | Overall | 0.8856 | 0.8728 | -1.4% |
| OneIG-EN | Overall | 0.5537 | 0.5750 | +3.8% |
| OneIG-ZH | Overall | 0.5208 | 0.5543 | +6.4% |
| LongTextBench | Average | 0.9636 | 0.9733 | +1.0% |
Key Findings:
- PE improves reasoning and style understanding: 3-6% improvement in OneIG benchmarks means more accurate prompt understanding and better style control
- PE improves long text rendering: 1% improvement in LongTextBench, crucial for poster design and multilingual layout scenarios
- PE slightly reduces object localization accuracy: 1.4% drop in GenEval, as PE-expanded prompts may introduce additional details
Default Deployment: Embedded PE Mode
The default ERNIE-Image deployment integrates PE and DiT together, running through Diffusers or SGLang:
from diffusers import ERNIEImagePipeline
pipe = ERNIEImagePipeline.from_pretrained("baidu/ERNIE-Image-Turbo")
image = pipe(
prompt="a cat sitting on a sofa",
use_pe=True, # PE enabled by default
num_inference_steps=8,
)
Problems with embedded mode:
- Sequential inference: PE runs first (~5 seconds), then DiT runs (~8 seconds), total time ~13 seconds
- Resource competition: PE and DiT share the same GPU memory, potentially triggering OOM
- Cannot optimize independently: PE and DiT use different inference frameworks and cannot be separately tuned
Independent Deployment: Method 2 Complete Guide
The official GitHub README provides Method 2: deploying PE and DiT as independent services.
Architecture Design
User Request → [PE Server (vLLM)] → Enhanced Prompt → [DiT Server (SGLang)] → Image
- PE Server: Deploys the 3B PE model using vLLM for prompt enhancement
- DiT Server: Deploys the 8B DiT model using SGLang for image generation
- Orchestration Layer: Python script or API Gateway coordinates both services
Step 1: Deploy DiT Server (SGLang)
# Install SGLang
git clone https://github.com/sgl-project/sglang.git
cd sglang
pip install -e .
Start ERNIE-Image-Turbo service
sglang serve --model-path baidu/ERNIE-Image-Turbo
After startup, SGLang serves at http://localhost:30000.
Key parameter tuning:
sglang serve --model-path baidu/ERNIE-Image-Turbo \
--mem-fraction-static 0.85 \
--chunked-prefill-size 4096 \
--schedule-conconcurrency 4
mem-fraction-static: Controls memory allocation ratio, 0.85 means 85% for static allocationchunked-prefill-size: Prefill chunk size, affects batch inference performanceschedule-conconcurrency: Scheduling concurrency, improves throughput
Step 2: Deploy PE Server (vLLM)
# Install vLLM
pip install vllm
Start PE service
vllm serve baidu/ERNIE-Image-PE
--port 8001
--max-model-len 4096
--gpu-memory-utilization 0.7
PE model selection:
baidu/ERNIE-Image-PE: Official 3B PE model- Requires ~6GB VRAM (BF16) or ~3GB VRAM (INT4 quantization)
Step 3: Orchestration Layer Implementation
import requests
import json
def generate_image_with_separate_pe(user_prompt: str) -> dict:
# Stage 1: PE enhances prompt
pe_response = requests.post(
"http://localhost:8001/v1/completions",
json={
"model": "ERNIE-Image-PE",
"prompt": f"Enhance this image generation prompt: {user_prompt}",
"max_tokens": 512,
"temperature": 0.7,
}
)
enhanced_prompt = pe_response.json()["choices"][0]["text"]
# Stage 2: DiT generates image
dit_response = requests.post(
"http://localhost:30000/v1/image/generate",
json={
"model": "ERNIE-Image-Turbo",
"prompt": enhanced_prompt,
"num_inference_steps": 8,
"guidance_scale": 1.0,
}
)
return dit_response.json()
Usage example
result = generate_image_with_separate_pe("a cat sitting on a sofa")
Step 4: Docker Containerized Deployment
version: '3.8'
services:
pe-server:
image: ernie-image-pe:latest
runtime: nvidia
ports: ["8001:8001"]
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: 1
capabilities: [gpu]
dit-server:
image: ernie-image-dit:latest
runtime: nvidia
ports: ["30000:30000"]
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: 1
capabilities: [gpu]
orchestrator:
build: ./orchestrator
ports: ["8080:8080"]
depends_on: [pe-server, dit-server]
Performance Comparison Analysis
Embedded Mode vs Independent Deployment
| Metric | Embedded Mode (Diffusers) | Independent (SGLang + vLLM) | Improvement |
|---|---|---|---|
| PE Inference Time | ~5 seconds | ~1.5 seconds | -70% |
| DiT Inference Time | ~8 seconds | ~6 seconds | -25% |
| Total Time | ~13 seconds | ~8.5 seconds | -35% |
| Peak VRAM | ~24 GB | ~18 GB (separated) | -25% |
| Concurrency Support | 1 request | 4+ requests | 4x+ |
Key Improvements:
- PE Acceleration: vLLM's PagedAttention technology significantly accelerates 3B language model inference
- DiT Acceleration: SGLang's RadixAttention optimizes DiT inference
- Memory Optimization: Two services can deploy on different GPUs, avoiding memory competition
- Concurrency Improvement: PE and DiT can process different requests in parallel
Multi-GPU Deployment Strategy
GPU 0: PE Server (vLLM) - Process 4 concurrent PE requests
GPU 1: DiT Server #1 (SGLang) - ERNIE-Image-Turbo
GPU 2: DiT Server #2 (SGLang) - ERNIE-Image-Turbo
With this configuration, theoretical throughput reaches:
- PE processing: 4 requests/second
- DiT processing: 2 × 0.167 requests/second = 0.333 requests/second
- System bottleneck at DiT, but with 2 DiT servers achieving ~0.333 requests/second throughput
Production Best Practices
1. Monitoring and Alerting
import time
import logging
logger = logging.getLogger("ernie-image-service")
def monitored_generate(user_prompt: str) -> dict:
start = time.time()
try:
result = generate_image_with_separate_pe(user_prompt)
elapsed = time.time() - start
logger.info(f"Generation completed in {elapsed:.2f}s")
return result
except Exception as e:
logger.error(f"Generation failed: {e}")
raise
2. Error Handling and Retry
import tenacity
@tenacity.retry(
stop=tenacity.stop_after_attempt(3),
wait=tenacity.wait_exponential(multiplier=1, min=1, max=10),
)
def generate_with_retry(user_prompt: str) -> dict:
return generate_image_with_separate_pe(user_prompt)
3. Caching Strategy
Cache PE enhancement results for repeated prompts:
from functools import lru_cache
@lru_cache(maxsize=1000)
def cached_pe_enhance(user_prompt: str) -> str:
"""Cache PE enhancement results, avoid redundant computation"""
response = requests.post(
"http://localhost:8001/v1/completions",
json={
"model": "ERNIE-Image-PE",
"prompt": f"Enhance: {user_prompt}",
"max_tokens": 512,
"temperature": 0.7,
}
)
return response.json()["choices"][0]["text"]
4. Load Balancing
Use Nginx or API Gateway for DiT server load balancing:
upstream dit_servers {
server localhost:30001 weight=1;
server localhost:30002 weight=1;
server localhost:30003 weight=1;
}
server {
listen 8080;
location /v1/image/generate {
proxy_pass http://dit_servers;
}
}
Common Troubleshooting
Q: PE Server fails to start, insufficient VRAM
Solution: Use INT4 quantization for PE model:
vllm serve baidu/ERNIE-Image-PE \
--port 8001 \
--quantization awq \
--gpu-memory-utilization 0.5
Q: DiT Server inference slower than expected
Solution: Check SGLang configuration:
- Ensure
mem-fraction-staticis set reasonably (0.80-0.90) - Use
--chunked-prefill-sizeto optimize batch inference - Consider Tensor Parallel for multi-GPU parallel inference
Q: High latency between PE and DiT
Solution:
- Deploy both services on the same machine to reduce network latency
- Use shared memory or gRPC instead of HTTP communication
- Optimize orchestration layer code to reduce serialization overhead
Summary
ERNIE-Image's PE independent deployment is a simple yet effective optimization:
- 35% faster: From ~13 seconds down to ~8.5 seconds
- 25% less VRAM: From ~24 GB down to ~18 GB
- 4x+ concurrency: PE and DiT can scale independently
For production environments requiring high throughput (such as e-commerce batch image generation, social media content creation), independent deployment is the recommended approach. For personal development and small-scale use, embedded mode remains sufficient.
Recommended Reading:
- EI-094: ERNIE-Image Prompt Enhancer Toggle Strategy and Best Practices
- EI-070: ERNIE-Image Inference Optimization Deep Guide
- EI-034: SGLang Production Deployment Guide