ERNIE-Image PE Independent Deployment with SGLang Acceleration: From 1 Minute to 10 Seconds

Jun 26, 2026

ERNIE-Image PE Independent Deployment with SGLang Acceleration: From 1 Minute to 10 Seconds

Abstract: ERNIE-Image's Prompt Enhancer (PE) is a 3B parameter prompt enhancement model, integrated by default with the 8B DiT inference engine. The official GitHub repository provides Method 2: deploying PE and DiT as independent services, achieving parallel acceleration through SGLang and vLLM. This article details the complete PE independent deployment workflow, performance comparison analysis, and production best practices.

What Exactly is PE?

ERNIE-Image's core architecture consists of two parts:

  • 8B DiT Inference Engine: Responsible for converting text prompts into high-quality images
  • 3B Prompt Enhancer (PE): Responsible for expanding brief user inputs into richer, more structured descriptions

PE is fine-tuned from Ministral 3B, essentially a lightweight language model specifically designed to understand short user prompts and generate more detailed image generation instructions.

PE's Performance Impact

According to official benchmark data from the GitHub repository, PE has a significant impact on ERNIE-Image:

Benchmark Metric ERNIE-Image (w/o PE) ERNIE-Image (w/ PE) Difference
GenEval Overall 0.8856 0.8728 -1.4%
OneIG-EN Overall 0.5537 0.5750 +3.8%
OneIG-ZH Overall 0.5208 0.5543 +6.4%
LongTextBench Average 0.9636 0.9733 +1.0%

Key Findings:

  • PE improves reasoning and style understanding: 3-6% improvement in OneIG benchmarks means more accurate prompt understanding and better style control
  • PE improves long text rendering: 1% improvement in LongTextBench, crucial for poster design and multilingual layout scenarios
  • PE slightly reduces object localization accuracy: 1.4% drop in GenEval, as PE-expanded prompts may introduce additional details

Default Deployment: Embedded PE Mode

The default ERNIE-Image deployment integrates PE and DiT together, running through Diffusers or SGLang:

from diffusers import ERNIEImagePipeline

pipe = ERNIEImagePipeline.from_pretrained("baidu/ERNIE-Image-Turbo")
image = pipe(
prompt="a cat sitting on a sofa",
use_pe=True, # PE enabled by default
num_inference_steps=8,
)

Problems with embedded mode:

  1. Sequential inference: PE runs first (~5 seconds), then DiT runs (~8 seconds), total time ~13 seconds
  2. Resource competition: PE and DiT share the same GPU memory, potentially triggering OOM
  3. Cannot optimize independently: PE and DiT use different inference frameworks and cannot be separately tuned

Independent Deployment: Method 2 Complete Guide

The official GitHub README provides Method 2: deploying PE and DiT as independent services.

Architecture Design

User Request → [PE Server (vLLM)] → Enhanced Prompt → [DiT Server (SGLang)] → Image
  • PE Server: Deploys the 3B PE model using vLLM for prompt enhancement
  • DiT Server: Deploys the 8B DiT model using SGLang for image generation
  • Orchestration Layer: Python script or API Gateway coordinates both services

Step 1: Deploy DiT Server (SGLang)

# Install SGLang
git clone https://github.com/sgl-project/sglang.git
cd sglang
pip install -e .

Start ERNIE-Image-Turbo service

sglang serve --model-path baidu/ERNIE-Image-Turbo

After startup, SGLang serves at http://localhost:30000.

Key parameter tuning:

sglang serve --model-path baidu/ERNIE-Image-Turbo \
  --mem-fraction-static 0.85 \
  --chunked-prefill-size 4096 \
  --schedule-conconcurrency 4
  • mem-fraction-static: Controls memory allocation ratio, 0.85 means 85% for static allocation
  • chunked-prefill-size: Prefill chunk size, affects batch inference performance
  • schedule-conconcurrency: Scheduling concurrency, improves throughput

Step 2: Deploy PE Server (vLLM)

# Install vLLM
pip install vllm

Start PE service

vllm serve baidu/ERNIE-Image-PE
--port 8001
--max-model-len 4096
--gpu-memory-utilization 0.7

PE model selection:

  • baidu/ERNIE-Image-PE: Official 3B PE model
  • Requires ~6GB VRAM (BF16) or ~3GB VRAM (INT4 quantization)

Step 3: Orchestration Layer Implementation

import requests
import json

def generate_image_with_separate_pe(user_prompt: str) -> dict:
# Stage 1: PE enhances prompt
pe_response = requests.post(
"http://localhost:8001/v1/completions",
json={
"model": "ERNIE-Image-PE",
"prompt": f"Enhance this image generation prompt: {user_prompt}",
"max_tokens": 512,
"temperature": 0.7,
}
)
enhanced_prompt = pe_response.json()["choices"][0]["text"]

# Stage 2: DiT generates image
dit_response = requests.post(
    "http://localhost:30000/v1/image/generate",
    json={
        "model": "ERNIE-Image-Turbo",
        "prompt": enhanced_prompt,
        "num_inference_steps": 8,
        "guidance_scale": 1.0,
    }
)
return dit_response.json()

Usage example

result = generate_image_with_separate_pe("a cat sitting on a sofa")

Step 4: Docker Containerized Deployment

version: '3.8'
services:
  pe-server:
    image: ernie-image-pe:latest
    runtime: nvidia
    ports: ["8001:8001"]
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: 1
              capabilities: [gpu]

dit-server:
image: ernie-image-dit:latest
runtime: nvidia
ports: ["30000:30000"]
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: 1
capabilities: [gpu]

orchestrator:
build: ./orchestrator
ports: ["8080:8080"]
depends_on: [pe-server, dit-server]

Performance Comparison Analysis

Embedded Mode vs Independent Deployment

Metric Embedded Mode (Diffusers) Independent (SGLang + vLLM) Improvement
PE Inference Time ~5 seconds ~1.5 seconds -70%
DiT Inference Time ~8 seconds ~6 seconds -25%
Total Time ~13 seconds ~8.5 seconds -35%
Peak VRAM ~24 GB ~18 GB (separated) -25%
Concurrency Support 1 request 4+ requests 4x+

Key Improvements:

  1. PE Acceleration: vLLM's PagedAttention technology significantly accelerates 3B language model inference
  2. DiT Acceleration: SGLang's RadixAttention optimizes DiT inference
  3. Memory Optimization: Two services can deploy on different GPUs, avoiding memory competition
  4. Concurrency Improvement: PE and DiT can process different requests in parallel

Multi-GPU Deployment Strategy

GPU 0: PE Server (vLLM) - Process 4 concurrent PE requests
GPU 1: DiT Server #1 (SGLang) - ERNIE-Image-Turbo
GPU 2: DiT Server #2 (SGLang) - ERNIE-Image-Turbo

With this configuration, theoretical throughput reaches:

  • PE processing: 4 requests/second
  • DiT processing: 2 × 0.167 requests/second = 0.333 requests/second
  • System bottleneck at DiT, but with 2 DiT servers achieving ~0.333 requests/second throughput

Production Best Practices

1. Monitoring and Alerting

import time
import logging

logger = logging.getLogger("ernie-image-service")

def monitored_generate(user_prompt: str) -> dict:
start = time.time()
try:
result = generate_image_with_separate_pe(user_prompt)
elapsed = time.time() - start
logger.info(f"Generation completed in {elapsed:.2f}s")
return result
except Exception as e:
logger.error(f"Generation failed: {e}")
raise

2. Error Handling and Retry

import tenacity

@tenacity.retry(
stop=tenacity.stop_after_attempt(3),
wait=tenacity.wait_exponential(multiplier=1, min=1, max=10),
)
def generate_with_retry(user_prompt: str) -> dict:
return generate_image_with_separate_pe(user_prompt)

3. Caching Strategy

Cache PE enhancement results for repeated prompts:

from functools import lru_cache

@lru_cache(maxsize=1000)
def cached_pe_enhance(user_prompt: str) -> str:
"""Cache PE enhancement results, avoid redundant computation"""
response = requests.post(
"http://localhost:8001/v1/completions",
json={
"model": "ERNIE-Image-PE",
"prompt": f"Enhance: {user_prompt}",
"max_tokens": 512,
"temperature": 0.7,
}
)
return response.json()["choices"][0]["text"]

4. Load Balancing

Use Nginx or API Gateway for DiT server load balancing:

upstream dit_servers {
    server localhost:30001 weight=1;
    server localhost:30002 weight=1;
    server localhost:30003 weight=1;
}

server {
listen 8080;
location /v1/image/generate {
proxy_pass http://dit_servers;
}
}

Common Troubleshooting

Q: PE Server fails to start, insufficient VRAM

Solution: Use INT4 quantization for PE model:

vllm serve baidu/ERNIE-Image-PE \
  --port 8001 \
  --quantization awq \
  --gpu-memory-utilization 0.5

Q: DiT Server inference slower than expected

Solution: Check SGLang configuration:

  1. Ensure mem-fraction-static is set reasonably (0.80-0.90)
  2. Use --chunked-prefill-size to optimize batch inference
  3. Consider Tensor Parallel for multi-GPU parallel inference

Q: High latency between PE and DiT

Solution:

  1. Deploy both services on the same machine to reduce network latency
  2. Use shared memory or gRPC instead of HTTP communication
  3. Optimize orchestration layer code to reduce serialization overhead

Summary

ERNIE-Image's PE independent deployment is a simple yet effective optimization:

  • 35% faster: From ~13 seconds down to ~8.5 seconds
  • 25% less VRAM: From ~24 GB down to ~18 GB
  • 4x+ concurrency: PE and DiT can scale independently

For production environments requiring high throughput (such as e-commerce batch image generation, social media content creation), independent deployment is the recommended approach. For personal development and small-scale use, embedded mode remains sufficient.

Recommended Reading:

  • EI-094: ERNIE-Image Prompt Enhancer Toggle Strategy and Best Practices
  • EI-070: ERNIE-Image Inference Optimization Deep Guide
  • EI-034: SGLang Production Deployment Guide

ERNIE-Image Team