ERNIE-Image Docker Containerization and Kubernetes Production Deployment Guide
From single-machine experiment to enterprise-grade production: Complete Docker containerization guide for ERNIE-Image, including Docker image building, Kubernetes orchestration, auto-scaling, and monitoring alerts.
Introduction
ERNIE-Image has become the go-to open-source text-to-image model with its excellent 8B-parameter performance and broad hardware compatibility. However, moving the model from experimental to production environments presents multiple challenges: environment consistency, resource isolation, elastic scaling, and high availability.
This guide provides a complete Docker containerization deployment solution and Kubernetes production orchestration guide, covering the full chain from basic image building to production-grade monitoring and alerting.
I. Docker Image Building
1.1 Base Image Design
# Dockerfile.ernie-image
FROM nvidia/cuda:12.4.1-devel-ubuntu22.04
Environment variables
ENV DEBIAN_FRONTEND=noninteractive
PYTHON_VERSION=3.11
PYTHONDONTWRITEBYTECODE=1
PIP_NO_CACHE_DIR=1
System dependencies
RUN apt-get update && apt-get install -y --no-install-recommends
python3${PYTHON_VERSION}
python3${PYTHON_VERSION}-venv
git
wget
curl
&& rm -rf /var/lib/apt/lists/*
Virtual environment
RUN python3 -m venv /opt/venv
ENV PATH="/opt/venv/bin:$PATH"
Python dependencies
RUN pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu121
RUN pip install diffusers transformers accelerate sentencepiece protobuf
SGLang (high-performance inference)
RUN pip install sglang
Working directory
WORKDIR /app
Application code
COPY scripts/ ./scripts/
COPY config/ ./config/
Model cache
RUN mkdir -p /models/ernie-image
Health check
HEALTHCHECK --interval=30s --timeout=10s --start-period=60s --retries=3
CMD python3 /app/scripts/health_check.py || exit 1
EXPOSE 8000
CMD ["python3", "/app/scripts/server.py"]
1.2 Multi-Stage Build Optimization
# Dockerfile.multi-stage
# Stage 1: Build dependencies
FROM nvidia/cuda:12.4.1-devel-ubuntu22.04 AS builder
ENV DEBIAN_FRONTEND=noninteractive
RUN apt-get update && apt-get install -y python3 python3-venv git
RUN python3 -m venv /opt/venv
ENV PATH="/opt/venv/bin:$PATH"
RUN pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu121
RUN pip install diffusers transformers accelerate sentencepiece protobuf sglang
Stage 2: Runtime
FROM nvidia/cuda:12.4.1-runtime-ubuntu22.04
ENV DEBIAN_FRONTEND=noninteractive
RUN apt-get update && apt-get install -y --no-install-recommends
python3 libgl1 && rm -rf /var/lib/apt/lists/*
COPY --from=builder /opt/venv /opt/venv
ENV PATH="/opt/venv/bin:$PATH"
WORKDIR /app
COPY scripts/ ./scripts/
COPY config/ ./config/
RUN mkdir -p /models/ernie-image
EXPOSE 8000
CMD ["python3", "/app/scripts/server.py"]
1.3 Image Tagging Strategy
| Tag | Purpose |
|---|---|
ernie-image-server:latest |
Latest stable |
ernie-image-server:1.0.0 |
Semantic versioning |
ernie-image-server:cu121-torch2.3 |
Specific CUDA/PyTorch version |
ernie-image-server:lite |
Lightweight (no training tools) |
II. Docker Compose Single-Node Deployment
2.1 docker-compose.yml
# docker-compose.yml
version: '3.8'
services:
ernie-image:
build:
context: .
dockerfile: Dockerfile.ernie-image
image: ernie-image-server:latest
container_name: ernie-image-server
restart: unless-stopped
ports:
- "8000:8000"
volumes:
- ./models:/models/ernie-image
- ./outputs:/app/outputs
- ./config:/app/config:ro
environment:
- MODEL_PATH=/models/ernie-image
- OUTPUT_DIR=/app/outputs
- MAX_BATCH_SIZE=8
- CFG_SCALE=4.0
- NUM_INFERENCE_STEPS=50
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: 1
capabilities: [gpu]
limits:
memory: 24g
healthcheck:
test: ["CMD", "python3", "/app/scripts/health_check.py"]
interval: 30s
timeout: 10s
retries: 3
networks:
- ernie-network
nginx:
image: nginx:alpine
container_name: ernie-image-nginx
restart: unless-stopped
ports:
- "80:80"
- "443:443"
volumes:
- ./nginx/nginx.conf:/etc/nginx/nginx.conf:ro
- ./nginx/ssl:/etc/nginx/ssl:ro
depends_on:
- ernie-image
networks:
- ernie-network
networks:
ernie-network:
driver: bridge
2.2 Start and Verify
# Pull model
huggingface-cli download baidu/ERNIE-Image --local-dir /path/to/models/ernie-image
Start services
docker compose up -d
Check status
docker compose ps
Expected: ernie-image-server (healthy), ernie-image-nginx (healthy)
View logs
docker compose logs -f ernie-image
Test endpoint
curl -X POST http://localhost:8000/generate
-H "Content-Type: application/json"
-d '{"prompt": "A cat sitting on a chair, photorealistic style", "steps": 50, "cfg_scale": 4.0}'
III. Kubernetes Production Deployment
3.1 Namespace and Resource Quotas
# k8s/namespace.yaml
apiVersion: v1
kind: Namespace
metadata:
name: ernie-image
labels:
team: ai-infra
environment: production
---
apiVersion: v1
kind: ResourceQuota
metadata:
name: ernie-image-quota
namespace: ernie-image
spec:
hard:
requests.cpu: "32"
requests.memory: 128Gi
requests.nvidia.com/gpu: "8"
limits.cpu: "64"
limits.memory: 256Gi
limits.nvidia.com/gpu: "8"
3.2 Persistent Storage
# k8s/pvc.yaml
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: ernie-models-pvc
namespace: ernie-image
spec:
accessModes:
- ReadOnlyMany
resources:
requests:
storage: 50Gi
storageClassName: gp3
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: ernie-outputs-pvc
namespace: ernie-image
spec:
accessModes:
- ReadWriteMany
resources:
requests:
storage: 100Gi
storageClassName: gp3
3.3 Deployment Configuration
# k8s/deployment.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
name: ernie-image-server
namespace: ernie-image
labels:
app: ernie-image
version: v1.0.0
spec:
replicas: 3
strategy:
type: RollingUpdate
rollingUpdate:
maxSurge: 1
maxUnavailable: 0
selector:
matchLabels:
app: ernie-image
template:
metadata:
labels:
app: ernie-image
version: v1.0.0
annotations:
prometheus.io/scrape: "true"
prometheus.io/port: "8000"
spec:
containers:
- name: ernie-image
image: ernie-image-server:1.0.0
ports:
- containerPort: 8000
name: http
env:
- name: MODEL_PATH
value: "/models/ernie-image"
- name: OUTPUT_DIR
value: "/outputs"
- name: MAX_BATCH_SIZE
value: "8"
- name: NUM_WORKERS
value: "4"
resources:
requests:
cpu: "4"
memory: 24Gi
nvidia.com/gpu: "1"
limits:
cpu: "8"
memory: 48Gi
nvidia.com/gpu: "1"
volumeMounts:
- name: models
mountPath: /models/ernie-image
readOnly: true
- name: outputs
mountPath: /outputs
livenessProbe:
httpGet:
path: /health
port: 8000
initialDelaySeconds: 60
periodSeconds: 30
timeoutSeconds: 10
failureThreshold: 3
readinessProbe:
httpGet:
path: /ready
port: 8000
initialDelaySeconds: 30
periodSeconds: 15
timeoutSeconds: 5
volumes:
- name: models
persistentVolumeClaim:
claimName: ernie-models-pvc
- name: outputs
persistentVolumeClaim:
claimName: ernie-outputs-pvc
affinity:
podAntiAffinity:
preferredDuringSchedulingIgnoredDuringExecution:
- weight: 100
podAffinityTerm:
labelSelector:
matchExpressions:
- key: app
operator: In
values:
- ernie-image
topologyKey: kubernetes.io/hostname
tolerations:
- key: nvidia.com/gpu
operator: Exists
effect: NoSchedule
3.4 Service and Ingress
# k8s/service.yaml
apiVersion: v1
kind: Service
metadata:
name: ernie-image-service
namespace: ernie-image
spec:
type: ClusterIP
ports:
- port: 80
targetPort: 8000
protocol: TCP
selector:
app: ernie-image
---
# k8s/ingress.yaml
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
name: ernie-image-ingress
namespace: ernie-image
annotations:
nginx.ingress.kubernetes.io/ssl-redirect: "true"
nginx.ingress.kubernetes.io/proxy-body-size: "50m"
cert-manager.io/cluster-issuer: "letsencrypt-prod"
spec:
ingressClassName: nginx
tls:
- hosts:
- api.ernie-image.example.com
secretName: ernie-image-tls
rules:
- host: api.ernie-image.example.com
http:
paths:
- path: /
pathType: Prefix
backend:
service:
name: ernie-image-service
port:
number: 80
3.5 Auto-Scaling (HPA)
# k8s/hpa.yaml
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: ernie-image-hpa
namespace: ernie-image
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: ernie-image-server
minReplicas: 1
maxReplicas: 10
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 70
- type: Resource
resource:
name: memory
target:
type: Utilization
averageUtilization: 80
- type: Pods
pods:
metric:
name: requests-per-pod
target:
type: AverageValue
averageValue: "100"
3.6 GPU Node Auto-Scaling (KEDA)
# k8s/keda-scaled-object.yaml
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
name: ernie-image-keda
namespace: ernie-image
spec:
scaleTargetRef:
name: ernie-image-server
minReplicaCount: 1
maxReplicaCount: 10
triggers:
- type: prometheus
metricName: ernie_queue_length
threshold: "20"
serverAddress: http://prometheus:9090
IV. Monitoring and Alerting
4.1 Prometheus Metrics
# metrics.py - Prometheus metrics in application
from prometheus_client import Counter, Histogram, Gauge
REQUEST_COUNT = Counter(
'ernie_image_requests_total',
'Total image generation requests',
['status', 'model_variant']
)
REQUEST_LATENCY = Histogram(
'ernie_image_request_latency_seconds',
'Request latency',
['model_variant']
)
QUEUE_LENGTH = Gauge(
'ernie_image_queue_length',
'Current queue length'
)
GPU_UTILIZATION = Gauge(
'ernie_image_gpu_utilization_percent',
'GPU utilization',
['gpu_id']
)
ACTIVE_REQUESTS = Gauge(
'ernie_image_active_requests',
'Currently active requests'
)
4.2 Grafana Dashboard
Recommended panel layout:
┌─────────────────┬─────────────────┬─────────────────┐
│ QPS (req/s) │ P99 Latency │ GPU Util % │
├─────────────────┼─────────────────┼─────────────────┤
│ Queue Length │ Active Pods │ Error Rate % │
├─────────────────┼─────────────────┼─────────────────┤
│ Memory (GB) │ VRAM (GB) │ Throughput │
└─────────────────┴─────────────────┴─────────────────┘
4.3 Alert Rules
# k8s/alerts.yaml
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: ernie-image-alerts
namespace: monitoring
spec:
groups:
- name: ernie-image
rules:
- alert: HighQueueLength
expr: ernie_image_queue_length > 50
for: 2m
labels:
severity: warning
- alert: HighLatency
expr: histogram_quantile(0.99, rate(ernie_image_request_latency_seconds_bucket[5m])) > 30
for: 5m
labels:
severity: warning
- alert: HighErrorRate
expr: rate(ernie_image_requests_total{status="error"}[5m]) / rate(ernie_image_requests_total[5m]) > 0.05
for: 5m
labels:
severity: critical
- alert: GPUOutOfMemory
expr: ernie_image_gpu_utilization_percent > 95
for: 1m
labels:
severity: critical
V. Configuration Management
5.1 Kubernetes ConfigMap
# k8s/configmap.yaml
apiVersion: v1
kind: ConfigMap
metadata:
name: ernie-image-config
namespace: ernie-image
data:
model_variant: "ERNIE-Image-Turbo"
max_batch_size: "8"
num_inference_steps: "8"
cfg_scale: "4.0"
output_format: "jpeg"
output_quality: "95"
max_image_size: "2048x2048"
enable_pe: "true"
timeout_seconds: "120"
max_queue_size: "100"
5.2 Multi-Environment Configuration
# environments:
# development:
# replicas: 1
# gpu: 1 (RTX 3090)
# max_batch_size: 2
# staging:
# replicas: 2
# gpu: 2 (A100 40GB)
# max_batch_size: 4
# production:
# replicas: 3-10 (HPA)
# gpu: 3-10 (A100 80GB)
# max_batch_size: 8
VI. CI/CD Pipeline
6.1 GitLab CI Example
# .gitlab-ci.yml
stages:
- build
- test
- deploy
variables:
IMAGE_NAME: registry.example.com/ernie-image-server
IMAGE_TAG: $CI_COMMIT_SHORT_SHA
build:
stage: build
script:
- docker build -t $IMAGE_NAME:$IMAGE_TAG .
- docker push $IMAGE_NAME:$IMAGE_TAG
tags:
- docker-builder
test:
stage: test
script:
- docker run --gpus all $IMAGE_NAME:$IMAGE_TAG python3 -m pytest tests/
tags:
- gpu-runner
deploy-staging:
stage: deploy
script:
- kubectl set image deployment/ernie-image-server
ernie-image=$IMAGE_NAME:$IMAGE_TAG -n ernie-image-staging
environment:
name: staging
tags:
- k8s-admin
deploy-production:
stage: deploy
script:
- kubectl set image deployment/ernie-image-server
ernie-image=$IMAGE_NAME:$IMAGE_TAG -n ernie-image
environment:
name: production
when: manual
tags:
- k8s-admin
VII. Troubleshooting
7.1 Common Issues
| Issue | Cause | Solution |
|---|---|---|
| GPU OOM | Batch size too large | Reduce MAX_BATCH_SIZE, use NVFP4 quantization |
| Inference timeout | Queue backlog | Check HPA config, increase replicas |
| Slow model loading | Cold start | Pre-load models to PVC, use readinessProbe |
| VRAM leak | Uncleaned cache | Restart pods periodically, torch.cuda.empty_cache() |
| Slow cross-GPU | No NVLink | Use NVLink or InfiniBand |
7.2 Log Viewing
# View pod logs
kubectl logs -n ernie-image deployment/ernie-image-server --tail=100
Real-time logs
kubectl logs -f -n ernie-image -l app=ernie-image
Previous pod logs (crash debugging)
kubectl logs -n ernie-image deployment/ernie-image-server --previous
VIII. Cost Optimization
8.1 Quantization Savings
| Deployment | VRAM | Speed | Cost |
|---|---|---|---|
| BF16 Full Precision | ~17GB | Baseline | 100% |
| FP8 Quantized | ~9GB | ~95% | ~60% |
| NVFP4 Quantized | ~5GB | ~90% | ~35% |
| INT8 Quantized | ~9GB | ~85% | ~50% |
8.2 GPU Spot Instances
Using Spot instances in Kubernetes saves 60-70% on GPU costs:
# k8s/node-pool-spot.yaml
apiVersion: karpenter.sh/v1alpha5
kind: NodePool
metadata:
name: ernie-gpu-spot
spec:
template:
spec:
requirements:
- key: nvidia.com/gpu
operator: In
values: ["A100-80GB"]
spot: true
IX. Summary
The Docker containerization + Kubernetes orchestration approach for ERNIE-Image provides an enterprise-grade production deployment solution:
- Docker images: Environment isolation, reproducibility, easy distribution
- Kubernetes: Auto-scaling, high availability, rolling updates
- Monitoring: Real-time metrics, intelligent alerts, failure prediction
- CI/CD: Automated building, testing, deployment
- Cost optimization: Quantization + Spot instances + auto-scaling
This approach works for everything from SMB single-node deployments to large enterprise multi-cluster production environments — the best practice for taking ERNIE-Image from experiment to production.
Further Reading: