ERNIE-Image Docker Containerization and Kubernetes Production Deployment Guide

Jun 11, 2026

ERNIE-Image Docker Containerization and Kubernetes Production Deployment Guide

From single-machine experiment to enterprise-grade production: Complete Docker containerization guide for ERNIE-Image, including Docker image building, Kubernetes orchestration, auto-scaling, and monitoring alerts.

Introduction

ERNIE-Image has become the go-to open-source text-to-image model with its excellent 8B-parameter performance and broad hardware compatibility. However, moving the model from experimental to production environments presents multiple challenges: environment consistency, resource isolation, elastic scaling, and high availability.

This guide provides a complete Docker containerization deployment solution and Kubernetes production orchestration guide, covering the full chain from basic image building to production-grade monitoring and alerting.

I. Docker Image Building

1.1 Base Image Design

# Dockerfile.ernie-image
FROM nvidia/cuda:12.4.1-devel-ubuntu22.04

Environment variables

ENV DEBIAN_FRONTEND=noninteractive
PYTHON_VERSION=3.11
PYTHONDONTWRITEBYTECODE=1
PIP_NO_CACHE_DIR=1

System dependencies

RUN apt-get update && apt-get install -y --no-install-recommends
python3${PYTHON_VERSION}
python3${PYTHON_VERSION}-venv
git
wget
curl
&& rm -rf /var/lib/apt/lists/*

Virtual environment

RUN python3 -m venv /opt/venv
ENV PATH="/opt/venv/bin:$PATH"

Python dependencies

RUN pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu121
RUN pip install diffusers transformers accelerate sentencepiece protobuf

SGLang (high-performance inference)

RUN pip install sglang

Working directory

WORKDIR /app

Application code

COPY scripts/ ./scripts/
COPY config/ ./config/

Model cache

RUN mkdir -p /models/ernie-image

Health check

HEALTHCHECK --interval=30s --timeout=10s --start-period=60s --retries=3
CMD python3 /app/scripts/health_check.py || exit 1

EXPOSE 8000
CMD ["python3", "/app/scripts/server.py"]

1.2 Multi-Stage Build Optimization

# Dockerfile.multi-stage
# Stage 1: Build dependencies
FROM nvidia/cuda:12.4.1-devel-ubuntu22.04 AS builder
ENV DEBIAN_FRONTEND=noninteractive
RUN apt-get update && apt-get install -y python3 python3-venv git
RUN python3 -m venv /opt/venv
ENV PATH="/opt/venv/bin:$PATH"
RUN pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu121
RUN pip install diffusers transformers accelerate sentencepiece protobuf sglang

Stage 2: Runtime

FROM nvidia/cuda:12.4.1-runtime-ubuntu22.04
ENV DEBIAN_FRONTEND=noninteractive
RUN apt-get update && apt-get install -y --no-install-recommends
python3 libgl1 && rm -rf /var/lib/apt/lists/*
COPY --from=builder /opt/venv /opt/venv
ENV PATH="/opt/venv/bin:$PATH"
WORKDIR /app
COPY scripts/ ./scripts/
COPY config/ ./config/
RUN mkdir -p /models/ernie-image
EXPOSE 8000
CMD ["python3", "/app/scripts/server.py"]

1.3 Image Tagging Strategy

Tag Purpose
ernie-image-server:latest Latest stable
ernie-image-server:1.0.0 Semantic versioning
ernie-image-server:cu121-torch2.3 Specific CUDA/PyTorch version
ernie-image-server:lite Lightweight (no training tools)

II. Docker Compose Single-Node Deployment

2.1 docker-compose.yml

# docker-compose.yml
version: '3.8'

services:
ernie-image:
build:
context: .
dockerfile: Dockerfile.ernie-image
image: ernie-image-server:latest
container_name: ernie-image-server
restart: unless-stopped
ports:
- "8000:8000"
volumes:
- ./models:/models/ernie-image
- ./outputs:/app/outputs
- ./config:/app/config:ro
environment:
- MODEL_PATH=/models/ernie-image
- OUTPUT_DIR=/app/outputs
- MAX_BATCH_SIZE=8
- CFG_SCALE=4.0
- NUM_INFERENCE_STEPS=50
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: 1
capabilities: [gpu]
limits:
memory: 24g
healthcheck:
test: ["CMD", "python3", "/app/scripts/health_check.py"]
interval: 30s
timeout: 10s
retries: 3
networks:
- ernie-network

nginx:
image: nginx:alpine
container_name: ernie-image-nginx
restart: unless-stopped
ports:
- "80:80"
- "443:443"
volumes:
- ./nginx/nginx.conf:/etc/nginx/nginx.conf:ro
- ./nginx/ssl:/etc/nginx/ssl:ro
depends_on:
- ernie-image
networks:
- ernie-network

networks:
ernie-network:
driver: bridge

2.2 Start and Verify

# Pull model
huggingface-cli download baidu/ERNIE-Image --local-dir /path/to/models/ernie-image

Start services

docker compose up -d

Check status

docker compose ps

Expected: ernie-image-server (healthy), ernie-image-nginx (healthy)

View logs

docker compose logs -f ernie-image

Test endpoint

curl -X POST http://localhost:8000/generate
-H "Content-Type: application/json"
-d '{"prompt": "A cat sitting on a chair, photorealistic style", "steps": 50, "cfg_scale": 4.0}'

III. Kubernetes Production Deployment

3.1 Namespace and Resource Quotas

# k8s/namespace.yaml
apiVersion: v1
kind: Namespace
metadata:
  name: ernie-image
  labels:
    team: ai-infra
    environment: production
---
apiVersion: v1
kind: ResourceQuota
metadata:
  name: ernie-image-quota
  namespace: ernie-image
spec:
  hard:
    requests.cpu: "32"
    requests.memory: 128Gi
    requests.nvidia.com/gpu: "8"
    limits.cpu: "64"
    limits.memory: 256Gi
    limits.nvidia.com/gpu: "8"

3.2 Persistent Storage

# k8s/pvc.yaml
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: ernie-models-pvc
  namespace: ernie-image
spec:
  accessModes:
    - ReadOnlyMany
  resources:
    requests:
      storage: 50Gi
  storageClassName: gp3
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: ernie-outputs-pvc
  namespace: ernie-image
spec:
  accessModes:
    - ReadWriteMany
  resources:
    requests:
      storage: 100Gi
  storageClassName: gp3

3.3 Deployment Configuration

# k8s/deployment.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
  name: ernie-image-server
  namespace: ernie-image
  labels:
    app: ernie-image
    version: v1.0.0
spec:
  replicas: 3
  strategy:
    type: RollingUpdate
    rollingUpdate:
      maxSurge: 1
      maxUnavailable: 0
  selector:
    matchLabels:
      app: ernie-image
  template:
    metadata:
      labels:
        app: ernie-image
        version: v1.0.0
      annotations:
        prometheus.io/scrape: "true"
        prometheus.io/port: "8000"
    spec:
      containers:
        - name: ernie-image
          image: ernie-image-server:1.0.0
          ports:
            - containerPort: 8000
              name: http
          env:
            - name: MODEL_PATH
              value: "/models/ernie-image"
            - name: OUTPUT_DIR
              value: "/outputs"
            - name: MAX_BATCH_SIZE
              value: "8"
            - name: NUM_WORKERS
              value: "4"
          resources:
            requests:
              cpu: "4"
              memory: 24Gi
              nvidia.com/gpu: "1"
            limits:
              cpu: "8"
              memory: 48Gi
              nvidia.com/gpu: "1"
          volumeMounts:
            - name: models
              mountPath: /models/ernie-image
              readOnly: true
            - name: outputs
              mountPath: /outputs
          livenessProbe:
            httpGet:
              path: /health
              port: 8000
            initialDelaySeconds: 60
            periodSeconds: 30
            timeoutSeconds: 10
            failureThreshold: 3
          readinessProbe:
            httpGet:
              path: /ready
              port: 8000
            initialDelaySeconds: 30
            periodSeconds: 15
            timeoutSeconds: 5
      volumes:
        - name: models
          persistentVolumeClaim:
            claimName: ernie-models-pvc
        - name: outputs
          persistentVolumeClaim:
            claimName: ernie-outputs-pvc
      affinity:
        podAntiAffinity:
          preferredDuringSchedulingIgnoredDuringExecution:
            - weight: 100
              podAffinityTerm:
                labelSelector:
                  matchExpressions:
                    - key: app
                      operator: In
                      values:
                        - ernie-image
                topologyKey: kubernetes.io/hostname
      tolerations:
        - key: nvidia.com/gpu
          operator: Exists
          effect: NoSchedule

3.4 Service and Ingress

# k8s/service.yaml
apiVersion: v1
kind: Service
metadata:
  name: ernie-image-service
  namespace: ernie-image
spec:
  type: ClusterIP
  ports:
    - port: 80
      targetPort: 8000
      protocol: TCP
  selector:
    app: ernie-image
---
# k8s/ingress.yaml
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
  name: ernie-image-ingress
  namespace: ernie-image
  annotations:
    nginx.ingress.kubernetes.io/ssl-redirect: "true"
    nginx.ingress.kubernetes.io/proxy-body-size: "50m"
    cert-manager.io/cluster-issuer: "letsencrypt-prod"
spec:
  ingressClassName: nginx
  tls:
    - hosts:
        - api.ernie-image.example.com
      secretName: ernie-image-tls
  rules:
    - host: api.ernie-image.example.com
      http:
        paths:
          - path: /
            pathType: Prefix
            backend:
              service:
                name: ernie-image-service
                port:
                  number: 80

3.5 Auto-Scaling (HPA)

# k8s/hpa.yaml
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: ernie-image-hpa
  namespace: ernie-image
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: ernie-image-server
  minReplicas: 1
  maxReplicas: 10
  metrics:
    - type: Resource
      resource:
        name: cpu
        target:
          type: Utilization
          averageUtilization: 70
    - type: Resource
      resource:
        name: memory
        target:
          type: Utilization
          averageUtilization: 80
    - type: Pods
      pods:
        metric:
          name: requests-per-pod
        target:
          type: AverageValue
          averageValue: "100"

3.6 GPU Node Auto-Scaling (KEDA)

# k8s/keda-scaled-object.yaml
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
  name: ernie-image-keda
  namespace: ernie-image
spec:
  scaleTargetRef:
    name: ernie-image-server
  minReplicaCount: 1
  maxReplicaCount: 10
  triggers:
    - type: prometheus
      metricName: ernie_queue_length
      threshold: "20"
      serverAddress: http://prometheus:9090

IV. Monitoring and Alerting

4.1 Prometheus Metrics

# metrics.py - Prometheus metrics in application
from prometheus_client import Counter, Histogram, Gauge

REQUEST_COUNT = Counter(
'ernie_image_requests_total',
'Total image generation requests',
['status', 'model_variant']
)

REQUEST_LATENCY = Histogram(
'ernie_image_request_latency_seconds',
'Request latency',
['model_variant']
)

QUEUE_LENGTH = Gauge(
'ernie_image_queue_length',
'Current queue length'
)

GPU_UTILIZATION = Gauge(
'ernie_image_gpu_utilization_percent',
'GPU utilization',
['gpu_id']
)

ACTIVE_REQUESTS = Gauge(
'ernie_image_active_requests',
'Currently active requests'
)

4.2 Grafana Dashboard

Recommended panel layout:

┌─────────────────┬─────────────────┬─────────────────┐
│  QPS (req/s)    │  P99 Latency    │  GPU Util %     │
├─────────────────┼─────────────────┼─────────────────┤
│  Queue Length   │  Active Pods    │  Error Rate %   │
├─────────────────┼─────────────────┼─────────────────┤
│  Memory (GB)    │  VRAM (GB)      │  Throughput     │
└─────────────────┴─────────────────┴─────────────────┘

4.3 Alert Rules

# k8s/alerts.yaml
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
  name: ernie-image-alerts
  namespace: monitoring
spec:
  groups:
    - name: ernie-image
      rules:
        - alert: HighQueueLength
          expr: ernie_image_queue_length > 50
          for: 2m
          labels:
            severity: warning
    - alert: HighLatency
      expr: histogram_quantile(0.99, rate(ernie_image_request_latency_seconds_bucket[5m])) > 30
      for: 5m
      labels:
        severity: warning

    - alert: HighErrorRate
      expr: rate(ernie_image_requests_total{status="error"}[5m]) / rate(ernie_image_requests_total[5m]) > 0.05
      for: 5m
      labels:
        severity: critical

    - alert: GPUOutOfMemory
      expr: ernie_image_gpu_utilization_percent > 95
      for: 1m
      labels:
        severity: critical

V. Configuration Management

5.1 Kubernetes ConfigMap

# k8s/configmap.yaml
apiVersion: v1
kind: ConfigMap
metadata:
  name: ernie-image-config
  namespace: ernie-image
data:
  model_variant: "ERNIE-Image-Turbo"
  max_batch_size: "8"
  num_inference_steps: "8"
  cfg_scale: "4.0"
  output_format: "jpeg"
  output_quality: "95"
  max_image_size: "2048x2048"
  enable_pe: "true"
  timeout_seconds: "120"
  max_queue_size: "100"

5.2 Multi-Environment Configuration

# environments:
#   development:
#     replicas: 1
#     gpu: 1 (RTX 3090)
#     max_batch_size: 2
#   staging:
#     replicas: 2
#     gpu: 2 (A100 40GB)
#     max_batch_size: 4
#   production:
#     replicas: 3-10 (HPA)
#     gpu: 3-10 (A100 80GB)
#     max_batch_size: 8

VI. CI/CD Pipeline

6.1 GitLab CI Example

# .gitlab-ci.yml
stages:
  - build
  - test
  - deploy

variables:
IMAGE_NAME: registry.example.com/ernie-image-server
IMAGE_TAG: $CI_COMMIT_SHORT_SHA

build:
stage: build
script:
- docker build -t $IMAGE_NAME:$IMAGE_TAG .
- docker push $IMAGE_NAME:$IMAGE_TAG
tags:
- docker-builder

test:
stage: test
script:
- docker run --gpus all $IMAGE_NAME:$IMAGE_TAG python3 -m pytest tests/
tags:
- gpu-runner

deploy-staging:
stage: deploy
script:
- kubectl set image deployment/ernie-image-server
ernie-image=$IMAGE_NAME:$IMAGE_TAG -n ernie-image-staging
environment:
name: staging
tags:
- k8s-admin

deploy-production:
stage: deploy
script:
- kubectl set image deployment/ernie-image-server
ernie-image=$IMAGE_NAME:$IMAGE_TAG -n ernie-image
environment:
name: production
when: manual
tags:
- k8s-admin

VII. Troubleshooting

7.1 Common Issues

Issue Cause Solution
GPU OOM Batch size too large Reduce MAX_BATCH_SIZE, use NVFP4 quantization
Inference timeout Queue backlog Check HPA config, increase replicas
Slow model loading Cold start Pre-load models to PVC, use readinessProbe
VRAM leak Uncleaned cache Restart pods periodically, torch.cuda.empty_cache()
Slow cross-GPU No NVLink Use NVLink or InfiniBand

7.2 Log Viewing

# View pod logs
kubectl logs -n ernie-image deployment/ernie-image-server --tail=100

Real-time logs

kubectl logs -f -n ernie-image -l app=ernie-image

Previous pod logs (crash debugging)

kubectl logs -n ernie-image deployment/ernie-image-server --previous

VIII. Cost Optimization

8.1 Quantization Savings

Deployment VRAM Speed Cost
BF16 Full Precision ~17GB Baseline 100%
FP8 Quantized ~9GB ~95% ~60%
NVFP4 Quantized ~5GB ~90% ~35%
INT8 Quantized ~9GB ~85% ~50%

8.2 GPU Spot Instances

Using Spot instances in Kubernetes saves 60-70% on GPU costs:

# k8s/node-pool-spot.yaml
apiVersion: karpenter.sh/v1alpha5
kind: NodePool
metadata:
  name: ernie-gpu-spot
spec:
  template:
    spec:
      requirements:
        - key: nvidia.com/gpu
          operator: In
          values: ["A100-80GB"]
      spot: true

IX. Summary

The Docker containerization + Kubernetes orchestration approach for ERNIE-Image provides an enterprise-grade production deployment solution:

  1. Docker images: Environment isolation, reproducibility, easy distribution
  2. Kubernetes: Auto-scaling, high availability, rolling updates
  3. Monitoring: Real-time metrics, intelligent alerts, failure prediction
  4. CI/CD: Automated building, testing, deployment
  5. Cost optimization: Quantization + Spot instances + auto-scaling

This approach works for everything from SMB single-node deployments to large enterprise multi-cluster production environments — the best practice for taking ERNIE-Image from experiment to production.


Further Reading:

ERNIE-Image Team