ERNIE-Image Docker 容器化部署与 Kubernetes 生产环境指南

Jun 11, 2026

ERNIE-Image Docker 容器化部署与 Kubernetes 生产环境指南

从单机实验到企业级生产:ERNIE-Image 容器化部署全攻略,含 Docker 镜像构建、Kubernetes 编排、自动扩缩容与监控告警。

引言

ERNIE-Image 以其出色的 8B 参数性能和广泛的硬件兼容性,已成为开源文生图领域的首选。但将模型从实验环境迁移到生产环境,面临着诸多挑战:环境一致性、资源隔离、弹性扩缩容、高可用性。

本文将提供完整的 Docker 容器化部署方案 和 Kubernetes 生产环境编排指南,涵盖从基础镜像构建到生产级监控告警的全链路。

一、Docker 镜像构建

1.1 基础镜像设计

# Dockerfile.ernie-image
FROM nvidia/cuda:12.4.1-devel-ubuntu22.04

设置环境变量

ENV DEBIAN_FRONTEND=noninteractive
PYTHON_VERSION=3.11
PYTHONDONTWRITEBYTECODE=1
PIP_NO_CACHE_DIR=1

安装系统依赖

RUN apt-get update && apt-get install -y --no-install-recommends
python3${PYTHON_VERSION}
python3${PYTHON_VERSION}-venv
git
wget
curl
&& rm -rf /var/lib/apt/lists/*

创建虚拟环境

RUN python3 -m venv /opt/venv
ENV PATH="/opt/venv/bin:$PATH"

安装 Python 依赖

RUN pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu121
RUN pip install diffusers transformers accelerate sentencepiece protobuf

安装 SGLang(高性能推理)

RUN pip install sglang

工作目录

WORKDIR /app

复制应用代码

COPY scripts/ ./scripts/
COPY config/ ./config/

创建模型缓存目录

RUN mkdir -p /models/ernie-image

健康检查

HEALTHCHECK --interval=30s --timeout=10s --start-period=60s --retries=3
CMD python3 /app/scripts/health_check.py || exit 1

暴露端口

EXPOSE 8000

启动命令

CMD ["python3", "/app/scripts/server.py"]

1.2 多阶段构建优化

# Dockerfile.multi-stage
# 阶段 1:构建依赖
FROM nvidia/cuda:12.4.1-devel-ubuntu22.04 AS builder
ENV DEBIAN_FRONTEND=noninteractive
RUN apt-get update && apt-get install -y python3 python3-venv git
RUN python3 -m venv /opt/venv
ENV PATH="/opt/venv/bin:$PATH"
RUN pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu121
RUN pip install diffusers transformers accelerate sentencepiece protobuf sglang

阶段 2:运行时

FROM nvidia/cuda:12.4.1-runtime-ubuntu22.04
ENV DEBIAN_FRONTEND=noninteractive
RUN apt-get update && apt-get install -y --no-install-recommends
python3 libgl1 && rm -rf /var/lib/apt/lists/*
COPY --from=builder /opt/venv /opt/venv
ENV PATH="/opt/venv/bin:$PATH"
WORKDIR /app
COPY scripts/ ./scripts/
COPY config/ ./config/
RUN mkdir -p /models/ernie-image
EXPOSE 8000
CMD ["python3", "/app/scripts/server.py"]

1.3 镜像标签策略

标签 用途
ernie-image-server:latest 最新稳定版
ernie-image-server:1.0.0 语义化版本
ernie-image-server:cu121-torch2.3 特定 CUDA/PyTorch 版本
ernie-image-server:lite 精简版(不含训练工具)

二、Docker Compose 单机部署

2.1 docker-compose.yml

# docker-compose.yml
version: '3.8'

services:
ernie-image:
build:
context: .
dockerfile: Dockerfile.ernie-image
image: ernie-image-server:latest
container_name: ernie-image-server
restart: unless-stopped
ports:
- "8000:8000"
volumes:
- ./models:/models/ernie-image
- ./outputs:/app/outputs
- ./config:/app/config:ro
environment:
- MODEL_PATH=/models/ernie-image
- OUTPUT_DIR=/app/outputs
- MAX_BATCH_SIZE=8
- CFG_SCALE=4.0
- NUM_INFERENCE_STEPS=50
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: 1
capabilities: [gpu]
limits:
memory: 24g
healthcheck:
test: ["CMD", "python3", "/app/scripts/health_check.py"]
interval: 30s
timeout: 10s
retries: 3
networks:
- ernie-network

nginx:
image: nginx:alpine
container_name: ernie-image-nginx
restart: unless-stopped
ports:
- "80:80"
- "443:443"
volumes:
- ./nginx/nginx.conf:/etc/nginx/nginx.conf:ro
- ./nginx/ssl:/etc/nginx/ssl:ro
depends_on:
- ernie-image
networks:
- ernie-network

networks:
ernie-network:
driver: bridge

2.2 启动与验证

# 拉取模型
huggingface-cli download baidu/ERNIE-Image --local-dir /path/to/models/ernie-image

启动服务

docker compose up -d

查看状态

docker compose ps

期望输出:ernie-image-server (healthy), ernie-image-nginx (healthy)

查看日志

docker compose logs -f ernie-image

测试接口

curl -X POST http://localhost:8000/generate
-H "Content-Type: application/json"
-d '{"prompt": "A cat sitting on a chair, photorealistic style", "steps": 50, "cfg_scale": 4.0}'

三、Kubernetes 生产环境部署

3.1 命名空间与资源配额

# k8s/namespace.yaml
apiVersion: v1
kind: Namespace
metadata:
  name: ernie-image
  labels:
    team: ai-infra
    environment: production
---
# ResourceQuota
apiVersion: v1
kind: ResourceQuota
metadata:
  name: ernie-image-quota
  namespace: ernie-image
spec:
  hard:
    requests.cpu: "32"
    requests.memory: 128Gi
    requests.nvidia.com/gpu: "8"
    limits.cpu: "64"
    limits.memory: 256Gi
    limits.nvidia.com/gpu: "8"

3.2 持久化存储

# k8s/pvc.yaml
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: ernie-models-pvc
  namespace: ernie-image
spec:
  accessModes:
    - ReadOnlyMany
  resources:
    requests:
      storage: 50Gi
  storageClassName: gp3
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: ernie-outputs-pvc
  namespace: ernie-image
spec:
  accessModes:
    - ReadWriteMany
  resources:
    requests:
      storage: 100Gi
  storageClassName: gp3

3.3 部署配置

# k8s/deployment.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
  name: ernie-image-server
  namespace: ernie-image
  labels:
    app: ernie-image
    version: v1.0.0
spec:
  replicas: 3
  strategy:
    type: RollingUpdate
    rollingUpdate:
      maxSurge: 1
      maxUnavailable: 0
  selector:
    matchLabels:
      app: ernie-image
  template:
    metadata:
      labels:
        app: ernie-image
        version: v1.0.0
      annotations:
        prometheus.io/scrape: "true"
        prometheus.io/port: "8000"
    spec:
      containers:
        - name: ernie-image
          image: ernie-image-server:1.0.0
          ports:
            - containerPort: 8000
              name: http
          env:
            - name: MODEL_PATH
              value: "/models/ernie-image"
            - name: OUTPUT_DIR
              value: "/outputs"
            - name: MAX_BATCH_SIZE
              value: "8"
            - name: NUM_WORKERS
              value: "4"
          resources:
            requests:
              cpu: "4"
              memory: 24Gi
              nvidia.com/gpu: "1"
            limits:
              cpu: "8"
              memory: 48Gi
              nvidia.com/gpu: "1"
          volumeMounts:
            - name: models
              mountPath: /models/ernie-image
              readOnly: true
            - name: outputs
              mountPath: /outputs
          livenessProbe:
            httpGet:
              path: /health
              port: 8000
            initialDelaySeconds: 60
            periodSeconds: 30
            timeoutSeconds: 10
            failureThreshold: 3
          readinessProbe:
            httpGet:
              path: /ready
              port: 8000
            initialDelaySeconds: 30
            periodSeconds: 15
            timeoutSeconds: 5
      volumes:
        - name: models
          persistentVolumeClaim:
            claimName: ernie-models-pvc
        - name: outputs
          persistentVolumeClaim:
            claimName: ernie-outputs-pvc
      affinity:
        podAntiAffinity:
          preferredDuringSchedulingIgnoredDuringExecution:
            - weight: 100
              podAffinityTerm:
                labelSelector:
                  matchExpressions:
                    - key: app
                      operator: In
                      values:
                        - ernie-image
                topologyKey: kubernetes.io/hostname
      tolerations:
        - key: nvidia.com/gpu
          operator: Exists
          effect: NoSchedule

3.4 服务与 Ingress

# k8s/service.yaml
apiVersion: v1
kind: Service
metadata:
  name: ernie-image-service
  namespace: ernie-image
  labels:
    app: ernie-image
spec:
  type: ClusterIP
  ports:
    - port: 80
      targetPort: 8000
      protocol: TCP
      name: http
  selector:
    app: ernie-image
---
# k8s/ingress.yaml
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
  name: ernie-image-ingress
  namespace: ernie-image
  annotations:
    nginx.ingress.kubernetes.io/ssl-redirect: "true"
    nginx.ingress.kubernetes.io/proxy-body-size: "50m"
    cert-manager.io/cluster-issuer: "letsencrypt-prod"
spec:
  ingressClassName: nginx
  tls:
    - hosts:
        - api.ernie-image.example.com
      secretName: ernie-image-tls
  rules:
    - host: api.ernie-image.example.com
      http:
        paths:
          - path: /
            pathType: Prefix
            backend:
              service:
                name: ernie-image-service
                port:
                  number: 80

3.5 自动扩缩容(HPA)

# k8s/hpa.yaml
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: ernie-image-hpa
  namespace: ernie-image
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: ernie-image-server
  minReplicas: 1
  maxReplicas: 10
  metrics:
    - type: Resource
      resource:
        name: cpu
        target:
          type: Utilization
          averageUtilization: 70
    - type: Resource
      resource:
        name: memory
        target:
          type: Utilization
          averageUtilization: 80
    - type: Pods
      pods:
        metric:
          name: requests-per-pod
        target:
          type: AverageValue
          averageValue: "100"

3.6 GPU 节点自动扩缩容(KEDA + Cluster Autoscaler)

对于 GPU 资源,KEDA 可以基于自定义指标触发扩缩容:

# k8s/keda-scaled-object.yaml
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
  name: ernie-image-keda
  namespace: ernie-image
spec:
  scaleTargetRef:
    name: ernie-image-server
  minReplicaCount: 1
  maxReplicaCount: 10
  triggers:
    - type: prometheus
      metricName: ernie_queue_length
      threshold: "20"
      serverAddress: http://prometheus:9090

四、监控与告警

4.1 Prometheus 指标

# metrics.py - 在应用中添加 Prometheus 指标
from prometheus_client import Counter, Histogram, Gauge

请求计数器

REQUEST_COUNT = Counter(
'ernie_image_requests_total',
'Total number of image generation requests',
['status', 'model_variant']
)

请求延迟直方图

REQUEST_LATENCY = Histogram(
'ernie_image_request_latency_seconds',
'Image generation request latency',
['model_variant']
)

队列长度

QUEUE_LENGTH = Gauge(
'ernie_image_queue_length',
'Current generation queue length'
)

GPU 使用率

GPU_UTILIZATION = Gauge(
'ernie_image_gpu_utilization_percent',
'GPU utilization percentage',
['gpu_id']
)

活跃请求数

ACTIVE_REQUESTS = Gauge(
'ernie_image_active_requests',
'Number of currently active generation requests'
)

4.2 Grafana 仪表板

推荐监控面板布局:

┌─────────────────┬─────────────────┬─────────────────┐
│  QPS (请求/秒)  │  P99 延迟 (ms)   │  GPU 使用率 (%)  │
├─────────────────┼─────────────────┼─────────────────┤
│  队列长度        │  活跃 Pod 数     │  错误率 (%)      │
├─────────────────┼─────────────────┼─────────────────┤
│  内存使用 (GB)   │  显存使用 (GB)   │  吞吐 (images/s) │
└─────────────────┴─────────────────┴─────────────────┘

4.3 告警规则

# k8s/alerts.yaml
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
  name: ernie-image-alerts
  namespace: monitoring
spec:
  groups:
    - name: ernie-image
      rules:
        - alert: HighQueueLength
          expr: ernie_image_queue_length > 50
          for: 2m
          labels:
            severity: warning
          annotations:
            summary: "ERNIE-Image queue length exceeds 50"
            description: "Current queue length: {{ $value }}"
    - alert: HighLatency
      expr: histogram_quantile(0.99, rate(ernie_image_request_latency_seconds_bucket[5m])) > 30
      for: 5m
      labels:
        severity: warning
      annotations:
        summary: "ERNIE-Image P99 latency exceeds 30 seconds"

    - alert: HighErrorRate
      expr: rate(ernie_image_requests_total{status="error"}[5m]) / rate(ernie_image_requests_total[5m]) > 0.05
      for: 5m
      labels:
        severity: critical
      annotations:
        summary: "ERNIE-Image error rate exceeds 5%"

    - alert: GPUOutOfMemory
      expr: ernie_image_gpu_utilization_percent > 95
      for: 1m
      labels:
        severity: critical
      annotations:
        summary: "GPU memory nearly full"

五、配置管理

5.1 Kubernetes ConfigMap

# k8s/configmap.yaml
apiVersion: v1
kind: ConfigMap
metadata:
  name: ernie-image-config
  namespace: ernie-image
data:
  model_variant: "ERNIE-Image-Turbo"
  max_batch_size: "8"
  num_inference_steps: "8"
  cfg_scale: "4.0"
  output_format: "jpeg"
  output_quality: "95"
  max_image_size: "2048x2048"
  enable_pe: "true"
  pe_model_path: "/models/ernie-image-pe"
  timeout_seconds: "120"
  max_queue_size: "100"

5.2 多环境配置

# environments:
#   development:
#     replicas: 1
#     gpu: 1 (RTX 3090)
#     max_batch_size: 2
#   staging:
#     replicas: 2
#     gpu: 2 (A100 40GB)
#     max_batch_size: 4
#   production:
#     replicas: 3-10 (HPA)
#     gpu: 3-10 (A100 80GB)
#     max_batch_size: 8

六、CI/CD 流水线

6.1 GitLab CI 示例

# .gitlab-ci.yml
stages:
  - build
  - test
  - deploy

variables:
IMAGE_NAME: registry.example.com/ernie-image-server
IMAGE_TAG: $CI_COMMIT_SHORT_SHA

build:
stage: build
script:
- docker build -t $IMAGE_NAME:$IMAGE_TAG .
- docker push $IMAGE_NAME:$IMAGE_TAG
tags:
- docker-builder

test:
stage: test
script:
- docker run --gpus all $IMAGE_NAME:$IMAGE_TAG python3 -m pytest tests/
tags:
- gpu-runner

deploy-staging:
stage: deploy
script:
- kubectl set image deployment/ernie-image-server
ernie-image=$IMAGE_NAME:$IMAGE_TAG
-n ernie-image-staging
environment:
name: staging
tags:
- k8s-admin

deploy-production:
stage: deploy
script:
- kubectl set image deployment/ernie-image-server
ernie-image=$IMAGE_NAME:$IMAGE_TAG
-n ernie-image
environment:
name: production
when: manual
tags:
- k8s-admin

七、故障排查

7.1 常见问题

问题 原因 解决方案
GPU OOM 批处理大小过大 减小 MAX_BATCH_SIZE,使用 NVFP4 量化
推理超时 队列积压 检查 HPA 配置,增加副本数
模型加载慢 冷启动 预加载模型到持久卷,使用 readinessProbe
显存泄漏 未清理缓存 定期重启 Pod,设置 torch.cuda.empty_cache()
跨 GPU 通信慢 非 NVLink 互联 使用 NVLink 或 InfiniBand

7.2 日志查看

# 查看 Pod 日志
kubectl logs -n ernie-image deployment/ernie-image-server --tail=100

查看实时日志

kubectl logs -f -n ernie-image -l app=ernie-image

查看历史 Pod 日志(崩溃排查)

kubectl logs -n ernie-image deployment/ernie-image-server --previous

八、成本优化

8.1 量化部署节省

部署方式 显存 速度 成本对比
BF16 全精度 ~17GB 基准 100%
FP8 量化 ~9GB ~95% 速度 ~60% 成本
NVFP4 量化 ~5GB ~90% 速度 ~35% 成本
INT8 量化 ~9GB ~85% 速度 ~50% 成本

8.2 GPU Spot 实例

在 Kubernetes 中使用 Spot 实例可以节省 60-70% 的 GPU 成本:

# k8s/node-pool-spot.yaml
apiVersion: karpenter.sh/v1alpha5
kind: NodePool
metadata:
  name: ernie-gpu-spot
spec:
  template:
    spec:
      requirements:
        - key: nvidia.com/gpu
          operator: In
          values: ["A100-80GB"]
      spot: true

九、总结

ERNIE-Image 的 Docker 容器化 + Kubernetes 编排方案提供了企业级生产部署的完整解决方案:

  1. Docker 镜像:环境隔离、可复现、易于分发
  2. Kubernetes:自动扩缩容、高可用、滚动更新
  3. 监控告警:实时指标、智能告警、故障预测
  4. CI/CD:自动化构建、测试、部署
  5. 成本优化:量化部署 + Spot 实例 + 自动扩缩容

这套方案既适用于中小企业的单机部署,也适用于大型企业的多集群生产环境,是 ERNIE-Image 从实验走向生产的最佳实践。


扩展阅读:

ERNIE-Image Team