ERNIE-Image Docker 容器化部署与 Kubernetes 生产环境指南
从单机实验到企业级生产:ERNIE-Image 容器化部署全攻略,含 Docker 镜像构建、Kubernetes 编排、自动扩缩容与监控告警。
引言
ERNIE-Image 以其出色的 8B 参数性能和广泛的硬件兼容性,已成为开源文生图领域的首选。但将模型从实验环境迁移到生产环境,面临着诸多挑战:环境一致性、资源隔离、弹性扩缩容、高可用性。
本文将提供完整的 Docker 容器化部署方案 和 Kubernetes 生产环境编排指南,涵盖从基础镜像构建到生产级监控告警的全链路。
一、Docker 镜像构建
1.1 基础镜像设计
# Dockerfile.ernie-image
FROM nvidia/cuda:12.4.1-devel-ubuntu22.04
设置环境变量
ENV DEBIAN_FRONTEND=noninteractive
PYTHON_VERSION=3.11
PYTHONDONTWRITEBYTECODE=1
PIP_NO_CACHE_DIR=1
安装系统依赖
RUN apt-get update && apt-get install -y --no-install-recommends
python3${PYTHON_VERSION}
python3${PYTHON_VERSION}-venv
git
wget
curl
&& rm -rf /var/lib/apt/lists/*
创建虚拟环境
RUN python3 -m venv /opt/venv
ENV PATH="/opt/venv/bin:$PATH"
安装 Python 依赖
RUN pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu121
RUN pip install diffusers transformers accelerate sentencepiece protobuf
安装 SGLang(高性能推理)
RUN pip install sglang
工作目录
WORKDIR /app
复制应用代码
COPY scripts/ ./scripts/
COPY config/ ./config/
创建模型缓存目录
RUN mkdir -p /models/ernie-image
健康检查
HEALTHCHECK --interval=30s --timeout=10s --start-period=60s --retries=3
CMD python3 /app/scripts/health_check.py || exit 1
暴露端口
EXPOSE 8000
启动命令
CMD ["python3", "/app/scripts/server.py"]
1.2 多阶段构建优化
# Dockerfile.multi-stage
# 阶段 1:构建依赖
FROM nvidia/cuda:12.4.1-devel-ubuntu22.04 AS builder
ENV DEBIAN_FRONTEND=noninteractive
RUN apt-get update && apt-get install -y python3 python3-venv git
RUN python3 -m venv /opt/venv
ENV PATH="/opt/venv/bin:$PATH"
RUN pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu121
RUN pip install diffusers transformers accelerate sentencepiece protobuf sglang
阶段 2:运行时
FROM nvidia/cuda:12.4.1-runtime-ubuntu22.04
ENV DEBIAN_FRONTEND=noninteractive
RUN apt-get update && apt-get install -y --no-install-recommends
python3 libgl1 && rm -rf /var/lib/apt/lists/*
COPY --from=builder /opt/venv /opt/venv
ENV PATH="/opt/venv/bin:$PATH"
WORKDIR /app
COPY scripts/ ./scripts/
COPY config/ ./config/
RUN mkdir -p /models/ernie-image
EXPOSE 8000
CMD ["python3", "/app/scripts/server.py"]
1.3 镜像标签策略
| 标签 | 用途 |
|---|---|
ernie-image-server:latest |
最新稳定版 |
ernie-image-server:1.0.0 |
语义化版本 |
ernie-image-server:cu121-torch2.3 |
特定 CUDA/PyTorch 版本 |
ernie-image-server:lite |
精简版(不含训练工具) |
二、Docker Compose 单机部署
2.1 docker-compose.yml
# docker-compose.yml
version: '3.8'
services:
ernie-image:
build:
context: .
dockerfile: Dockerfile.ernie-image
image: ernie-image-server:latest
container_name: ernie-image-server
restart: unless-stopped
ports:
- "8000:8000"
volumes:
- ./models:/models/ernie-image
- ./outputs:/app/outputs
- ./config:/app/config:ro
environment:
- MODEL_PATH=/models/ernie-image
- OUTPUT_DIR=/app/outputs
- MAX_BATCH_SIZE=8
- CFG_SCALE=4.0
- NUM_INFERENCE_STEPS=50
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: 1
capabilities: [gpu]
limits:
memory: 24g
healthcheck:
test: ["CMD", "python3", "/app/scripts/health_check.py"]
interval: 30s
timeout: 10s
retries: 3
networks:
- ernie-network
nginx:
image: nginx:alpine
container_name: ernie-image-nginx
restart: unless-stopped
ports:
- "80:80"
- "443:443"
volumes:
- ./nginx/nginx.conf:/etc/nginx/nginx.conf:ro
- ./nginx/ssl:/etc/nginx/ssl:ro
depends_on:
- ernie-image
networks:
- ernie-network
networks:
ernie-network:
driver: bridge
2.2 启动与验证
# 拉取模型
huggingface-cli download baidu/ERNIE-Image --local-dir /path/to/models/ernie-image
启动服务
docker compose up -d
查看状态
docker compose ps
期望输出:ernie-image-server (healthy), ernie-image-nginx (healthy)
查看日志
docker compose logs -f ernie-image
测试接口
curl -X POST http://localhost:8000/generate
-H "Content-Type: application/json"
-d '{"prompt": "A cat sitting on a chair, photorealistic style", "steps": 50, "cfg_scale": 4.0}'
三、Kubernetes 生产环境部署
3.1 命名空间与资源配额
# k8s/namespace.yaml
apiVersion: v1
kind: Namespace
metadata:
name: ernie-image
labels:
team: ai-infra
environment: production
---
# ResourceQuota
apiVersion: v1
kind: ResourceQuota
metadata:
name: ernie-image-quota
namespace: ernie-image
spec:
hard:
requests.cpu: "32"
requests.memory: 128Gi
requests.nvidia.com/gpu: "8"
limits.cpu: "64"
limits.memory: 256Gi
limits.nvidia.com/gpu: "8"
3.2 持久化存储
# k8s/pvc.yaml
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: ernie-models-pvc
namespace: ernie-image
spec:
accessModes:
- ReadOnlyMany
resources:
requests:
storage: 50Gi
storageClassName: gp3
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: ernie-outputs-pvc
namespace: ernie-image
spec:
accessModes:
- ReadWriteMany
resources:
requests:
storage: 100Gi
storageClassName: gp3
3.3 部署配置
# k8s/deployment.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
name: ernie-image-server
namespace: ernie-image
labels:
app: ernie-image
version: v1.0.0
spec:
replicas: 3
strategy:
type: RollingUpdate
rollingUpdate:
maxSurge: 1
maxUnavailable: 0
selector:
matchLabels:
app: ernie-image
template:
metadata:
labels:
app: ernie-image
version: v1.0.0
annotations:
prometheus.io/scrape: "true"
prometheus.io/port: "8000"
spec:
containers:
- name: ernie-image
image: ernie-image-server:1.0.0
ports:
- containerPort: 8000
name: http
env:
- name: MODEL_PATH
value: "/models/ernie-image"
- name: OUTPUT_DIR
value: "/outputs"
- name: MAX_BATCH_SIZE
value: "8"
- name: NUM_WORKERS
value: "4"
resources:
requests:
cpu: "4"
memory: 24Gi
nvidia.com/gpu: "1"
limits:
cpu: "8"
memory: 48Gi
nvidia.com/gpu: "1"
volumeMounts:
- name: models
mountPath: /models/ernie-image
readOnly: true
- name: outputs
mountPath: /outputs
livenessProbe:
httpGet:
path: /health
port: 8000
initialDelaySeconds: 60
periodSeconds: 30
timeoutSeconds: 10
failureThreshold: 3
readinessProbe:
httpGet:
path: /ready
port: 8000
initialDelaySeconds: 30
periodSeconds: 15
timeoutSeconds: 5
volumes:
- name: models
persistentVolumeClaim:
claimName: ernie-models-pvc
- name: outputs
persistentVolumeClaim:
claimName: ernie-outputs-pvc
affinity:
podAntiAffinity:
preferredDuringSchedulingIgnoredDuringExecution:
- weight: 100
podAffinityTerm:
labelSelector:
matchExpressions:
- key: app
operator: In
values:
- ernie-image
topologyKey: kubernetes.io/hostname
tolerations:
- key: nvidia.com/gpu
operator: Exists
effect: NoSchedule
3.4 服务与 Ingress
# k8s/service.yaml
apiVersion: v1
kind: Service
metadata:
name: ernie-image-service
namespace: ernie-image
labels:
app: ernie-image
spec:
type: ClusterIP
ports:
- port: 80
targetPort: 8000
protocol: TCP
name: http
selector:
app: ernie-image
---
# k8s/ingress.yaml
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
name: ernie-image-ingress
namespace: ernie-image
annotations:
nginx.ingress.kubernetes.io/ssl-redirect: "true"
nginx.ingress.kubernetes.io/proxy-body-size: "50m"
cert-manager.io/cluster-issuer: "letsencrypt-prod"
spec:
ingressClassName: nginx
tls:
- hosts:
- api.ernie-image.example.com
secretName: ernie-image-tls
rules:
- host: api.ernie-image.example.com
http:
paths:
- path: /
pathType: Prefix
backend:
service:
name: ernie-image-service
port:
number: 80
3.5 自动扩缩容(HPA)
# k8s/hpa.yaml
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: ernie-image-hpa
namespace: ernie-image
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: ernie-image-server
minReplicas: 1
maxReplicas: 10
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 70
- type: Resource
resource:
name: memory
target:
type: Utilization
averageUtilization: 80
- type: Pods
pods:
metric:
name: requests-per-pod
target:
type: AverageValue
averageValue: "100"
3.6 GPU 节点自动扩缩容(KEDA + Cluster Autoscaler)
对于 GPU 资源,KEDA 可以基于自定义指标触发扩缩容:
# k8s/keda-scaled-object.yaml
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
name: ernie-image-keda
namespace: ernie-image
spec:
scaleTargetRef:
name: ernie-image-server
minReplicaCount: 1
maxReplicaCount: 10
triggers:
- type: prometheus
metricName: ernie_queue_length
threshold: "20"
serverAddress: http://prometheus:9090
四、监控与告警
4.1 Prometheus 指标
# metrics.py - 在应用中添加 Prometheus 指标
from prometheus_client import Counter, Histogram, Gauge
请求计数器
REQUEST_COUNT = Counter(
'ernie_image_requests_total',
'Total number of image generation requests',
['status', 'model_variant']
)
请求延迟直方图
REQUEST_LATENCY = Histogram(
'ernie_image_request_latency_seconds',
'Image generation request latency',
['model_variant']
)
队列长度
QUEUE_LENGTH = Gauge(
'ernie_image_queue_length',
'Current generation queue length'
)
GPU 使用率
GPU_UTILIZATION = Gauge(
'ernie_image_gpu_utilization_percent',
'GPU utilization percentage',
['gpu_id']
)
活跃请求数
ACTIVE_REQUESTS = Gauge(
'ernie_image_active_requests',
'Number of currently active generation requests'
)
4.2 Grafana 仪表板
推荐监控面板布局:
┌─────────────────┬─────────────────┬─────────────────┐
│ QPS (请求/秒) │ P99 延迟 (ms) │ GPU 使用率 (%) │
├─────────────────┼─────────────────┼─────────────────┤
│ 队列长度 │ 活跃 Pod 数 │ 错误率 (%) │
├─────────────────┼─────────────────┼─────────────────┤
│ 内存使用 (GB) │ 显存使用 (GB) │ 吞吐 (images/s) │
└─────────────────┴─────────────────┴─────────────────┘
4.3 告警规则
# k8s/alerts.yaml
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: ernie-image-alerts
namespace: monitoring
spec:
groups:
- name: ernie-image
rules:
- alert: HighQueueLength
expr: ernie_image_queue_length > 50
for: 2m
labels:
severity: warning
annotations:
summary: "ERNIE-Image queue length exceeds 50"
description: "Current queue length: {{ $value }}"
- alert: HighLatency
expr: histogram_quantile(0.99, rate(ernie_image_request_latency_seconds_bucket[5m])) > 30
for: 5m
labels:
severity: warning
annotations:
summary: "ERNIE-Image P99 latency exceeds 30 seconds"
- alert: HighErrorRate
expr: rate(ernie_image_requests_total{status="error"}[5m]) / rate(ernie_image_requests_total[5m]) > 0.05
for: 5m
labels:
severity: critical
annotations:
summary: "ERNIE-Image error rate exceeds 5%"
- alert: GPUOutOfMemory
expr: ernie_image_gpu_utilization_percent > 95
for: 1m
labels:
severity: critical
annotations:
summary: "GPU memory nearly full"
五、配置管理
5.1 Kubernetes ConfigMap
# k8s/configmap.yaml
apiVersion: v1
kind: ConfigMap
metadata:
name: ernie-image-config
namespace: ernie-image
data:
model_variant: "ERNIE-Image-Turbo"
max_batch_size: "8"
num_inference_steps: "8"
cfg_scale: "4.0"
output_format: "jpeg"
output_quality: "95"
max_image_size: "2048x2048"
enable_pe: "true"
pe_model_path: "/models/ernie-image-pe"
timeout_seconds: "120"
max_queue_size: "100"
5.2 多环境配置
# environments:
# development:
# replicas: 1
# gpu: 1 (RTX 3090)
# max_batch_size: 2
# staging:
# replicas: 2
# gpu: 2 (A100 40GB)
# max_batch_size: 4
# production:
# replicas: 3-10 (HPA)
# gpu: 3-10 (A100 80GB)
# max_batch_size: 8
六、CI/CD 流水线
6.1 GitLab CI 示例
# .gitlab-ci.yml
stages:
- build
- test
- deploy
variables:
IMAGE_NAME: registry.example.com/ernie-image-server
IMAGE_TAG: $CI_COMMIT_SHORT_SHA
build:
stage: build
script:
- docker build -t $IMAGE_NAME:$IMAGE_TAG .
- docker push $IMAGE_NAME:$IMAGE_TAG
tags:
- docker-builder
test:
stage: test
script:
- docker run --gpus all $IMAGE_NAME:$IMAGE_TAG python3 -m pytest tests/
tags:
- gpu-runner
deploy-staging:
stage: deploy
script:
- kubectl set image deployment/ernie-image-server
ernie-image=$IMAGE_NAME:$IMAGE_TAG
-n ernie-image-staging
environment:
name: staging
tags:
- k8s-admin
deploy-production:
stage: deploy
script:
- kubectl set image deployment/ernie-image-server
ernie-image=$IMAGE_NAME:$IMAGE_TAG
-n ernie-image
environment:
name: production
when: manual
tags:
- k8s-admin
七、故障排查
7.1 常见问题
| 问题 | 原因 | 解决方案 |
|---|---|---|
| GPU OOM | 批处理大小过大 | 减小 MAX_BATCH_SIZE,使用 NVFP4 量化 |
| 推理超时 | 队列积压 | 检查 HPA 配置,增加副本数 |
| 模型加载慢 | 冷启动 | 预加载模型到持久卷,使用 readinessProbe |
| 显存泄漏 | 未清理缓存 | 定期重启 Pod,设置 torch.cuda.empty_cache() |
| 跨 GPU 通信慢 | 非 NVLink 互联 | 使用 NVLink 或 InfiniBand |
7.2 日志查看
# 查看 Pod 日志
kubectl logs -n ernie-image deployment/ernie-image-server --tail=100
查看实时日志
kubectl logs -f -n ernie-image -l app=ernie-image
查看历史 Pod 日志(崩溃排查)
kubectl logs -n ernie-image deployment/ernie-image-server --previous
八、成本优化
8.1 量化部署节省
| 部署方式 | 显存 | 速度 | 成本对比 |
|---|---|---|---|
| BF16 全精度 | ~17GB | 基准 | 100% |
| FP8 量化 | ~9GB | ~95% 速度 | ~60% 成本 |
| NVFP4 量化 | ~5GB | ~90% 速度 | ~35% 成本 |
| INT8 量化 | ~9GB | ~85% 速度 | ~50% 成本 |
8.2 GPU Spot 实例
在 Kubernetes 中使用 Spot 实例可以节省 60-70% 的 GPU 成本:
# k8s/node-pool-spot.yaml
apiVersion: karpenter.sh/v1alpha5
kind: NodePool
metadata:
name: ernie-gpu-spot
spec:
template:
spec:
requirements:
- key: nvidia.com/gpu
operator: In
values: ["A100-80GB"]
spot: true
九、总结
ERNIE-Image 的 Docker 容器化 + Kubernetes 编排方案提供了企业级生产部署的完整解决方案:
- Docker 镜像:环境隔离、可复现、易于分发
- Kubernetes:自动扩缩容、高可用、滚动更新
- 监控告警:实时指标、智能告警、故障预测
- CI/CD:自动化构建、测试、部署
- 成本优化:量化部署 + Spot 实例 + 自动扩缩容
这套方案既适用于中小企业的单机部署,也适用于大型企业的多集群生产环境,是 ERNIE-Image 从实验走向生产的最佳实践。
扩展阅读: