ERNIE-Image on AMD GPUs: Official Day-0 Support for Instinct MI355X and Radeon AI PRO R9700

Jul 20, 2026

ERNIE-Image on AMD GPUs: Official Day-0 Support for Instinct MI355X and Radeon AI PRO R9700

Date: 2026-07-20
Target: ernie-image.app
Category: AI Technology


For as long as text-to-image models have been popular, there's been an unwritten rule: you need an NVIDIA GPU. CUDA, cuDNN, TensorRT — these technologies have made NVIDIA the de facto standard for AI image generation. Without one, you're either running slowly or not running at all.

That changed on April 23, 2026, when AMD published an official technical article announcing Day-0 support for Baidu's ERNIE-Image on two AMD GPUs: the Instinct MI355X (data center) and Radeon AI PRO R9700 (professional workstation). The headline finding: zero modifications to core inference code — ERNIE-Image runs on AMD hardware using the standard HuggingFace Diffusers + ROCm stack, with no manual CUDA-to-HIP porting required.

This matters especially for the Chinese market, where NVIDIA GPUs are supply-constrained and expensive, while AMD GPUs have a large existing install base in both servers (Instinct series) and workstations (Radeon series). ERNIE-Image, Baidu's open-source 8B-parameter text-to-image model, is the first Chinese open-source image generation model to receive official AMD Day-0 support.

This article provides a comprehensive deployment guide based on AMD's official technical article, covering memory optimization strategies, CUDA-to-ROCm migration, and practical recommendations for both GPUs.

Model Architecture and Hardware

ERNIE-Image Component Breakdown

ERNIE-Image is built on a single-stream Diffusion Transformer (DiT) architecture with four core components:

Component Type Size Role
Transformer ErnieImageTransformer2DModel 15 GB Diffusion backbone
Text Encoder Mistral3Model 7.2 GB Text encoding
Prompt Enhancer (PE) Ministral3ForCausalLM 7.2 GB Automatic prompt enrichment
VAE AutoencoderKLFlux2 161 MB Variational autoencoder
Scheduler FlowMatchEulerDiscreteScheduler — Sampling scheduler
Total ~29.5 GB

The Prompt Enhancer automatically expands short inputs into detailed Chinese descriptions, a key differentiator from other open-source image generation models.

AMD Instinct MI355X

The MI355X is AMD's latest data center AI accelerator, built on the CDNA 4 architecture (gfx950). It features 288 GB of HBM3e memory with 8 TB/s bandwidth, directly competing with NVIDIA's Blackwell generation in low-precision (FP4/FP6/FP8) compute throughput.

For ERNIE-Image, 288 GB of VRAM means the entire 29.5 GB model fits with enormous headroom — over 250 GB remains available for inference intermediate tensors after loading all components.

AMD Radeon AI PRO R9700

The R9700 is AMD's professional workstation AI accelerator, built on the RDNA 4 architecture (gfx1201) with 32 GB of GDDR6 memory. Key specifications:

  • 64 compute units, 4,096 stream processors
  • 128 AI accelerators
  • 640 GB/s peak bandwidth
  • 64 MB Infinity Cache
  • PCIe 5.0 x16 interface
  • 300W TBP
  • ECC memory support on Linux

The challenge for the R9700: 32 GB of VRAM cannot hold the full 29.5 GB model plus inference intermediate tensors. AMD's testing showed only about 1.17 GiB of headroom after loading all components — insufficient for inference. The solution is CPU offloading.

Environment Setup

Software Stack

AMD's official testing used the following software versions:

Component Version
Docker rocm/pytorch:rocm7.2.1_ubuntu24.04_py3.12_pytorch_release_2.9.1
PyTorch 2.9.1+rocm7.2.1
ROCm (HIP) 7.2.53211
Diffusers 0.38.0.dev0
Transformers 5.5.3
Accelerate 1.13.0

The Diffusers installation uses the add-ernie-image branch from the HsiaWinter/diffusers repository, as ERNIE-Image's official Diffusers support was merged via PR #13432.

Deployment Steps

Step 1: Pull the Docker Image

docker pull rocm/pytorch:rocm7.2.1_ubuntu24.04_py3.12_pytorch_release_2.9.1

Step 2: Create Container with GPU Passthrough

docker run -d --name ernie-image-test \
  --device=/dev/kfd --device=/dev/dri \
  --group-add video --group-add render \
  --shm-size=64G \
  -v /path/to/model:/workspace \
  rocm/pytorch:rocm7.2.1_ubuntu24.04_py3.12_pytorch_release_2.9.1 \
  sleep infinity

Key parameters:

  • --device=/dev/kfd --device=/dev/dri: Pass AMD GPU devices
  • --group-add video --group-add render: Grant GPU permissions
  • --shm-size=64G: Avoid data loading bottlenecks

Step 3: Verify GPU Availability

ROCm provides CUDA API compatibility through HIP, so standard torch.cuda interfaces work seamlessly:

import torch
print(f'CUDA available: {torch.cuda.is_available()}')
print(f'Device: {torch.cuda.get_device_name(0)}')
print(f'VRAM: {torch.cuda.get_device_properties(0).total_memory / 1024**3:.0f} GB')

Expected output:

CUDA available: True
Device: AMD Instinct MI355X
VRAM: 288 GB

Step 4: Install Diffusers and Dependencies

git clone https://github.com/HsiaWinter/diffusers /workspace/diffusers-ernie
cd /workspace/diffusers-ernie
git checkout add-ernie-image
pip install -e .
pip install accelerate Pillow transformers

Step 5: Extract Model Weights

Organize the ~29.5 GB of model files in the following structure:

ERNIE-Image/
├── model_index.json
├── transformer/      # 15 GB
├── text_encoder/     # 7.2 GB
├── pe/               # 7.2 GB
├── pe_tokenizer/
├── tokenizer/
├── scheduler/
└── vae/              # 161 MB

Inference on MI355X: Full-GPU Loading

On the MI355X, 288 GB of VRAM comfortably holds the entire model. The inference script is straightforward:

from diffusers import ErnieImagePipeline
import torch

pipe = ErnieImagePipeline.from_pretrained(
"/workspace/ERNIE-Image",
torch_dtype=torch.bfloat16
)
pipe = pipe.to("cuda")
pipe.transformer.eval()
pipe.vae.eval()
pipe.text_encoder.eval()
pipe.pe.eval()

generator = torch.Generator(device="cuda").manual_seed(42)
output = pipe(
prompt="A black and white Chinese rural dog running on grass",
height=1024,
width=1024,
num_inference_steps=50,
guidance_scale=5.0
)
output.images[0].save("ernie_output.png")

AMD's testing confirmed that all three prompts (English and Chinese) successfully generated 1024×1024 images. The Prompt Enhancer automatically expanded English inputs into detailed Chinese prompts, improving image quality further.

CUDA migration note: The only changes needed were removing CUBLAS_WORKSPACE_CONFIG and torch.backends.cudnn.deterministic settings — the core inference code remained completely unchanged.

Inference on R9700: CPU Offloading Strategy

The R9700's 32 GB VRAM presents a more realistic scenario — many content creators and professional designers use workstations with this class of GPU.

VRAM Analysis

Component-by-component memory analysis:

Component VRAM Usage Cumulative
Text Encoder 7.18 GiB 7.18 GiB
Prompt Enhancer +6.38 GiB 13.56 GiB
VAE +0.17 GiB 13.73 GiB
Transformer +14.96 GiB 28.69 GiB

After loading all components, only ~1.17 GiB of VRAM remains — far short of what's needed for 1024×1024 inference.

enable_model_cpu_offload

The solution uses HuggingFace Accelerate's enable_model_cpu_offload() method, which loads model components to the GPU on-demand during inference and immediately offloads them back to CPU after use:

from diffusers import ErnieImagePipeline
import torch

pipe = ErnieImagePipeline.from_pretrained(
"/workspace/ERNIE-Image",
torch_dtype=torch.bfloat16
)
pipe.enable_model_cpu_offload() # Key: enable CPU offloading
pipe.transformer.eval()
pipe.vae.eval()
pipe.text_encoder.eval()
pipe.pe.eval()

Inference code is identical to MI355X

output = pipe(
prompt="A black and white Chinese rural dog running on grass",
height=1024,
width=1024,
num_inference_steps=50,
guidance_scale=5.0
)
output.images[0].save("ernie_output_r9700.png")

Key change: Replace pipe.to("cuda") with pipe.enable_model_cpu_offload().

With CPU offloading enabled, peak VRAM drops to approximately 15 GB — well within the R9700's 32 GB limit, leaving about 17 GB of headroom.

The Cost of CPU Offloading

CPU offloading solves the VRAM limitation but adds latency — each inference requires transferring model components between CPU and GPU over PCIe 5.0 x16 (~64 GB/s bidirectional). For single image generation, this overhead is a few seconds, which is acceptable for non-batch scenarios.

CUDA to ROCm Migration Reference

For users migrating from NVIDIA CUDA environments, AMD provides this migration reference:

CUDA-Specific Item ROCm Handling
CUBLAS_WORKSPACE_CONFIG Not applicable — remove
torch.backends.cudnn.* Uses MIOpen — remove related settings
torch.use_deterministic_algorithms Partial support — remove if needed
torch.cuda.* API No changes needed (HIP compatibility layer)
Flash Attention / cuDNN AOTriton (auto-selected)

The bottom line: if your inference script only uses standard torch.cuda APIs, migrating from CUDA to ROCm requires almost no code changes. ERNIE-Image's Diffusers implementation meets this standard perfectly, enabling zero-modification porting.

Market Significance of Day-0 Support

ERNIE-Image's Unique Position

ERNIE-Image is Baidu's first open-source text-to-image DiT model, with 77.5K+ downloads and 660+ likes on HuggingFace. It is also the first Chinese open-source image generation model to receive official AMD Day-0 support.

Day-0 support means AMD validated and adapted ERNIE-Image on its GPUs immediately upon the model's open-source release. This reflects AMD's commitment to the Chinese open-source AI ecosystem.

The China Market Context

The Chinese AI market faces a unique hardware landscape: NVIDIA H100/B200 and other high-end GPUs are subject to export restrictions, making them difficult to acquire and expensive. AMD Instinct MI355X faces no similar restrictions, but CUDA ecosystem lock-in has historically hindered migration.

ERNIE-Image's AMD Day-0 support breaks this deadlock. An 8B-parameter model achieving state-of-the-art results among open-weight models, running on AMD GPUs with zero code modifications, means:

  • Data centers: Existing AMD Instinct infrastructure can directly deploy ERNIE-Image
  • Workstations: Radeon AI PRO users don't need to purchase additional NVIDIA hardware
  • Developers: ROCm ecosystem maturity is validated — full-stack compatibility from PyTorch through Diffusers

Performance Considerations

While AMD's article does not provide detailed end-to-end inference speed comparisons (e.g., seconds per image), based on published specifications:

  • The MI355X's 8 TB/s bandwidth exceeds the H100's 3.35 TB/s, offering advantages for large model inference
  • The R9700's 640 GB/s bandwidth is ~63% of an RTX 4090 (1,008 GB/s), but its 32 GB VRAM is 33% larger than the 4090's 24 GB
  • CPU offloading makes the full model accessible on 32 GB GPUs

Deployment Recommendations

GPU Selection Guide

  • MI355X (288 GB): Ideal for production deployment, batch inference, and high-concurrency services. Full GPU loading delivers maximum speed
  • R9700 (32 GB): Suitable for development, testing, and personal workstations. Requires CPU offloading; single-image latency is acceptable
  • Other AMD GPUs: For GPUs with less than 24 GB VRAM, consider GGUF or NVFP4 quantization (see EI-028/EI-015)

Troubleshooting

  1. torch.cuda.is_available() returns False: Verify ROCm installation with rocminfo to confirm GPU visibility
  2. Out of memory: Enable enable_model_cpu_offload() or upgrade to a higher-VRAM GPU
  3. Slow inference: Check if CPU offloading is active (adds PCIe transfer latency); use MI355X for batch workloads
  4. Precision issues: AOTriton automatically selects the Attention backend under ROCm; manual configuration is rarely needed

Looking Ahead

ERNIE-Image's AMD Day-0 support marks an important milestone for the Chinese open-source AI image generation ecosystem. As ROCm 7.x matures and AMD GPUs gain further market penetration in China, we can expect:

  • More Chinese open-source AI models to receive AMD Day-0 support
  • Continued ROCm optimization to improve inference performance on AMD hardware
  • The upcoming ERNIE-Image editing model (expected July 2026 Beta) to likely continue Day-0 support

For content creators, developers, and enterprise users, the ERNIE-Image + AMD GPU combination offers a high-quality text-to-image solution that doesn't depend on the NVIDIA ecosystem — a true embodiment of the open-source spirit.


Reference: AMD Developer Resources — "Day-0 Support for Baidu ERNIE-Image on AMD GPUs" (2026-04-23)

ERNIE-Image Team