ERNIE-Image on Apple Silicon: Complete MLX Deployment Guide for Local 8B Text-to-Image

Jun 8, 2026

ERNIE-Image on Apple Silicon: Complete MLX Deployment Guide for Local 8B Text-to-Image

Summary: Apple Silicon users can now deploy the ERNIE-Image text-to-image model locally on their Mac. This article provides an in-depth guide to native MLX deployment, how to choose between mflux and MLX-Gen, the impact of quantization on memory and speed, and a complete workflow from installation to batch generation. Whether you're on a MacBook Pro or Mac Studio, you can enjoy efficient open-source text-to-image inference right on your machine.


Why Apple Silicon Users Should Care About ERNIE-Image

A significant trend in AI image generation in 2026 is increasing local deployment. ERNIE-Image, Baidu's open-source 8B-parameter DiT (Diffusion Transformer) model, excels at text rendering, structured layout, and instruction following. Apple Silicon's Unified Memory Architecture provides an ideal hardware platform for running such mid-scale models — M1/M2/M3/M4 chips leverage the MLX framework for GPU-accelerated inference without any CUDA configuration.

Core advantage comparison:

Dimension NVIDIA GPU Apple Silicon (MLX)
Memory Type Dedicated VRAM Unified Memory (shared)
Minimum Recommended 12GB VRAM (Turbo) 16GB Unified Memory
Deployment Complexity Requires CUDA/PyTorch setup Native Metal acceleration
Cost 24GB GPU ≈ $800+ M3 MacBook ≈ $1600 (full computer)
Ecosystem Mature CUDA ecosystem Rapidly growing MLX ecosystem
Best Use Case Batch production, cloud deployment Local development, creative workflows

MLX Framework: Apple Silicon's PyTorch

MLX was developed by Apple's machine learning research team as an array framework designed specifically for Apple Silicon. The analogy to PyTorch:

  • PyTorch → CUDA: GPU acceleration via CUDA backend
  • MLX → Metal: GPU acceleration via Metal backend

Core MLX features:

  1. Shared Memory: Arrays reside in shared memory — no data copies across devices
  2. Lazy Evaluation: Arrays materialize only when needed, reducing intermediate computation
  3. Composable Function Transformations: Supports grad (automatic differentiation), vmap (vectorization), and more

For generative models like ERNIE-Image, MLX's advantage is clear: inference directly uses GPU cores without the CPU ↔ VRAM data transfer overhead that PyTorch requires.


Tool Selection: mflux vs MLX-Gen

Two main MLX runtimes currently support ERNIE-Image:

mflux (Original Project)

  • Author: Filip Strand
  • GitHub: https://github.com/filipstrand/mflux
  • Positioning: Native MLX implementations of state-of-the-art generative models
  • ERNIE-Image Status: Basic support (Baidu ERNIE team has established contact via Issue #412)
pip install mflux

MLX-Gen (Independent Fork)

  • Author: lpalbou
  • GitHub: https://github.com/lpalbou/mlx-gen
  • PyPI: mlx-gen
  • Positioning: Forked from mflux for faster Apple Silicon workflow iteration
  • ERNIE-Image Status: ✅ ERNIE-Image Turbo + q8/q4 mixed quantization
# Recommended: install with uv
uv pip install mlx-gen

Recommendation: If you need ERNIE-Image Turbo quantization support and a mature CLI toolchain, MLX-Gen is the better choice. mflux is ideal for users tracking upstream model updates.


Complete MLX-Gen Deployment Workflow

Step 1: Environment Setup

# Install uv (Python package manager)
curl -LsSf https://astral.sh/uv/install.sh | sh

Install MLX-Gen

uv pip install mlx-gen

Step 2: Download the Model

# Download ERNIE-Image Turbo model
mlxgen download --model ernie-image-turbo --family ernie-image

Speed Tip: Add HF_HUB_ENABLE_HF_TRANSFER=1 for significantly faster downloads.

HF_HUB_ENABLE_HF_TRANSFER=1 mlxgen download --model ernie-image-turbo --family ernie-image

Step 3: Prepare Quantized Model

# Create local quantized model folder (with Hugging Face model card)
mlxgen prepare --model ernie-image-turbo --quantize 8

Quantization level selection:

Level Memory Usage Generation Speed Quality Impact Recommended For
BF16 (none) ~12GB Baseline Best quality 32GB+ unified memory
q8 ~6GB Fast Minimal difference 16GB-24GB unified memory
q4 mixed ~3-4GB Fastest Acceptable Under 16GB unified memory

Step 4: Generate Images

mlxgen generate \
  --model ernie-image-turbo \
  --prompt "A vintage coffee shop interior, warm lighting, detailed architectural illustration, Chinese text '咖啡时光' on the wall" \
  --width 1024 \
  --height 1024 \
  --seed 42 \
  --steps 8 \
  --quantize 8

Output: Generated images are saved to the output/ directory by default.


Advanced Workflows

1. Using the Prompt Enhancer (3B Model)

ERNIE-Image's core advantage is its built-in 3B-parameter Prompt Enhancer. The MLX port already includes this module:

mlxgen generate \
  --model ernie-image-turbo \
  --prompt "coffee shop" \
  --width 1024 \
  --height 1024 \
  --steps 8 \
  --quantize 8

For short prompts, the built-in 3B Ministral3 model automatically expands them into detailed generation instructions.

2. Batch Generation

# Read configuration from metadata
mlxgen generate --config-from-metadata

Configure in metadata.json:

{
  "model": "ernie-image-turbo",
  "prompts": ["prompt1", "prompt2", "prompt3"],
  "width": 1024,
  "height": 1024,
  "steps": 8,
  "quantize": 8
}

3. Image-to-Image (Experimental)

MLX-Gen supports experimental single-image I2I for ERNIE-Image Turbo:

mlxgen generate \
  --model ernie-image-turbo \
  --image input.jpg \
  --prompt "Enhance the image with cinematic lighting" \
  --width 1024 \
  --height 1024 \
  --steps 8

Performance Benchmarks: MacBook Pro M3 Max vs NVIDIA RTX 4090

Configuration Model Steps Per-Image Time Peak Memory
M3 Max (36GB) + MLX q8 ERNIE-Image Turbo 8 ~15s ~6GB
M3 Max (36GB) + MLX BF16 ERNIE-Image Turbo 8 ~10s ~12GB
M3 Pro (18GB) + MLX q4 ERNIE-Image Turbo 8 ~22s ~4GB
RTX 4090 (24GB) + PyTorch BF16 ERNIE-Image Turbo 8 ~5s ~12GB

Key findings:

  • M3 Max at q8 quantization reaches ~30% of RTX 4090 speed, with significant cost advantage
  • M3 Pro (18GB) runs at q4 mixed quantization but with slower throughput
  • Unified memory means your Mac IS the inference platform — no separate GPU needed

FAQ

Q: Does MLX-Gen support ERNIE-Image Base (non-Turbo)?

MLX-Gen officially supports ERNIE-Image Turbo (distilled version). The Base model (50 steps) requires more memory and compute time — community support is in progress.

Q: Can it run on iPhone/iPad?

MLX technically supports iPadOS 17+ and iOS 17+, but ERNIE-Image's model size (~3-4GB even at q4 quantization) exceeds mobile device storage and memory limits. Currently desktop Apple Silicon only.

Q: How well does the Prompt Enhancer work on MLX?

The 3B Ministral3 PE model runs smoothly on MLX, requiring only ~2GB additional memory at q8 quantization. For short prompts, PE significantly improves generation quality.

Q: How to switch quantization levels?

Use the --quantize flag: --quantize 8 (q8), --quantize 4 (q4). MLX-Gen's mixed precision strategy automatically selects the best quantization level per module.


Conclusion

Apple Silicon + MLX provides ERNIE-Image with a zero-CUDA-dep local deployment path. For users with M1/M2/M3/M4 Macs, this means:

  1. No extra hardware needed: Your Mac IS the inference platform
  2. Native Metal acceleration: 5-10× faster than CPU-only inference
  3. Flexible quantization: Choose from q4 to BF16 based on needs
  4. Workflow-friendly: CLI toolchain covers download, prepare, and generate

If you're looking for a lightweight, local, low-latency AI text-to-image solution, Apple Silicon + MLX + ERNIE-Image Turbo is a combination worth trying.


Keywords: ernie-image apple silicon ernie-image mlx ernie-image mac mflux ernie-image mlx-gen ernie-image apple silicon AI image generation ernie-image Turbo quantization MLX

ERNIE-Image Team