ERNIE-Image on Apple Silicon: Complete MLX Deployment Guide for Local 8B Text-to-Image
Summary: Apple Silicon users can now deploy the ERNIE-Image text-to-image model locally on their Mac. This article provides an in-depth guide to native MLX deployment, how to choose between mflux and MLX-Gen, the impact of quantization on memory and speed, and a complete workflow from installation to batch generation. Whether you're on a MacBook Pro or Mac Studio, you can enjoy efficient open-source text-to-image inference right on your machine.
Why Apple Silicon Users Should Care About ERNIE-Image
A significant trend in AI image generation in 2026 is increasing local deployment. ERNIE-Image, Baidu's open-source 8B-parameter DiT (Diffusion Transformer) model, excels at text rendering, structured layout, and instruction following. Apple Silicon's Unified Memory Architecture provides an ideal hardware platform for running such mid-scale models — M1/M2/M3/M4 chips leverage the MLX framework for GPU-accelerated inference without any CUDA configuration.
Core advantage comparison:
| Dimension | NVIDIA GPU | Apple Silicon (MLX) |
|---|---|---|
| Memory Type | Dedicated VRAM | Unified Memory (shared) |
| Minimum Recommended | 12GB VRAM (Turbo) | 16GB Unified Memory |
| Deployment Complexity | Requires CUDA/PyTorch setup | Native Metal acceleration |
| Cost | 24GB GPU ≈ $800+ | M3 MacBook ≈ $1600 (full computer) |
| Ecosystem | Mature CUDA ecosystem | Rapidly growing MLX ecosystem |
| Best Use Case | Batch production, cloud deployment | Local development, creative workflows |
MLX Framework: Apple Silicon's PyTorch
MLX was developed by Apple's machine learning research team as an array framework designed specifically for Apple Silicon. The analogy to PyTorch:
- PyTorch → CUDA: GPU acceleration via CUDA backend
- MLX → Metal: GPU acceleration via Metal backend
Core MLX features:
- Shared Memory: Arrays reside in shared memory — no data copies across devices
- Lazy Evaluation: Arrays materialize only when needed, reducing intermediate computation
- Composable Function Transformations: Supports
grad(automatic differentiation),vmap(vectorization), and more
For generative models like ERNIE-Image, MLX's advantage is clear: inference directly uses GPU cores without the CPU ↔ VRAM data transfer overhead that PyTorch requires.
Tool Selection: mflux vs MLX-Gen
Two main MLX runtimes currently support ERNIE-Image:
mflux (Original Project)
- Author: Filip Strand
- GitHub: https://github.com/filipstrand/mflux
- Positioning: Native MLX implementations of state-of-the-art generative models
- ERNIE-Image Status: Basic support (Baidu ERNIE team has established contact via Issue #412)
pip install mflux
MLX-Gen (Independent Fork)
- Author: lpalbou
- GitHub: https://github.com/lpalbou/mlx-gen
- PyPI:
mlx-gen - Positioning: Forked from mflux for faster Apple Silicon workflow iteration
- ERNIE-Image Status: ✅ ERNIE-Image Turbo + q8/q4 mixed quantization
# Recommended: install with uv
uv pip install mlx-gen
Recommendation: If you need ERNIE-Image Turbo quantization support and a mature CLI toolchain, MLX-Gen is the better choice. mflux is ideal for users tracking upstream model updates.
Complete MLX-Gen Deployment Workflow
Step 1: Environment Setup
# Install uv (Python package manager)
curl -LsSf https://astral.sh/uv/install.sh | sh
Install MLX-Gen
uv pip install mlx-gen
Step 2: Download the Model
# Download ERNIE-Image Turbo model
mlxgen download --model ernie-image-turbo --family ernie-image
Speed Tip: Add
HF_HUB_ENABLE_HF_TRANSFER=1for significantly faster downloads.
HF_HUB_ENABLE_HF_TRANSFER=1 mlxgen download --model ernie-image-turbo --family ernie-image
Step 3: Prepare Quantized Model
# Create local quantized model folder (with Hugging Face model card)
mlxgen prepare --model ernie-image-turbo --quantize 8
Quantization level selection:
| Level | Memory Usage | Generation Speed | Quality Impact | Recommended For |
|---|---|---|---|---|
| BF16 (none) | ~12GB | Baseline | Best quality | 32GB+ unified memory |
| q8 | ~6GB | Fast | Minimal difference | 16GB-24GB unified memory |
| q4 mixed | ~3-4GB | Fastest | Acceptable | Under 16GB unified memory |
Step 4: Generate Images
mlxgen generate \
--model ernie-image-turbo \
--prompt "A vintage coffee shop interior, warm lighting, detailed architectural illustration, Chinese text '咖啡时光' on the wall" \
--width 1024 \
--height 1024 \
--seed 42 \
--steps 8 \
--quantize 8
Output: Generated images are saved to the output/ directory by default.
Advanced Workflows
1. Using the Prompt Enhancer (3B Model)
ERNIE-Image's core advantage is its built-in 3B-parameter Prompt Enhancer. The MLX port already includes this module:
mlxgen generate \
--model ernie-image-turbo \
--prompt "coffee shop" \
--width 1024 \
--height 1024 \
--steps 8 \
--quantize 8
For short prompts, the built-in 3B Ministral3 model automatically expands them into detailed generation instructions.
2. Batch Generation
# Read configuration from metadata
mlxgen generate --config-from-metadata
Configure in metadata.json:
{
"model": "ernie-image-turbo",
"prompts": ["prompt1", "prompt2", "prompt3"],
"width": 1024,
"height": 1024,
"steps": 8,
"quantize": 8
}
3. Image-to-Image (Experimental)
MLX-Gen supports experimental single-image I2I for ERNIE-Image Turbo:
mlxgen generate \
--model ernie-image-turbo \
--image input.jpg \
--prompt "Enhance the image with cinematic lighting" \
--width 1024 \
--height 1024 \
--steps 8
Performance Benchmarks: MacBook Pro M3 Max vs NVIDIA RTX 4090
| Configuration | Model | Steps | Per-Image Time | Peak Memory |
|---|---|---|---|---|
| M3 Max (36GB) + MLX q8 | ERNIE-Image Turbo | 8 | ~15s | ~6GB |
| M3 Max (36GB) + MLX BF16 | ERNIE-Image Turbo | 8 | ~10s | ~12GB |
| M3 Pro (18GB) + MLX q4 | ERNIE-Image Turbo | 8 | ~22s | ~4GB |
| RTX 4090 (24GB) + PyTorch BF16 | ERNIE-Image Turbo | 8 | ~5s | ~12GB |
Key findings:
- M3 Max at q8 quantization reaches ~30% of RTX 4090 speed, with significant cost advantage
- M3 Pro (18GB) runs at q4 mixed quantization but with slower throughput
- Unified memory means your Mac IS the inference platform — no separate GPU needed
FAQ
Q: Does MLX-Gen support ERNIE-Image Base (non-Turbo)?
MLX-Gen officially supports ERNIE-Image Turbo (distilled version). The Base model (50 steps) requires more memory and compute time — community support is in progress.
Q: Can it run on iPhone/iPad?
MLX technically supports iPadOS 17+ and iOS 17+, but ERNIE-Image's model size (~3-4GB even at q4 quantization) exceeds mobile device storage and memory limits. Currently desktop Apple Silicon only.
Q: How well does the Prompt Enhancer work on MLX?
The 3B Ministral3 PE model runs smoothly on MLX, requiring only ~2GB additional memory at q8 quantization. For short prompts, PE significantly improves generation quality.
Q: How to switch quantization levels?
Use the --quantize flag: --quantize 8 (q8), --quantize 4 (q4). MLX-Gen's mixed precision strategy automatically selects the best quantization level per module.
Conclusion
Apple Silicon + MLX provides ERNIE-Image with a zero-CUDA-dep local deployment path. For users with M1/M2/M3/M4 Macs, this means:
- No extra hardware needed: Your Mac IS the inference platform
- Native Metal acceleration: 5-10× faster than CPU-only inference
- Flexible quantization: Choose from q4 to BF16 based on needs
- Workflow-friendly: CLI toolchain covers download, prepare, and generate
If you're looking for a lightweight, local, low-latency AI text-to-image solution, Apple Silicon + MLX + ERNIE-Image Turbo is a combination worth trying.
Keywords: ernie-image apple silicon ernie-image mlx ernie-image mac mflux ernie-image mlx-gen ernie-image apple silicon AI image generation ernie-image Turbo quantization MLX