ERNIE-Image × mflux: Native MLX Inference on Apple Silicon — Complete Guide
Published: 2026-07-09
Platform: ernie-image.app
1. Background: A New Option for Apple Silicon Users
Since ERNIE-Image's release three months ago, the community has explored many local deployment approaches — from SGLang inference servers to GGUF quantization, AMD ROCm to Google Colab. For Apple Silicon Mac users, the primary option has been running through Diffusers + PyTorch MPS backend, which has a significant pain point: MPS backend performance is suboptimal, and it adds an extra abstraction layer through PyTorch.
In July 2026, the mflux project officially added ERNIE-Image support, offering Mac users a new path: no Diffusers, no PyTorch — run ERNIE-Image and ERNIE-Image-Turbo natively on the MLX framework.
mflux is a "line-by-line MLX port" project that directly ports generative image models from HuggingFace Diffusers and Transformers to Apple's MLX framework. The benefit is clear — removing the middle layer lets the model run directly on Metal Performance Shaders, with better performance and lower overhead.
mflux currently supports both ERNIE-Image (Base) and ERNIE-Image-Turbo, covering text-to-image, image-to-image, LoRA fine-tuning, and quantized inference.
2. What is mflux and Why Choose It?
mflux's core philosophy is simple: all models are implemented from scratch in MLX, using only HuggingFace Transformers' tokenizer components. The code style is influenced by Andrej Karpathy's "build from scratch" philosophy — minimal and explicit.
For Mac users, mflux offers several advantages over Diffusers + MPS:
- Native MLX inference: Model weights are directly converted from PyTorch to MLX format (NCHW → NHWC), no PyTorch MPS bridge needed
- Lower disk usage: Quantization (q8/q4) significantly reduces model file sizes
- LoRA training support: Built-in mflux-train tool supports LoRA fine-tuning
- Batch generation: Generate multiple images in one go
- HTTP API server: Start a remote inference server for API access
Community developer treadon has also pre-converted ERNIE-Image-Turbo's MLX weights and uploaded them to HuggingFace (treadon/ERNIE-Image-Turbo-MLX) — ready to use out of the box.
3. Installation and Quick Start
Installation
mflux uses uv as its package manager:
# Install uv (if not already installed)
curl -LsSf https://astral.sh/uv/install.sh | sh
Install mflux
uv pip install mflux
For faster downloads, add the hf_transfer extra:
uv pip install "mflux[hf_transfer]"
Basic Usage: Text-to-Image
Generate your first image with ERNIE-Image-Turbo:
mflux-generate-ernie-image-turbo \
--prompt "A black and white Chinese rural dog running on grassland, sunny" \
--width 1024 \
--height 1024 \
--seed 42 \
--steps 8 \
-q 8
The first run automatically downloads model weights (~12GB q8 quantized). Subsequent runs use local cache.
Parameter notes:
-q 8: 8-bit quantization (~12GB disk), recommended for most users-q 4: 4-bit quantization (~6.2GB disk), for limited storage--steps: Turbo version recommends 8 steps, Base version 28-50 steps--guidance: Base supports CFG (recommended 4.0), Turbo is fixed at 1.0
Using the Base Model
The Base (non-distilled) ERNIE-Image needs more steps but supports classifier-free guidance:
mflux-generate-ernie-image \
--prompt "A retro cafe sign with Japanese menu text, warm lighting" \
--width 1024 \
--height 1024 \
--seed 42 \
--steps 28 \
--guidance 4.0
Image-to-Image
mflux supports img2img via --image-path and --image-strength:
mflux-generate-ernie-image-turbo \
--prompt "Transform this sketch into a beautiful watercolor painting" \
--image-path /path/to/sketch.png \
--image-strength 0.6 \
--steps 8 \
-q 8
--image-strength ranges from 0.0 to 1.0 — higher values preserve more of the original image.
4. Performance Benchmarks
Based on tests by community developer treadon on an M4 Pro with 64GB:
| Pipeline | Total Time | Per Step |
|---|---|---|
| PyTorch/MPS (diffusers) | 137.0s | 17.1s/step |
| MLX (mflux) | 134.2s | 16.0s/step |
Component breakdown:
| Component | Time |
|---|---|
| Text encode (PyTorch) | 0.1s |
| Denoise (MLX) | 128s |
| VAE decode (MLX) | 6s |
| Total | ~134s |
This means generating a 1024×1024 image on an M4 Pro takes about 2 minutes. While not as fast as cloud GPU solutions (~9s on fal.ai), it's practical for offline use and privacy-sensitive work.
Expected performance on different Mac configurations:
- M1 Max (64GB): ~180-220s
- M2 Ultra (128GB): ~100-120s
- M4 Pro (64GB): ~134s (reference)
- M4 Max (128GB): ~80-100s
5. Quantization and Disk Usage
Full-precision ERNIE-Image weights are about 22GB. mflux supports two quantization levels:
| Quantization | Disk Usage | Notes |
|---|---|---|
| None (float16) | ~22GB | Full precision, high storage requirement |
| q8 (8-bit) | ~12GB | Recommended, minimal quality loss |
| q4 (4-bit) | ~6.2GB | For limited storage scenarios |
Save quantized models with mflux-save:
# Download and save as q8
mflux-save --model ernie-image-turbo -q 8
Download and save as q4
mflux-save --model ernie-image-turbo -q 4
6. LoRA Fine-Tuning on Your Mac
mflux includes the mflux-train tool for LoRA fine-tuning directly on Apple Silicon:
mflux-train --config train_ernie_image_turbo.json
Example config (train_ernie_image_turbo.json):
{
"model": "ernie-image-turbo",
"dataset_path": "./my_dataset",
"lora_rank": 16,
"lora_alpha": 32,
"learning_rate": 1e-4,
"num_epochs": 10,
"batch_size": 1,
"quantize": 8
}
Supported target layers:
mlp.*— MLP layerstime_embedding.*— Time embedding layersadaln_modulation— Adaptive Layer Norm modulationfinal_norm.linear— Final normalization linear layer
Use trained LoRAs during inference:
mflux-generate-ernie-image-turbo \
--prompt "my_style: a cat" \
--lora-weights /path/to/lora.safetensors \
--lora-scale 0.8
7. Python API and HTTP Server
Python Script
# /// script
# dependencies = ["mflux"]
# ///
from mflux import mflux_generate
mflux_generate(
model="ernie-image-turbo",
prompt="A puffin standing on a cliff",
width=1280,
height=500,
seed=42,
steps=8,
quantize=8
)
HTTP Server Mode
mflux-server --model ernie-image-turbo -q 8 --port 8080
REST API access:
curl -X POST http://localhost:8080/generate \
-H "Content-Type: application/json" \
-d '{"prompt": "Cyberpunk Tokyo night scene", "width": 1024, "height": 1024}'
8. Use Cases and Considerations
Best Use Cases
- Offline environments: Generate high-quality images without internet
- Privacy-sensitive projects: Data never leaves your machine
- Prototyping and testing: Quickly test prompt effects
- LoRA experimentation: Train and test style LoRAs at zero GPU cost
- Educational demos: Demonstrate ERNIE-Image capabilities on a MacBook
Important Notes
- Memory requirements: ERNIE-Image-Turbo q8 needs ~16GB unified memory; 24GB+ recommended
- First download: ~12GB (q8), first run takes time for download
- Speed vs cloud: Local inference is much slower than cloud GPU (~9s on fal.ai), suitable for low-throughput scenarios
- Turbo first: Use Turbo version (8 steps) on Mac; Base version (50 steps) is very slow
9. Summary
mflux provides Apple Silicon Mac users with a practical path to run ERNIE-Image without cloud GPU dependency. While local inference speed can't match cloud H100s, the offline availability, data privacy, and native LoRA training support make it a valuable complement for developers and creators using ERNIE-Image on Mac.
Key takeaways:
- mflux natively supports ERNIE-Image and ERNIE-Image-Turbo
- 8-bit quantization: ~12GB disk, ~134s per image on M4 Pro
- Supports LoRA training, img2img, and HTTP API server
- Pre-converted MLX weights available (treadon/ERNIE-Image-Turbo-MLX)
- Recommended config: q8 quantization + Turbo version
Keywords: ernie-image mflux ernie-image apple silicon ernie-image mlx ernie-image mac mflux ernie-image turbo mlx apple silicon ai image generation