ERNIE-Image × mflux: Native MLX Inference on Apple Silicon — Complete Guide

Jul 9, 2026

ERNIE-Image × mflux: Native MLX Inference on Apple Silicon — Complete Guide

Published: 2026-07-09
Platform: ernie-image.app

1. Background: A New Option for Apple Silicon Users

Since ERNIE-Image's release three months ago, the community has explored many local deployment approaches — from SGLang inference servers to GGUF quantization, AMD ROCm to Google Colab. For Apple Silicon Mac users, the primary option has been running through Diffusers + PyTorch MPS backend, which has a significant pain point: MPS backend performance is suboptimal, and it adds an extra abstraction layer through PyTorch.

In July 2026, the mflux project officially added ERNIE-Image support, offering Mac users a new path: no Diffusers, no PyTorch — run ERNIE-Image and ERNIE-Image-Turbo natively on the MLX framework.

mflux is a "line-by-line MLX port" project that directly ports generative image models from HuggingFace Diffusers and Transformers to Apple's MLX framework. The benefit is clear — removing the middle layer lets the model run directly on Metal Performance Shaders, with better performance and lower overhead.

mflux currently supports both ERNIE-Image (Base) and ERNIE-Image-Turbo, covering text-to-image, image-to-image, LoRA fine-tuning, and quantized inference.

2. What is mflux and Why Choose It?

mflux's core philosophy is simple: all models are implemented from scratch in MLX, using only HuggingFace Transformers' tokenizer components. The code style is influenced by Andrej Karpathy's "build from scratch" philosophy — minimal and explicit.

For Mac users, mflux offers several advantages over Diffusers + MPS:

  • Native MLX inference: Model weights are directly converted from PyTorch to MLX format (NCHW → NHWC), no PyTorch MPS bridge needed
  • Lower disk usage: Quantization (q8/q4) significantly reduces model file sizes
  • LoRA training support: Built-in mflux-train tool supports LoRA fine-tuning
  • Batch generation: Generate multiple images in one go
  • HTTP API server: Start a remote inference server for API access

Community developer treadon has also pre-converted ERNIE-Image-Turbo's MLX weights and uploaded them to HuggingFace (treadon/ERNIE-Image-Turbo-MLX) — ready to use out of the box.

3. Installation and Quick Start

Installation

mflux uses uv as its package manager:

# Install uv (if not already installed)
curl -LsSf https://astral.sh/uv/install.sh | sh

Install mflux

uv pip install mflux

For faster downloads, add the hf_transfer extra:

uv pip install "mflux[hf_transfer]"

Basic Usage: Text-to-Image

Generate your first image with ERNIE-Image-Turbo:

mflux-generate-ernie-image-turbo \
  --prompt "A black and white Chinese rural dog running on grassland, sunny" \
  --width 1024 \
  --height 1024 \
  --seed 42 \
  --steps 8 \
  -q 8

The first run automatically downloads model weights (~12GB q8 quantized). Subsequent runs use local cache.

Parameter notes:

  • -q 8: 8-bit quantization (~12GB disk), recommended for most users
  • -q 4: 4-bit quantization (~6.2GB disk), for limited storage
  • --steps: Turbo version recommends 8 steps, Base version 28-50 steps
  • --guidance: Base supports CFG (recommended 4.0), Turbo is fixed at 1.0

Using the Base Model

The Base (non-distilled) ERNIE-Image needs more steps but supports classifier-free guidance:

mflux-generate-ernie-image \
  --prompt "A retro cafe sign with Japanese menu text, warm lighting" \
  --width 1024 \
  --height 1024 \
  --seed 42 \
  --steps 28 \
  --guidance 4.0

Image-to-Image

mflux supports img2img via --image-path and --image-strength:

mflux-generate-ernie-image-turbo \
  --prompt "Transform this sketch into a beautiful watercolor painting" \
  --image-path /path/to/sketch.png \
  --image-strength 0.6 \
  --steps 8 \
  -q 8

--image-strength ranges from 0.0 to 1.0 — higher values preserve more of the original image.

4. Performance Benchmarks

Based on tests by community developer treadon on an M4 Pro with 64GB:

Pipeline Total Time Per Step
PyTorch/MPS (diffusers) 137.0s 17.1s/step
MLX (mflux) 134.2s 16.0s/step

Component breakdown:

Component Time
Text encode (PyTorch) 0.1s
Denoise (MLX) 128s
VAE decode (MLX) 6s
Total ~134s

This means generating a 1024×1024 image on an M4 Pro takes about 2 minutes. While not as fast as cloud GPU solutions (~9s on fal.ai), it's practical for offline use and privacy-sensitive work.

Expected performance on different Mac configurations:

  • M1 Max (64GB): ~180-220s
  • M2 Ultra (128GB): ~100-120s
  • M4 Pro (64GB): ~134s (reference)
  • M4 Max (128GB): ~80-100s

5. Quantization and Disk Usage

Full-precision ERNIE-Image weights are about 22GB. mflux supports two quantization levels:

Quantization Disk Usage Notes
None (float16) ~22GB Full precision, high storage requirement
q8 (8-bit) ~12GB Recommended, minimal quality loss
q4 (4-bit) ~6.2GB For limited storage scenarios

Save quantized models with mflux-save:

# Download and save as q8
mflux-save --model ernie-image-turbo -q 8

Download and save as q4

mflux-save --model ernie-image-turbo -q 4

6. LoRA Fine-Tuning on Your Mac

mflux includes the mflux-train tool for LoRA fine-tuning directly on Apple Silicon:

mflux-train --config train_ernie_image_turbo.json

Example config (train_ernie_image_turbo.json):

{
  "model": "ernie-image-turbo",
  "dataset_path": "./my_dataset",
  "lora_rank": 16,
  "lora_alpha": 32,
  "learning_rate": 1e-4,
  "num_epochs": 10,
  "batch_size": 1,
  "quantize": 8
}

Supported target layers:

  • mlp.* — MLP layers
  • time_embedding.* — Time embedding layers
  • adaln_modulation — Adaptive Layer Norm modulation
  • final_norm.linear — Final normalization linear layer

Use trained LoRAs during inference:

mflux-generate-ernie-image-turbo \
  --prompt "my_style: a cat" \
  --lora-weights /path/to/lora.safetensors \
  --lora-scale 0.8

7. Python API and HTTP Server

Python Script

# /// script
# dependencies = ["mflux"]
# ///

from mflux import mflux_generate

mflux_generate(
model="ernie-image-turbo",
prompt="A puffin standing on a cliff",
width=1280,
height=500,
seed=42,
steps=8,
quantize=8
)

HTTP Server Mode

mflux-server --model ernie-image-turbo -q 8 --port 8080

REST API access:

curl -X POST http://localhost:8080/generate \
  -H "Content-Type: application/json" \
  -d '{"prompt": "Cyberpunk Tokyo night scene", "width": 1024, "height": 1024}'

8. Use Cases and Considerations

Best Use Cases

  • Offline environments: Generate high-quality images without internet
  • Privacy-sensitive projects: Data never leaves your machine
  • Prototyping and testing: Quickly test prompt effects
  • LoRA experimentation: Train and test style LoRAs at zero GPU cost
  • Educational demos: Demonstrate ERNIE-Image capabilities on a MacBook

Important Notes

  1. Memory requirements: ERNIE-Image-Turbo q8 needs ~16GB unified memory; 24GB+ recommended
  2. First download: ~12GB (q8), first run takes time for download
  3. Speed vs cloud: Local inference is much slower than cloud GPU (~9s on fal.ai), suitable for low-throughput scenarios
  4. Turbo first: Use Turbo version (8 steps) on Mac; Base version (50 steps) is very slow

9. Summary

mflux provides Apple Silicon Mac users with a practical path to run ERNIE-Image without cloud GPU dependency. While local inference speed can't match cloud H100s, the offline availability, data privacy, and native LoRA training support make it a valuable complement for developers and creators using ERNIE-Image on Mac.

Key takeaways:

  • mflux natively supports ERNIE-Image and ERNIE-Image-Turbo
  • 8-bit quantization: ~12GB disk, ~134s per image on M4 Pro
  • Supports LoRA training, img2img, and HTTP API server
  • Pre-converted MLX weights available (treadon/ERNIE-Image-Turbo-MLX)
  • Recommended config: q8 quantization + Turbo version

Keywords: ernie-image mflux ernie-image apple silicon ernie-image mlx ernie-image mac mflux ernie-image turbo mlx apple silicon ai image generation

ERNIE-Image Team