ERNIE-Image Diffusers LoRA Training and Loading Complete Guide: From Error to ComfyUI Deployment
Abstract: The biggest pain point for ERNIE-Image in the Diffusers ecosystem is that
ErnieImagePipelinedoesn't support theload_lora_weights()method. Starting from the GitHub #13501 error reproduction, this article deep-dives into the differences between DiT single-stream architecture and standard UNet LoRA injection, provides the AI Toolkit training workflow, PEFT manual injection solutions, ComfyUI loading pipeline, and a complete production-grade LoRA training-to-deployment guide. Whether you want to train character consistency LoRAs or style customization LoRAs, this article bridges the last mile from training to deployment.
Starting from the Error: Why Diffusers LoRA Loading Fails
If you've tried loading an ERNIE-Image LoRA in Diffusers, you've definitely hit this error:
from diffusers import ErnieImagePipeline
pipe = ErnieImagePipeline.from_pretrained("baidu/ERNIE-Image")
pipe.load_lora_weights("path/to/lora.safetensors") # ❌ Error!
Error message:
AttributeError: 'ErnieImagePipeline' object has no attribute 'load_lora_weights'
This error comes from GitHub diffusers #13501, one of the most frequently asked ERNIE-Image questions in the community.
Root Cause: DiT Single-Stream vs Standard UNet
LoRA (Low-Rank Adaptation) approximates full parameter updates ΔW as the product of two low-rank matrices ΔW ≈ BA. Different architectures require LoRA injection at different layers:
| Architecture | LoRA Injection Point | Diffusers Support |
|---|---|---|
| UNet (SDXL) | cross_attn + self_attn |
✅ Native |
| MMDiT (FLUX.1) | context_attn + img_attn |
✅ Native |
| DiT Single-Stream (ERNIE-Image) | single_transformer_blocks.attn |
❌ Not implemented |
ERNIE-Image uses a single-stream DiT (Diffusion Transformer) architecture, different from FLUX's dual-stream MMDiT and SDXL's UNet. Diffusers' load_lora_weights() is designed for UNet namespace, while ERNIE-Image's weight namespace is:
transformer.single_transformer_blocks.0.attn.to_q.lora_A.weight
transformer.single_transformer_blocks.0.attn.to_k.lora_B.weight
This is why the standard API call fails.
Solution 1: AI Toolkit Training + ComfyUI Loading (Recommended)
Currently the most mature approach: train LoRA with AI Toolkit, load via ComfyUI. AI Toolkit has Day-0 support for ERNIE-Image.
Installing AI Toolkit
git clone https://github.com/ostris/ai-toolkit.git
cd ai-toolkit
pip install -r requirements.txt
Preparing the Dataset
dataset/
├── 001.jpg # images
├── 002.jpg
├── 003.jpg
└── captions.txt # one caption per line
Key configuration (config.yaml):
model: ernie-image
output_dir: ./lora_output
rank: 16 # LoRA rank, recommended 8-32
alpha: 16 # scaling factor, usually equals rank
target_modules:
- "attn.to_q"
- "attn.to_k"
- "attn.to_v"
- "attn.to_out.0"
learning_rate: 1e-4
num_epochs: 20
batch_size: 1
Training Command
python train.py --config config.yaml --dataset dataset/
After training completes, output lora.safetensors.
ComfyUI LoRA Loading
ComfyUI natively supports ERNIE-Image LoRA:
[Load Checkpoint (ernie-image.safetensors)]
↓
[LoraLoader (lora.safetensors, strength=0.8)]
↓
[CLIP Text Encode]
↓
[KSampler] → [VAE Decode] → [Save Image]
Key parameters:
strength: LoRA intensity, typically 0.6-1.0- Too high may cause image distortion
Solution 2: PEFT Manual Injection into Diffusers (Advanced)
If you insist on using Diffusers Pipeline, manually inject LoRA via PEFT:
import torch
from diffusers import ErnieImagePipeline
from peft import LoraConfig, get_peft_model
1. Load base model
pipe = ErnieImagePipeline.from_pretrained("baidu/ERNIE-Image")
2. Configure LoRA
lora_config = LoraConfig(
r=16,
lora_alpha=16,
target_modules=["to_q", "to_k", "to_v", "to_out.0"],
lora_dropout=0.0,
)
3. Inject into DiT model
dit_model = pipe.transformer # ERNIE-Image uses transformer, not unet
peft_model = get_peft_model(dit_model, lora_config)
4. Load pre-trained LoRA weights
lora_state_dict = torch.load("path/to/lora.pt", map_location="cpu")
peft_model.load_state_dict(lora_state_dict, strict=False)
5. Generate image (LoRA is now active)
image = pipe(
prompt="a photorealistic portrait",
num_inference_steps=50,
guidance_scale=7.0,
).images[0]
Notes:
- Weight keys must match training-time namespace (
single_transformer_blocks) - For safetensors format, use
safetensors.torch.load_file() - This bypasses diffusers' standard API — backward compatibility not guaranteed
Solution 3: Manual Loading from safetensors
from safetensors.torch import load_file
lora_weights = load_file("lora.safetensors")
for key, value in lora_weights.items():
if "lora_A" in key or "lora_B" in key:
# Parse target module path and inject
# e.g.: "transformer.single_transformer_blocks.0.attn.to_q.lora_A.weight"
pass
Tip: Want to quickly verify LoRA effects? Use ComfyUI's
LoraLoadernode directly.
LoRA Training Best Practices
Dataset Selection
| Goal | Image Count | Epochs | Recommended Rank |
|---|---|---|---|
| Character consistency | 15-30 | 20-30 | 16 |
| Style transfer | 20-50 | 15-25 | 32 |
| Concept injection | 10-20 | 25-40 | 8-16 |
Prompt Writing Principles
- Consistent prefix: Use the same identifier for all images, e.g.,
a photo of [V] character - Descriptive suffix: Add scene description per image
- Avoid overfitting: Don't repeat the same modifiers in every prompt
Training Parameter Tuning
learning_rate: 1e-4 # Too high = overfitting
num_epochs: 20 # 20 epochs usually enough for character LoRA
rank: 16 # Too small = can't capture details, too large = overfits
alpha: 16 # Usually equals rank
FAQ
Q1: ComfyUI reports "lora key not loaded"
lora key not loaded: transformer.single_transformer_blocks.0.attn.to_k.lora_A.weight
Cause: LoRA trained on different architecture (e.g., FLUX), namespace incompatible.
Solution: Train ERNIE-Image-specific LoRA with AI Toolkit. Don't mix FLUX/SDXL LoRAs.
Q2: Image quality degrades after loading LoRA
Cause: Strength set too high.
Solution: Start from 0.5 and gradually increase. Typically 0.6-0.8 is optimal.
Q3: PEFT injection has no effect
Cause: Weight keys don't match actual layer names.
Solution: Print pipe.transformer.state_dict().keys() to verify namespace.
LoRA Ecosystem Comparison
| Model | Diffusers Native LoRA | AI Toolkit | ComfyUI | Ecosystem Maturity |
|---|---|---|---|---|
| SDXL | ✅ | ✅ | ✅ | ⭐⭐⭐⭐⭐ |
| FLUX.1 | ✅ | ✅ | ✅ | ⭐⭐⭐⭐ |
| ERNIE-Image | ❌ | ✅ | ✅ | ⭐⭐⭐ |
ERNIE-Image's LoRA ecosystem is growing rapidly. AI Toolkit + ComfyUI already covers most scenarios, and it'll become more complete once Diffusers officially implements load_lora_weights().
Summary
ERNIE-Image's single-stream DiT architecture makes its LoRA loading path different from standard Diffusers Pipelines, but through the AI Toolkit training + ComfyUI loading approach, you can complete the full training-to-deployment workflow.
Core recommendations:
- Training: AI Toolkit (Day-0 ERNIE-Image support, outputs safetensors)
- Loading: ComfyUI
LoraLoadernode (most stable) - Advanced users: PEFT manual injection into Diffusers
- Dataset: 15-30 images for good results
Looking ahead: Diffusers is working on ERNIE-Image LoRA support (#13501), which will eventually enable standard pipe.load_lora_weights() API calls.
Keywords: ernie-image diffusers lora ernie-image lora training ernie-image ai toolkit ernie-image load_lora_weights ernie-image comfyui lora ernie-image DiT lora ernie-image lora safetensors