ERNIE-Image Diffusers LoRA Training and Loading Complete Guide: From Error to ComfyUI Deployment

Jun 13, 2026

ERNIE-Image Diffusers LoRA Training and Loading Complete Guide: From Error to ComfyUI Deployment

Abstract: The biggest pain point for ERNIE-Image in the Diffusers ecosystem is that ErnieImagePipeline doesn't support the load_lora_weights() method. Starting from the GitHub #13501 error reproduction, this article deep-dives into the differences between DiT single-stream architecture and standard UNet LoRA injection, provides the AI Toolkit training workflow, PEFT manual injection solutions, ComfyUI loading pipeline, and a complete production-grade LoRA training-to-deployment guide. Whether you want to train character consistency LoRAs or style customization LoRAs, this article bridges the last mile from training to deployment.


Starting from the Error: Why Diffusers LoRA Loading Fails

If you've tried loading an ERNIE-Image LoRA in Diffusers, you've definitely hit this error:

from diffusers import ErnieImagePipeline

pipe = ErnieImagePipeline.from_pretrained("baidu/ERNIE-Image")
pipe.load_lora_weights("path/to/lora.safetensors") # ❌ Error!

Error message:

AttributeError: 'ErnieImagePipeline' object has no attribute 'load_lora_weights'

This error comes from GitHub diffusers #13501, one of the most frequently asked ERNIE-Image questions in the community.

Root Cause: DiT Single-Stream vs Standard UNet

LoRA (Low-Rank Adaptation) approximates full parameter updates ΔW as the product of two low-rank matrices ΔW ≈ BA. Different architectures require LoRA injection at different layers:

Architecture LoRA Injection Point Diffusers Support
UNet (SDXL) cross_attn + self_attn ✅ Native
MMDiT (FLUX.1) context_attn + img_attn ✅ Native
DiT Single-Stream (ERNIE-Image) single_transformer_blocks.attn ❌ Not implemented

ERNIE-Image uses a single-stream DiT (Diffusion Transformer) architecture, different from FLUX's dual-stream MMDiT and SDXL's UNet. Diffusers' load_lora_weights() is designed for UNet namespace, while ERNIE-Image's weight namespace is:

transformer.single_transformer_blocks.0.attn.to_q.lora_A.weight
transformer.single_transformer_blocks.0.attn.to_k.lora_B.weight

This is why the standard API call fails.


Solution 1: AI Toolkit Training + ComfyUI Loading (Recommended)

Currently the most mature approach: train LoRA with AI Toolkit, load via ComfyUI. AI Toolkit has Day-0 support for ERNIE-Image.

Installing AI Toolkit

git clone https://github.com/ostris/ai-toolkit.git
cd ai-toolkit
pip install -r requirements.txt

Preparing the Dataset

dataset/
├── 001.jpg    # images
├── 002.jpg
├── 003.jpg
└── captions.txt  # one caption per line

Key configuration (config.yaml):

model: ernie-image
output_dir: ./lora_output
rank: 16            # LoRA rank, recommended 8-32
alpha: 16           # scaling factor, usually equals rank
target_modules:
  - "attn.to_q"
  - "attn.to_k"
  - "attn.to_v"
  - "attn.to_out.0"
learning_rate: 1e-4
num_epochs: 20
batch_size: 1

Training Command

python train.py --config config.yaml --dataset dataset/

After training completes, output lora.safetensors.

ComfyUI LoRA Loading

ComfyUI natively supports ERNIE-Image LoRA:

[Load Checkpoint (ernie-image.safetensors)]
       ↓
[LoraLoader (lora.safetensors, strength=0.8)]
       ↓
[CLIP Text Encode]
       ↓
[KSampler] → [VAE Decode] → [Save Image]

Key parameters:

  • strength: LoRA intensity, typically 0.6-1.0
  • Too high may cause image distortion

Solution 2: PEFT Manual Injection into Diffusers (Advanced)

If you insist on using Diffusers Pipeline, manually inject LoRA via PEFT:

import torch
from diffusers import ErnieImagePipeline
from peft import LoraConfig, get_peft_model

1. Load base model

pipe = ErnieImagePipeline.from_pretrained("baidu/ERNIE-Image")

2. Configure LoRA

lora_config = LoraConfig(
r=16,
lora_alpha=16,
target_modules=["to_q", "to_k", "to_v", "to_out.0"],
lora_dropout=0.0,
)

3. Inject into DiT model

dit_model = pipe.transformer # ERNIE-Image uses transformer, not unet
peft_model = get_peft_model(dit_model, lora_config)

4. Load pre-trained LoRA weights

lora_state_dict = torch.load("path/to/lora.pt", map_location="cpu")
peft_model.load_state_dict(lora_state_dict, strict=False)

5. Generate image (LoRA is now active)

image = pipe(
prompt="a photorealistic portrait",
num_inference_steps=50,
guidance_scale=7.0,
).images[0]

Notes:

  • Weight keys must match training-time namespace (single_transformer_blocks)
  • For safetensors format, use safetensors.torch.load_file()
  • This bypasses diffusers' standard API — backward compatibility not guaranteed

Solution 3: Manual Loading from safetensors

from safetensors.torch import load_file

lora_weights = load_file("lora.safetensors")

for key, value in lora_weights.items():
if "lora_A" in key or "lora_B" in key:
# Parse target module path and inject
# e.g.: "transformer.single_transformer_blocks.0.attn.to_q.lora_A.weight"
pass

Tip: Want to quickly verify LoRA effects? Use ComfyUI's LoraLoader node directly.


LoRA Training Best Practices

Dataset Selection

Goal Image Count Epochs Recommended Rank
Character consistency 15-30 20-30 16
Style transfer 20-50 15-25 32
Concept injection 10-20 25-40 8-16

Prompt Writing Principles

  1. Consistent prefix: Use the same identifier for all images, e.g., a photo of [V] character
  2. Descriptive suffix: Add scene description per image
  3. Avoid overfitting: Don't repeat the same modifiers in every prompt

Training Parameter Tuning

learning_rate: 1e-4        # Too high = overfitting
num_epochs: 20             # 20 epochs usually enough for character LoRA
rank: 16                   # Too small = can't capture details, too large = overfits
alpha: 16                  # Usually equals rank

FAQ

Q1: ComfyUI reports "lora key not loaded"

lora key not loaded: transformer.single_transformer_blocks.0.attn.to_k.lora_A.weight

Cause: LoRA trained on different architecture (e.g., FLUX), namespace incompatible.

Solution: Train ERNIE-Image-specific LoRA with AI Toolkit. Don't mix FLUX/SDXL LoRAs.

Q2: Image quality degrades after loading LoRA

Cause: Strength set too high.

Solution: Start from 0.5 and gradually increase. Typically 0.6-0.8 is optimal.

Q3: PEFT injection has no effect

Cause: Weight keys don't match actual layer names.

Solution: Print pipe.transformer.state_dict().keys() to verify namespace.


LoRA Ecosystem Comparison

Model Diffusers Native LoRA AI Toolkit ComfyUI Ecosystem Maturity
SDXL ✅ ✅ ✅ ⭐⭐⭐⭐⭐
FLUX.1 ✅ ✅ ✅ ⭐⭐⭐⭐
ERNIE-Image ❌ ✅ ✅ ⭐⭐⭐

ERNIE-Image's LoRA ecosystem is growing rapidly. AI Toolkit + ComfyUI already covers most scenarios, and it'll become more complete once Diffusers officially implements load_lora_weights().


Summary

ERNIE-Image's single-stream DiT architecture makes its LoRA loading path different from standard Diffusers Pipelines, but through the AI Toolkit training + ComfyUI loading approach, you can complete the full training-to-deployment workflow.

Core recommendations:

  1. Training: AI Toolkit (Day-0 ERNIE-Image support, outputs safetensors)
  2. Loading: ComfyUI LoraLoader node (most stable)
  3. Advanced users: PEFT manual injection into Diffusers
  4. Dataset: 15-30 images for good results

Looking ahead: Diffusers is working on ERNIE-Image LoRA support (#13501), which will eventually enable standard pipe.load_lora_weights() API calls.


Keywords: ernie-image diffusers lora ernie-image lora training ernie-image ai toolkit ernie-image load_lora_weights ernie-image comfyui lora ernie-image DiT lora ernie-image lora safetensors

ERNIE-Image Team