ERNIE-Image Custom Style Fine-Tuning: Train Your Own LoRA from Scratch

Jun 3, 2026

ERNIE-Image Custom Style Fine-Tuning: Train Your Own LoRA from Scratch

Summary: ERNIE-Image's LoRA fine-tuning capability lets individual creators train their own artistic styles. This guide takes you from zero to deploying a custom style LoRA — dataset preparation, environment setup, training, and loading in ComfyUI. All local, completely free.


Why Style Fine-Tuning?

ERNIE-Image's base model already covers a wide range of styles — from photorealism to anime, watercolor to oil painting. But when you have unique visual style requirements, the base model often falls short:

  • Brand visual consistency: Fixed visual style for e-commerce products
  • Personal art signature: Your unique painting style or color palette
  • IP character consistency: Consistent character appearance across a series
  • Specific media styles: Cyberpunk, vintage posters, minimalism, etc.

LoRA (Low-Rank Adaptation) fine-tuning trains just 0.5%~1.5% of model parameters to teach ERNIE-Image your signature style. Compared to full fine-tuning, it's faster, cheaper, and switchable like a plugin.

Hardware Requirements

Configuration Minimum Recommended
GPU NVIDIA 12GB VRAM NVIDIA 24GB+ VRAM
Quantized version 8GB VRAM (GGUF/NVFP4) —
RAM 16GB 32GB
Disk 20GB free 40GB+

Tip: If your GPU has less than 12GB VRAM, consider cloud GPU services (RunPod, Vast.ai) at roughly $0.20~$0.50/hour.

Step 1: Prepare Your Style Dataset

Dataset Requirements

  • Image count: 10~50 (more consistent = better results)
  • Resolution:统一 scale to 1024×1024 or 1024×768
  • Format: JPG or PNG
  • Style consistency: All images should share the same visual style

Example Dataset Structure

my-style-dataset/
├── images/
│   ├── style_001.jpg
│   ├── style_002.jpg
│   ├── ...
│   └── style_050.jpg
├── captions/
│   ├── style_001.txt  # "a digital painting of a landscape in watercolor style"
│   ├── style_002.txt  # "a digital painting of a portrait in watercolor style"
│   └── ...
└── metadata.json

Generating Captions

Captions tell the model "what this image depicts," while style information is learned from visual features. Use this format:

a [style_keyword] of [subject]

Examples:

  • a watercolor painting of a cat
  • a cyberpunk digital art of a city street
  • a vintage poster illustration of a woman

Tip: You can auto-generate captions using CLIP or BLIP models, then manually adjust style keywords.

Step 2: Set Up the Training Environment

Method 1: Official Diffusers + PEFT

# Create virtual environment
python3 -m venv ernie-lora-env
source ernie-lora-env/bin/activate

Install dependencies

pip install torch torchvision --index-url https://download.pytorch.org/whl/cu121
pip install diffusers transformers accelerate peft bitsandbytes
pip install datasets pillow timm

Clone ERNIE-Image repo

git clone https://github.com/baidu/ERNIE-Image.git
cd ERNIE-Image

Method 2: DreamBooth Training Script

ERNIE-Image is built on the Diffusers architecture, so you can directly use HuggingFace's official train_dreambooth.py script:

git clone https://github.com/huggingface/diffusers.git
cd examples/dreambooth
pip install -r requirements.txt

Step 3: Run LoRA Fine-Tuning

Training Command

accelerate launch train_dreambooth_lora.py \
  --pretrained_model_name_or_path=baidu/ERNIE-Image \
  --instance_data_dir=./my-style-dataset/images \
  --instance_prompt="a {style_keyword} of subject" \
  --output_dir=./lora-outputs \
  --learning_rate=1e-4 \
  --max_train_steps=1000 \
  --train_batch_size=1 \
  --gradient_accumulation_steps=4 \
  --resolution=1024 \
  --mixed_precision=bf16 \
  --lora_rank=32 \
  --lora_alpha=16 \
  --lora_target_modules=["to_q","to_k","to_v","to_out.0"] \
  --checkpointing_steps=200 \
  --seed=42

Key Parameters

Parameter Recommended Notes
learning_rate 1e-4 ~ 5e-4 Too high = overfitting, too low = undertrained
max_train_steps 500~2000 Fewer images = fewer steps, more images = more steps
lora_rank 16~64 Higher = more expressive but larger file
lora_alpha lora_rank/2 Typically half of rank
mixed_precision bf16 Needs Ampere+ GPU; older GPUs use fp16

Training Time Reference

Images Steps Learning Rate RTX 3090 RTX 4090
15 500 3e-4 ~15 min ~10 min
30 1000 2e-4 ~30 min ~20 min
50 2000 1e-4 ~60 min ~40 min

Step 4: Validate Fine-Tuning Results

Quick Test

After training, find pytorch_lora_weights.safetensors in your lora-outputs directory. Test with:

from diffusers import DiffusionPipeline

Load base model

pipe = DiffusionPipeline.from_pretrained(
"baidu/ERNIE-Image",
torch_dtype=torch.float16
)
pipe.to("cuda")

Load LoRA

pipe.load_lora_weights("./lora-outputs", weight_name="pytorch_lora_weights.safetensors")

Generate test image

prompt = "a watercolor painting of a mountain landscape, soft colors, flowing brush strokes"
image = pipe(prompt, num_inference_steps=50, guidance_scale=7.5).images[0]
image.save("test_output.png")

Evaluation Checklist

  1. Style consistency: Do different subject prompts maintain the unified style?
  2. Overfitting: Is the model just copying training images instead of learning the style?
  3. Prompt responsiveness: When changing the subject in the prompt, does the style follow?

Step 5: Use LoRA in ComfyUI

Installation Steps

  1. Copy pytorch_lora_weights.safetensors to ComfyUI/models/loras/
  2. Launch ComfyUI
  3. Add a "Load LoRA" node to your workflow
  4. Select your LoRA model file
  5. Set LoRA strength (recommended: 0.6~0.9)

ComfyUI Workflow Example

[Load Checkpoint: ERNIE-Image]
       ↓
[Load LoRA: your-style-lora.safetensors, strength=0.8]
       ↓
[CLIP Text Encode (Positive Prompt)]
       ↓
[KSampler] → [VAE Decode] → [Save Image]

Tip: A LoRA strength of 0.8 typically gives the best results. Too high (>1.0) may over-stylize, too low (<0.5) has minimal effect.

Common Issues and Solutions

Issue 1: CUDA Out of Memory (OOM)

RuntimeError: CUDA out of memory

Solutions:

  • Lower resolution to 768 or 512
  • Enable gradient checkpointing: --gradient_checkpointing
  • Reduce train_batch_size or increase gradient_accumulation_steps
  • Use 8-bit optimizer: --optimizer_type=8bit_adam

Issue 2: Output too similar to training images

This indicates overfitting.

Solutions:

  • Lower learning rate (from 3e-4 to 1e-4)
  • Reduce training steps
  • Add more variation in instance prompts
  • Use regularization images

Issue 3: Style not prominent enough

Solutions:

  • Increase lora_rank to 64
  • Increase training steps
  • Ensure training data style is distinct and unified
  • Raise LoRA strength to 0.9~1.0 in ComfyUI

Issue 4: Stacking Multiple LoRAs

ERNIE-Image supports loading multiple LoRAs simultaneously for style + character combinations:

pipe.load_lora_weights("./style-lora", weight_name="pytorch_lora_weights.safetensors", adapter_name="style")
pipe.load_lora_weights("./character-lora", weight_name="pytorch_lora_weights.safetensors", adapter_name="character")
pipe.set_adapters(["style", "character"], adapter_weights=[0.8, 0.7])

Summary

ERNIE-Image's LoRA style fine-tuning is a zero-cost personalization capability. With 1050 style-consistent images and 1560 minutes of training on a consumer GPU, you can own a custom artistic style model.

Workflow Recap:

  1. Prepare 10~50 style-consistent images + captions
  2. Set up Diffusers + PEFT training environment
  3. Run LoRA fine-tuning (lr=1e-4, steps=1000, rank=32)
  4. Validate style results
  5. Load and use in ComfyUI

Compared to Midjourney's --sref (style reference) or FLUX's reference image features, ERNIE-Image LoRA offers more precise, controllable style transfer — and it's completely free, running locally.


Originally published on ernie-image.app. Please credit when sharing.

ERNIE-Image Team