ERNIE-Image Custom Style Fine-Tuning: Train Your Own LoRA from Scratch
Summary: ERNIE-Image's LoRA fine-tuning capability lets individual creators train their own artistic styles. This guide takes you from zero to deploying a custom style LoRA — dataset preparation, environment setup, training, and loading in ComfyUI. All local, completely free.
Why Style Fine-Tuning?
ERNIE-Image's base model already covers a wide range of styles — from photorealism to anime, watercolor to oil painting. But when you have unique visual style requirements, the base model often falls short:
- Brand visual consistency: Fixed visual style for e-commerce products
- Personal art signature: Your unique painting style or color palette
- IP character consistency: Consistent character appearance across a series
- Specific media styles: Cyberpunk, vintage posters, minimalism, etc.
LoRA (Low-Rank Adaptation) fine-tuning trains just 0.5%~1.5% of model parameters to teach ERNIE-Image your signature style. Compared to full fine-tuning, it's faster, cheaper, and switchable like a plugin.
Hardware Requirements
| Configuration | Minimum | Recommended |
|---|---|---|
| GPU | NVIDIA 12GB VRAM | NVIDIA 24GB+ VRAM |
| Quantized version | 8GB VRAM (GGUF/NVFP4) | — |
| RAM | 16GB | 32GB |
| Disk | 20GB free | 40GB+ |
Tip: If your GPU has less than 12GB VRAM, consider cloud GPU services (RunPod, Vast.ai) at roughly $0.20~$0.50/hour.
Step 1: Prepare Your Style Dataset
Dataset Requirements
- Image count: 10~50 (more consistent = better results)
- Resolution:统一 scale to 1024×1024 or 1024×768
- Format: JPG or PNG
- Style consistency: All images should share the same visual style
Example Dataset Structure
my-style-dataset/
├── images/
│ ├── style_001.jpg
│ ├── style_002.jpg
│ ├── ...
│ └── style_050.jpg
├── captions/
│ ├── style_001.txt # "a digital painting of a landscape in watercolor style"
│ ├── style_002.txt # "a digital painting of a portrait in watercolor style"
│ └── ...
└── metadata.json
Generating Captions
Captions tell the model "what this image depicts," while style information is learned from visual features. Use this format:
a [style_keyword] of [subject]
Examples:
a watercolor painting of a cata cyberpunk digital art of a city streeta vintage poster illustration of a woman
Tip: You can auto-generate captions using CLIP or BLIP models, then manually adjust style keywords.
Step 2: Set Up the Training Environment
Method 1: Official Diffusers + PEFT
# Create virtual environment
python3 -m venv ernie-lora-env
source ernie-lora-env/bin/activate
Install dependencies
pip install torch torchvision --index-url https://download.pytorch.org/whl/cu121
pip install diffusers transformers accelerate peft bitsandbytes
pip install datasets pillow timm
Clone ERNIE-Image repo
git clone https://github.com/baidu/ERNIE-Image.git
cd ERNIE-Image
Method 2: DreamBooth Training Script
ERNIE-Image is built on the Diffusers architecture, so you can directly use HuggingFace's official train_dreambooth.py script:
git clone https://github.com/huggingface/diffusers.git
cd examples/dreambooth
pip install -r requirements.txt
Step 3: Run LoRA Fine-Tuning
Training Command
accelerate launch train_dreambooth_lora.py \
--pretrained_model_name_or_path=baidu/ERNIE-Image \
--instance_data_dir=./my-style-dataset/images \
--instance_prompt="a {style_keyword} of subject" \
--output_dir=./lora-outputs \
--learning_rate=1e-4 \
--max_train_steps=1000 \
--train_batch_size=1 \
--gradient_accumulation_steps=4 \
--resolution=1024 \
--mixed_precision=bf16 \
--lora_rank=32 \
--lora_alpha=16 \
--lora_target_modules=["to_q","to_k","to_v","to_out.0"] \
--checkpointing_steps=200 \
--seed=42
Key Parameters
| Parameter | Recommended | Notes |
|---|---|---|
learning_rate |
1e-4 ~ 5e-4 | Too high = overfitting, too low = undertrained |
max_train_steps |
500~2000 | Fewer images = fewer steps, more images = more steps |
lora_rank |
16~64 | Higher = more expressive but larger file |
lora_alpha |
lora_rank/2 | Typically half of rank |
mixed_precision |
bf16 | Needs Ampere+ GPU; older GPUs use fp16 |
Training Time Reference
| Images | Steps | Learning Rate | RTX 3090 | RTX 4090 |
|---|---|---|---|---|
| 15 | 500 | 3e-4 | ~15 min | ~10 min |
| 30 | 1000 | 2e-4 | ~30 min | ~20 min |
| 50 | 2000 | 1e-4 | ~60 min | ~40 min |
Step 4: Validate Fine-Tuning Results
Quick Test
After training, find pytorch_lora_weights.safetensors in your lora-outputs directory. Test with:
from diffusers import DiffusionPipeline
Load base model
pipe = DiffusionPipeline.from_pretrained(
"baidu/ERNIE-Image",
torch_dtype=torch.float16
)
pipe.to("cuda")
Load LoRA
pipe.load_lora_weights("./lora-outputs", weight_name="pytorch_lora_weights.safetensors")
Generate test image
prompt = "a watercolor painting of a mountain landscape, soft colors, flowing brush strokes"
image = pipe(prompt, num_inference_steps=50, guidance_scale=7.5).images[0]
image.save("test_output.png")
Evaluation Checklist
- Style consistency: Do different subject prompts maintain the unified style?
- Overfitting: Is the model just copying training images instead of learning the style?
- Prompt responsiveness: When changing the subject in the prompt, does the style follow?
Step 5: Use LoRA in ComfyUI
Installation Steps
- Copy
pytorch_lora_weights.safetensorstoComfyUI/models/loras/ - Launch ComfyUI
- Add a "Load LoRA" node to your workflow
- Select your LoRA model file
- Set LoRA strength (recommended: 0.6~0.9)
ComfyUI Workflow Example
[Load Checkpoint: ERNIE-Image]
↓
[Load LoRA: your-style-lora.safetensors, strength=0.8]
↓
[CLIP Text Encode (Positive Prompt)]
↓
[KSampler] → [VAE Decode] → [Save Image]
Tip: A LoRA strength of 0.8 typically gives the best results. Too high (>1.0) may over-stylize, too low (<0.5) has minimal effect.
Common Issues and Solutions
Issue 1: CUDA Out of Memory (OOM)
RuntimeError: CUDA out of memory
Solutions:
- Lower
resolutionto 768 or 512 - Enable gradient checkpointing:
--gradient_checkpointing - Reduce
train_batch_sizeor increasegradient_accumulation_steps - Use 8-bit optimizer:
--optimizer_type=8bit_adam
Issue 2: Output too similar to training images
This indicates overfitting.
Solutions:
- Lower learning rate (from 3e-4 to 1e-4)
- Reduce training steps
- Add more variation in instance prompts
- Use regularization images
Issue 3: Style not prominent enough
Solutions:
- Increase
lora_rankto 64 - Increase training steps
- Ensure training data style is distinct and unified
- Raise LoRA strength to 0.9~1.0 in ComfyUI
Issue 4: Stacking Multiple LoRAs
ERNIE-Image supports loading multiple LoRAs simultaneously for style + character combinations:
pipe.load_lora_weights("./style-lora", weight_name="pytorch_lora_weights.safetensors", adapter_name="style")
pipe.load_lora_weights("./character-lora", weight_name="pytorch_lora_weights.safetensors", adapter_name="character")
pipe.set_adapters(["style", "character"], adapter_weights=[0.8, 0.7])
Summary
ERNIE-Image's LoRA style fine-tuning is a zero-cost personalization capability. With 1050 style-consistent images and 1560 minutes of training on a consumer GPU, you can own a custom artistic style model.
Workflow Recap:
- Prepare 10~50 style-consistent images + captions
- Set up Diffusers + PEFT training environment
- Run LoRA fine-tuning (lr=1e-4, steps=1000, rank=32)
- Validate style results
- Load and use in ComfyUI
Compared to Midjourney's --sref (style reference) or FLUX's reference image features, ERNIE-Image LoRA offers more precise, controllable style transfer — and it's completely free, running locally.
Originally published on ernie-image.app. Please credit when sharing.