fal.ai ERNIE-Image LoRA Cloud Training API: Zero-GPU Fine-Tuning at $1.2 per Run

Jun 9, 2026

fal.ai ERNIE-Image LoRA Cloud Training API: Zero-GPU Fine-Tuning at $1.2 per Run

Summary: fal.ai has launched the official ERNIE-Image LoRA cloud training API (fal-ai/ernie-image-trainer) at $1.2 per 1,000 training steps. No local GPU required, no environment setup — upload a ZIP archive and train your custom LoRA. This article covers the complete API usage, dataset preparation, parameter tuning, output file handling, and a thorough comparison with local training, enabling you to fine-tune ERNIE-Image styles and characters without owning a GPU.


1. Why Cloud LoRA Training?

ERNIE-Image's LoRA fine-tuning capability has been extensively validated by the community (EI-071, EI-055), but local training faces three major barriers:

  1. Hardware requirement: 8B-parameter DiT model LoRA training needs 24GB+ VRAM (RTX 3090/4090 or A100)
  2. Environment setup: Aligning Diffusers, Accelerate, PEFT library versions is complex
  3. Training time: 2,000 steps typically take 30-60 minutes locally

fal.ai's cloud API solves all three: no GPU needed, no setup, results in minutes, with commercial use allowed.


2. fal.ai ERNIE-Image Trainer API Overview

Endpoint: fal-ai/ernie-image-trainer

Feature Details
Base Model ERNIE-Image 8B (DiT)
Training Type LoRA (Diffusers format)
Pricing $1.2 / 1,000 steps
License ✅ Apache 2.0 / Commercial use OK
Recommended Steps 1,000-3,000
Output Format .safetensors LoRA + config file

Cost Breakdown

Training Steps Cost Use Case
1,000 steps $1.20 Style exploration / quick testing
2,000 steps $2.40 Standard character/style LoRA (recommended)
3,000 steps $3.60 High-quality / complex concept LoRA

3. Dataset Preparation: ZIP Archive Format

Directory Structure

training_data.zip
├── 001.jpg
├── 001.txt    ← "a portrait photo of [trigger word] wearing a blue shirt"
├── 002.jpg
├── 002.txt    ← "a portrait photo of [trigger word] smiling"
├── 003.jpg
├── 003.txt    ← "a portrait photo of [trigger word] in a suit"
└── ...

Key Rules

  • Images: Any common format (JPG/PNG), recommended 1024×1024 or higher
  • Captions: .txt file with the same name as the image
  • Trigger word: Use a unique identifier (e.g., spfg person) to avoid conflicts with common words
  • Image count: Recommend 15-30 images, covering different angles, lighting, and backgrounds

Caption Writing Tips

# ✅ Recommended: descriptive + trigger word
"a professional portrait of spfg person, studio lighting, sharp focus, 8K detail"

❌ Avoid: too brief

"person"

✅ Character training: include identity/style description

"a digital painting of spfg warrior, fantasy armor, dynamic pose, dramatic lighting"

Tip: If some images lack .txt files, set the default_caption parameter as a fallback.


4. Complete API Usage Workflow

4.1 Install Client

npm install --save @fal-ai/client

4.2 Configure API Key

export FAL_KEY="your_fal_api_key_here"

4.3 JavaScript Training Example

import { fal } from "@fal-ai/client";

// 1. Upload training data ZIP
const zipFile = new File([/* your ZIP data */], "training.zip", {
type: "application/zip",
});
const zipUrl = await fal.storage.upload(zipFile);

// 2. Submit training task (using queue protocol)
const { request_id } = await fal.queue.submit(
"fal-ai/ernie-image-trainer",
{
input: {
images_data_url: zipUrl,
steps: 2000, // 2000 steps = $2.40
learning_rate: 0.0005, // default learning rate
default_caption: null, // no default caption
},
webhookUrl: "https://your-server.com/training-complete", // optional: callback
}
);

console.log("Training submitted, Request ID:", request_id);

// 3. Poll status (if not using webhook)
const checkStatus = async () => {
const status = await fal.queue.status("fal-ai/ernie-image-trainer", {
requestId: request_id,
logs: true,
});
console.log("Current status:", status.status);
return status.status === "COMPLETED";
};

4.4 Python Training Example

import fal_client

Upload ZIP file

with open("training_data.zip", "rb") as f:
zip_url = fal_client.upload_file(f)

Submit training task

request = fal_client.queue.submit(
"fal-ai/ernie-image-trainer",
arguments={
"images_data_url": zip_url,
"steps": 2000,
"learning_rate": 0.0005,
}
)

print(f"Training submitted: {request['request_id']}")

Wait for completion

result = fal_client.queue.result(
"fal-ai/ernie-image-trainer",
request_id=request["request_id"]
)

print("LoRA file:", result["diffusers_lora_file"]["url"])
print("Config file:", result["config_file"]["url"])


5. Training Parameters Deep Dive

5.1 Learning Rate

Learning Rate Use Case Risk
0.0001 Conservative training / large concepts May underfit
0.0005 Default / general purpose Balanced choice
0.001 Fast convergence / strong style May overfit
0.002+ Special cases High overfit risk

5.2 Training Steps

Steps Cost Effect
500-1,000 $0.60-$1.20 Initial learning, weak concept
1,500-2,000 $1.80-$2.40 Standard LoRA, recommended
2,500-3,000 $3.00-$3.60 Deep memory, watch for overfit
3,000+ $3.60+ Usually unnecessary

Rule of thumb: Start with 2,000 steps + 0.0005 learning rate, then adjust based on output quality.


6. Post-Training: Using the LoRA Files

6.1 Download Output Files

After training, the API returns two files:

  • diffusers_lora_file: LoRA weights (.safetensors format)
  • config_file: Model configuration

6.2 Using in Diffusers

from diffusers import ErnieImagePipeline

pipeline = ErnieImagePipeline.from_pretrained("baidu/ERNIE-Image")

Load LoRA

pipeline.load_lora_weights(
"path/to/diffusers_lora_file.safetensors",
adapter_name="my_custom_style"
)

Generate image

image = pipeline(
prompt="a spfg portrait, studio lighting",
lora_scale=0.8, # LoRA strength
num_inference_steps=50,
).images[0]
image.save("output.jpg")

6.3 Using in ComfyUI

  1. Download diffusers_lora_file to ComfyUI/models/loras/
  2. Add a "Load LoRA" node in ComfyUI
  3. Select the LoRA file, set strength (0.6-1.0)
  4. Connect to the ERNIE-Image checkpoint node

7. fal.ai vs Local Training: Complete Comparison

Dimension fal.ai Cloud API Local Training
GPU Required ❌ None ✅ 24GB+ VRAM
Initial Cost $0 (pay-per-use) $1,000+ (GPU)
Per Training $1.20-$3.60 ~$0.10 electricity
Training Time 3-10 minutes 30-60 minutes
Setup Zero config Complex (Diffusers/PEFT/Accelerate)
Parameter Flexibility Limited (4 params) Full control
Commercial License ✅ ✅
Best For Beginners / occasional training Pro users / high-frequency

When to choose fal.ai?

  • No GPU or insufficient VRAM
  • Quick LoRA concept testing
  • Occasional training (< 10 runs/month)
  • Don't want to handle environment setup

When to choose local training?

  • High-frequency training (> 10 runs/month, lower cost)
  • Need full control over training parameters
  • High dataset privacy requirements
  • Already have GPU resources

8. Practical Case: Training a Character Consistency LoRA

Background

Train a LoRA for a specific character, used for comic character consistency.

Step 1: Prepare Dataset

  • Collect 20 character images (different expressions/angles/clothing)
  • Caption each: "a comic illustration of spfg character, [scene description]"
  • Package as ZIP

Step 2: Submit Training

const { request_id } = await fal.queue.submit(
  "fal-ai/ernie-image-trainer",
  {
    input: {
      images_data_url: zipUrl,
      steps: 2000,
      learning_rate: 0.0005,
    },
  }
);

Step 3: Validate Results

  • Generate images with different prompts, check character consistency
  • If character features aren't prominent enough → increase steps or adjust learning rate
  • If overfitted (background/clothing memorized) → reduce steps

9. FAQ

Q: How long does training take?

A: 2,000 steps typically takes 3-10 minutes, depending on fal.ai server load.

Q: Can I train multi-concept LoRAs?

A: Yes, but use different trigger words to distinguish concepts.

Q: Can I use the output LoRA on other platforms?

A: Output is standard Diffusers LoRA format, compatible with ComfyUI, Diffusers API, etc.

Q: What if training fails?

A: Check ZIP format, ensure images have valid captions. fal.ai returns error logs.

Q: How does this differ from Z-Image Turbo Trainer?

A: ERNIE-Image Trainer is based on Baidu's 8B DiT ($1.2/1K steps), Z-Image Turbo Trainer on Alibaba's 6B ($0.85/1K steps). ERNIE-Image excels in text rendering and instruction following.


10. Conclusion

fal.ai's ERNIE-Image LoRA cloud training API opens the door for zero-GPU users to fine-tune ERNIE-Image. Starting at $1.2, training in minutes, zero config — style and character LoRA training has never been more accessible.

For occasional trainers, this is the most cost-effective choice; for high-frequency users, combine with local training (EI-071) and cloud training based on your needs.

Next step: Try your first LoRA training on fal.ai/ernie-image-trainer!


This article is based on fal.ai official API documentation and the ERNIE-Image technical report. API pricing and parameters may change — refer to the fal.ai official docs for the latest.

ERNIE-Image Team