fal.ai ERNIE-Image LoRA Cloud Training API: Zero-GPU Fine-Tuning at $1.2 per Run
Summary: fal.ai has launched the official ERNIE-Image LoRA cloud training API (
fal-ai/ernie-image-trainer) at $1.2 per 1,000 training steps. No local GPU required, no environment setup — upload a ZIP archive and train your custom LoRA. This article covers the complete API usage, dataset preparation, parameter tuning, output file handling, and a thorough comparison with local training, enabling you to fine-tune ERNIE-Image styles and characters without owning a GPU.
1. Why Cloud LoRA Training?
ERNIE-Image's LoRA fine-tuning capability has been extensively validated by the community (EI-071, EI-055), but local training faces three major barriers:
- Hardware requirement: 8B-parameter DiT model LoRA training needs 24GB+ VRAM (RTX 3090/4090 or A100)
- Environment setup: Aligning Diffusers, Accelerate, PEFT library versions is complex
- Training time: 2,000 steps typically take 30-60 minutes locally
fal.ai's cloud API solves all three: no GPU needed, no setup, results in minutes, with commercial use allowed.
2. fal.ai ERNIE-Image Trainer API Overview
Endpoint: fal-ai/ernie-image-trainer
| Feature | Details |
|---|---|
| Base Model | ERNIE-Image 8B (DiT) |
| Training Type | LoRA (Diffusers format) |
| Pricing | $1.2 / 1,000 steps |
| License | ✅ Apache 2.0 / Commercial use OK |
| Recommended Steps | 1,000-3,000 |
| Output Format | .safetensors LoRA + config file |
Cost Breakdown
| Training Steps | Cost | Use Case |
|---|---|---|
| 1,000 steps | $1.20 | Style exploration / quick testing |
| 2,000 steps | $2.40 | Standard character/style LoRA (recommended) |
| 3,000 steps | $3.60 | High-quality / complex concept LoRA |
3. Dataset Preparation: ZIP Archive Format
Directory Structure
training_data.zip
├── 001.jpg
├── 001.txt ← "a portrait photo of [trigger word] wearing a blue shirt"
├── 002.jpg
├── 002.txt ← "a portrait photo of [trigger word] smiling"
├── 003.jpg
├── 003.txt ← "a portrait photo of [trigger word] in a suit"
└── ...
Key Rules
- Images: Any common format (JPG/PNG), recommended 1024×1024 or higher
- Captions:
.txtfile with the same name as the image - Trigger word: Use a unique identifier (e.g.,
spfg person) to avoid conflicts with common words - Image count: Recommend 15-30 images, covering different angles, lighting, and backgrounds
Caption Writing Tips
# ✅ Recommended: descriptive + trigger word
"a professional portrait of spfg person, studio lighting, sharp focus, 8K detail"
❌ Avoid: too brief
"person"
✅ Character training: include identity/style description
"a digital painting of spfg warrior, fantasy armor, dynamic pose, dramatic lighting"
Tip: If some images lack
.txtfiles, set thedefault_captionparameter as a fallback.
4. Complete API Usage Workflow
4.1 Install Client
npm install --save @fal-ai/client
4.2 Configure API Key
export FAL_KEY="your_fal_api_key_here"
4.3 JavaScript Training Example
import { fal } from "@fal-ai/client";
// 1. Upload training data ZIP
const zipFile = new File([/* your ZIP data */], "training.zip", {
type: "application/zip",
});
const zipUrl = await fal.storage.upload(zipFile);
// 2. Submit training task (using queue protocol)
const { request_id } = await fal.queue.submit(
"fal-ai/ernie-image-trainer",
{
input: {
images_data_url: zipUrl,
steps: 2000, // 2000 steps = $2.40
learning_rate: 0.0005, // default learning rate
default_caption: null, // no default caption
},
webhookUrl: "https://your-server.com/training-complete", // optional: callback
}
);
console.log("Training submitted, Request ID:", request_id);
// 3. Poll status (if not using webhook)
const checkStatus = async () => {
const status = await fal.queue.status("fal-ai/ernie-image-trainer", {
requestId: request_id,
logs: true,
});
console.log("Current status:", status.status);
return status.status === "COMPLETED";
};
4.4 Python Training Example
import fal_client
Upload ZIP file
with open("training_data.zip", "rb") as f:
zip_url = fal_client.upload_file(f)
Submit training task
request = fal_client.queue.submit(
"fal-ai/ernie-image-trainer",
arguments={
"images_data_url": zip_url,
"steps": 2000,
"learning_rate": 0.0005,
}
)
print(f"Training submitted: {request['request_id']}")
Wait for completion
result = fal_client.queue.result(
"fal-ai/ernie-image-trainer",
request_id=request["request_id"]
)
print("LoRA file:", result["diffusers_lora_file"]["url"])
print("Config file:", result["config_file"]["url"])
5. Training Parameters Deep Dive
5.1 Learning Rate
| Learning Rate | Use Case | Risk |
|---|---|---|
| 0.0001 | Conservative training / large concepts | May underfit |
| 0.0005 | Default / general purpose | Balanced choice |
| 0.001 | Fast convergence / strong style | May overfit |
| 0.002+ | Special cases | High overfit risk |
5.2 Training Steps
| Steps | Cost | Effect |
|---|---|---|
| 500-1,000 | $0.60-$1.20 | Initial learning, weak concept |
| 1,500-2,000 | $1.80-$2.40 | Standard LoRA, recommended |
| 2,500-3,000 | $3.00-$3.60 | Deep memory, watch for overfit |
| 3,000+ | $3.60+ | Usually unnecessary |
Rule of thumb: Start with 2,000 steps + 0.0005 learning rate, then adjust based on output quality.
6. Post-Training: Using the LoRA Files
6.1 Download Output Files
After training, the API returns two files:
diffusers_lora_file: LoRA weights (.safetensorsformat)config_file: Model configuration
6.2 Using in Diffusers
from diffusers import ErnieImagePipeline
pipeline = ErnieImagePipeline.from_pretrained("baidu/ERNIE-Image")
Load LoRA
pipeline.load_lora_weights(
"path/to/diffusers_lora_file.safetensors",
adapter_name="my_custom_style"
)
Generate image
image = pipeline(
prompt="a spfg portrait, studio lighting",
lora_scale=0.8, # LoRA strength
num_inference_steps=50,
).images[0]
image.save("output.jpg")
6.3 Using in ComfyUI
- Download
diffusers_lora_filetoComfyUI/models/loras/ - Add a "Load LoRA" node in ComfyUI
- Select the LoRA file, set strength (0.6-1.0)
- Connect to the ERNIE-Image checkpoint node
7. fal.ai vs Local Training: Complete Comparison
| Dimension | fal.ai Cloud API | Local Training |
|---|---|---|
| GPU Required | ❌ None | ✅ 24GB+ VRAM |
| Initial Cost | $0 (pay-per-use) | $1,000+ (GPU) |
| Per Training | $1.20-$3.60 | ~$0.10 electricity |
| Training Time | 3-10 minutes | 30-60 minutes |
| Setup | Zero config | Complex (Diffusers/PEFT/Accelerate) |
| Parameter Flexibility | Limited (4 params) | Full control |
| Commercial License | ✅ | ✅ |
| Best For | Beginners / occasional training | Pro users / high-frequency |
When to choose fal.ai?
- No GPU or insufficient VRAM
- Quick LoRA concept testing
- Occasional training (< 10 runs/month)
- Don't want to handle environment setup
When to choose local training?
- High-frequency training (> 10 runs/month, lower cost)
- Need full control over training parameters
- High dataset privacy requirements
- Already have GPU resources
8. Practical Case: Training a Character Consistency LoRA
Background
Train a LoRA for a specific character, used for comic character consistency.
Step 1: Prepare Dataset
- Collect 20 character images (different expressions/angles/clothing)
- Caption each:
"a comic illustration of spfg character, [scene description]" - Package as ZIP
Step 2: Submit Training
const { request_id } = await fal.queue.submit(
"fal-ai/ernie-image-trainer",
{
input: {
images_data_url: zipUrl,
steps: 2000,
learning_rate: 0.0005,
},
}
);
Step 3: Validate Results
- Generate images with different prompts, check character consistency
- If character features aren't prominent enough → increase steps or adjust learning rate
- If overfitted (background/clothing memorized) → reduce steps
9. FAQ
Q: How long does training take?
A: 2,000 steps typically takes 3-10 minutes, depending on fal.ai server load.
Q: Can I train multi-concept LoRAs?
A: Yes, but use different trigger words to distinguish concepts.
Q: Can I use the output LoRA on other platforms?
A: Output is standard Diffusers LoRA format, compatible with ComfyUI, Diffusers API, etc.
Q: What if training fails?
A: Check ZIP format, ensure images have valid captions. fal.ai returns error logs.
Q: How does this differ from Z-Image Turbo Trainer?
A: ERNIE-Image Trainer is based on Baidu's 8B DiT ($1.2/1K steps), Z-Image Turbo Trainer on Alibaba's 6B ($0.85/1K steps). ERNIE-Image excels in text rendering and instruction following.
10. Conclusion
fal.ai's ERNIE-Image LoRA cloud training API opens the door for zero-GPU users to fine-tune ERNIE-Image. Starting at $1.2, training in minutes, zero config — style and character LoRA training has never been more accessible.
For occasional trainers, this is the most cost-effective choice; for high-frequency users, combine with local training (EI-071) and cloud training based on your needs.
Next step: Try your first LoRA training on fal.ai/ernie-image-trainer!
This article is based on fal.ai official API documentation and the ERNIE-Image technical report. API pricing and parameters may change — refer to the fal.ai official docs for the latest.