ERNIE-Image vs Recraft V3: Text Rendering Showdown — 8B Open-Source Challenger Takes on the Closed-Source Champion
Abstract: Recraft V3 tops the HuggingFace Text-to-Image Leaderboard with an ELO of 1172, claiming to be "the only model in the world that can generate images with long texts." ERNIE-Image 8B holds LongTextBench 0.9733, leading open-source text rendering. These two models represent two approaches to text-in-image generation: closed-source API precision vs open-source freedom. This article compares them across architecture, accuracy, cost, and real-world use cases.
Background: Text Rendering — The Last Mountain of AI Image Generation
From Stable Diffusion to FLUX, almost every open-source image generation model faces the same bottleneck: the model can generate beautiful images, but as soon as you ask it to render readable text inside the image, quality collapses. Letters turn into gibberish, CJK characters become blocks, and even simple logos fail to render correctly.
In late 2024, Recraft V3 emerged, claiming the #1 spot on HuggingFace's Text-to-Image Leaderboard with an ELO of 1172, declaring itself "the only model in the world that can generate images with long texts." In April 2026, Baidu open-sourced ERNIE-Image 8B, scoring 0.9733 on LongTextBench and establishing a leading position in open-source text rendering.
Two models, two paths: closed-source API precision vs open-source deployment freedom. Which one truly solves the "holy grail problem" of text rendering in AI image generation?
Architecture Comparison: Two Different Philosophies
Recraft V3: ControlNet-Style Text Layout Conditioning
Recraft V3's text rendering pipeline is a multi-stage system:
- Custom OCR Model: Trained based on the paper "Bridging the Gap Between End-to-End and Two-Step Text Spotting," addressing the failure of open-source OCR models under distribution mismatch
- Conditioned Captioning Model: Standard image captions rarely mention text. Recraft trained a specialized image captioning model that generates descriptions including text layout information
- LLM Text Layout Generator: Transforms text position, size, and font information into layout maps
- ControlNet-Style Text Conditioning: Feeds the text layout map as an additional condition into the image generation model, guiding precise typography placement
According to their technical blog: "JSON output was 10x slower than the final format" — Recraft optimized the LLM's JSON output into a lightweight custom format, significantly improving inference speed.
ERNIE-Image: Native Text Rendering via Single-Stream DiT
ERNIE-Image takes a completely different route: instead of relying on extra text layout conditioning, it improves training data and strategies so the DiT natively possesses text rendering capability.
Key technical points:
- Qwen3 VLM Auto-Captioning: Uses Qwen3 vision-language model to extract structural descriptions from training data, including text content and layout within images
- Aesthetic Alignment Training (DPO + Flow Matching): Integrates aesthetic evaluation through Direct Preference Optimization
- Single-Stream Architecture: 8B DiT parameters, using FLUX.2 VAE, without needing extra ControlNet or text layout conditions
Key quote from the ERNIE-Image technical report:
"ERNIE-Image aims to realize our team's ultimate goal: to build a model that is open-source, powerful, and easy for everyone to use. We adopt the FLUX.2 VAE, providing a strong open-source latent space for high-fidelity image generation."
Architectural Philosophy Comparison
| Dimension | Recraft V3 | ERNIE-Image |
|---|---|---|
| Rendering Method | Extra text layout condition (ControlNet-style) | Native DiT training |
| Parameters | Undisclosed | 8B |
| Open Source | ❌ Closed | ✅ Apache 2.0 |
| Additional Models | OCR + Captioning + Layout LLM + Image Gen | Main model only (PE optional) |
| Inference Pipeline | Multi-stage (slower) | Single-stage (faster) |
Text Rendering Accuracy Comparison
Short Text (1-5 Words)
Recraft V3 is nearly perfect for short text rendering — its HuggingFace ELO 1172 score is primarily driven by extremely high accuracy on these tasks. ERNIE-Image also performs excellently, with LongTextBench 0.9733 proving its strong text recognition capability.
Based on community feedback and official benchmarks:
- Recraft V3: Short text accuracy approaching 99%, with precise font, size, and position control
- ERNIE-Image: Short text accuracy approximately 90-95%, with bilingual Chinese-English text rendering support
Long Text (10+ Words/Sentences)
This is Recraft V3's core selling point. Their official claim: "the only model in the world that can generate images with long texts." Recraft V3 can render entire paragraphs of text in a single image while maintaining readability and layout consistency.
ERNIE-Image also performs reasonably well with long text, relying on the powerful instruction-following capability of its 8B parameters. The LongTextBench 0.9733 score covers various lengths from phrases to sentences.
Key difference: Recraft V3's text rendering is "precise control" (users specify position and size), while ERNIE-Image's is "auto-layout" (the model decides text placement).
Multilingual Text Support
| Language | Recraft V3 | ERNIE-Image |
|---|---|---|
| English | ✅ Excellent | ✅ Excellent |
| Chinese | ⚠️ Limited | ✅ Native support |
| Japanese | ⚠️ Limited | ✅ Supported |
| Spanish | ✅ Supported | ✅ Supported |
| Arabic | ❌ Untested | ⚠️ Limited |
ERNIE-Image's text rendering has significant advantages with CJK (Chinese/Japanese/Korean) characters — a capability most English-trained image generation models cannot match.
Cost and Accessibility Comparison
Recraft V3
- Free tier: 50 credits/day
- Basic plan: €10/month (1,000 credits)
- API access: RESTful API available
- Deployment: Web and API only, no local deployment option
- Vector graphics: ✅ Native support (SVG output)
ERNIE-Image
- Local deployment: Completely free, Apache 2.0 license
- GPU requirements: BF16 ~16GB+ VRAM, FP8 ~8GB+, GGUF ~12GB+
- Free cloud: Google Colab free T4 GPU works
- Third-party APIs: SiliconFlow (~¥0.11/image), WaveSpeed, Civitai, etc.
- Commercial use: ✅ Apache 2.0, unrestricted
Cost Comparison (Generating 1,000 Text-Inclusive Images)
| Option | Estimated Cost |
|---|---|
| Recraft V3 Basic Plan | €10-30 (depends on complexity) |
| ERNIE-Image Local GPU | ¥0 (existing GPU) + electricity |
| ERNIE-Image SiliconFlow | ≈ ¥110 |
| ERNIE-Image Colab Free | ¥0 (T4 GPU) |
Conclusion: For heavy users, ERNIE-Image's cost advantage is significant. For occasional designers, Recraft V3's free tier might suffice.
Real-World Use Case Comparison
Use Case 1: Poster Design
- Recraft V3 advantage: Precise text positioning and sizing, ideal for professional poster design
- ERNIE-Image advantage: Native Chinese poster support, no prompt translation needed
Use Case 2: E-commerce Product Photos
- Recraft V3 advantage: Precise brand text rendering on product packaging
- ERNIE-Image advantage: Local deployment protects business privacy, extremely low batch generation cost
Use Case 3: Social Media Content
- Recraft V3 advantage: Vector output, directly usable in design software
- ERNIE-Image advantage: Direct Chinese social media (Xiaohongshu, Weibo) content generation
Use Case 4: Academic Research / Education
- Recraft V3 disadvantage: Closed source, cannot reproduce or improve
- ERNIE-Image advantage: Open source, usable for research, education, fine-tuning
Recommendation Matrix
| Your Need | Recommendation |
|---|---|
| Precise text positioning and layout control | Recraft V3 |
| Chinese text rendering | ERNIE-Image |
| Budget-conscious or high-volume use | ERNIE-Image |
| Local deployment (data privacy/offline) | ERNIE-Image |
| Vector graphic output | Recraft V3 |
| Academic research or secondary development | ERNIE-Image |
| Occasional use, want plug-and-play | Recraft V3 |
| Open, auditable model | ERNIE-Image |
Conclusion
Recraft V3 and ERNIE-Image represent two extremes of text rendering: ultimate precision vs open freedom.
Recraft V3's advantage lies in its professional design tool positioning — precise text positioning, vector output, one-stop editing capabilities. For professional designers, it's a "plug-and-play" text rendering solution.
ERNIE-Image's advantage is its open-source nature and cost-effectiveness. An 8B parameter compact model, Apache 2.0 licensed, supporting local deployment, free Colab operation, and third-party APIs as low as ¥0.11/image. More importantly, it natively supports CJK text rendering — something most English-trained models can't do.
If you need a "tool" to quickly generate text-inclusive images, Recraft V3 is currently the best choice.
If you need a "platform" to build your own text rendering workflow, ERNIE-Image is currently the best choice.
In the 2026 AI image generation landscape, these two models are not competitors but complementary tools — serving different user groups and use cases.
Sources: Recraft V3 technical blog, ERNIE-Image technical report (arXiv 2605.25347), HuggingFace Leaderboard, community testing feedback.