ERNIE-Image vs Recraft V3: Text Rendering Showdown — 8B Open-Source Challenger Takes on the Closed-Source Champion

Jul 6, 2026

ERNIE-Image vs Recraft V3: Text Rendering Showdown — 8B Open-Source Challenger Takes on the Closed-Source Champion

Abstract: Recraft V3 tops the HuggingFace Text-to-Image Leaderboard with an ELO of 1172, claiming to be "the only model in the world that can generate images with long texts." ERNIE-Image 8B holds LongTextBench 0.9733, leading open-source text rendering. These two models represent two approaches to text-in-image generation: closed-source API precision vs open-source freedom. This article compares them across architecture, accuracy, cost, and real-world use cases.

Background: Text Rendering — The Last Mountain of AI Image Generation

From Stable Diffusion to FLUX, almost every open-source image generation model faces the same bottleneck: the model can generate beautiful images, but as soon as you ask it to render readable text inside the image, quality collapses. Letters turn into gibberish, CJK characters become blocks, and even simple logos fail to render correctly.

In late 2024, Recraft V3 emerged, claiming the #1 spot on HuggingFace's Text-to-Image Leaderboard with an ELO of 1172, declaring itself "the only model in the world that can generate images with long texts." In April 2026, Baidu open-sourced ERNIE-Image 8B, scoring 0.9733 on LongTextBench and establishing a leading position in open-source text rendering.

Two models, two paths: closed-source API precision vs open-source deployment freedom. Which one truly solves the "holy grail problem" of text rendering in AI image generation?

Architecture Comparison: Two Different Philosophies

Recraft V3: ControlNet-Style Text Layout Conditioning

Recraft V3's text rendering pipeline is a multi-stage system:

  1. Custom OCR Model: Trained based on the paper "Bridging the Gap Between End-to-End and Two-Step Text Spotting," addressing the failure of open-source OCR models under distribution mismatch
  2. Conditioned Captioning Model: Standard image captions rarely mention text. Recraft trained a specialized image captioning model that generates descriptions including text layout information
  3. LLM Text Layout Generator: Transforms text position, size, and font information into layout maps
  4. ControlNet-Style Text Conditioning: Feeds the text layout map as an additional condition into the image generation model, guiding precise typography placement

According to their technical blog: "JSON output was 10x slower than the final format" — Recraft optimized the LLM's JSON output into a lightweight custom format, significantly improving inference speed.

ERNIE-Image: Native Text Rendering via Single-Stream DiT

ERNIE-Image takes a completely different route: instead of relying on extra text layout conditioning, it improves training data and strategies so the DiT natively possesses text rendering capability.

Key technical points:

  • Qwen3 VLM Auto-Captioning: Uses Qwen3 vision-language model to extract structural descriptions from training data, including text content and layout within images
  • Aesthetic Alignment Training (DPO + Flow Matching): Integrates aesthetic evaluation through Direct Preference Optimization
  • Single-Stream Architecture: 8B DiT parameters, using FLUX.2 VAE, without needing extra ControlNet or text layout conditions

Key quote from the ERNIE-Image technical report:

"ERNIE-Image aims to realize our team's ultimate goal: to build a model that is open-source, powerful, and easy for everyone to use. We adopt the FLUX.2 VAE, providing a strong open-source latent space for high-fidelity image generation."

Architectural Philosophy Comparison

Dimension Recraft V3 ERNIE-Image
Rendering Method Extra text layout condition (ControlNet-style) Native DiT training
Parameters Undisclosed 8B
Open Source ❌ Closed ✅ Apache 2.0
Additional Models OCR + Captioning + Layout LLM + Image Gen Main model only (PE optional)
Inference Pipeline Multi-stage (slower) Single-stage (faster)

Text Rendering Accuracy Comparison

Short Text (1-5 Words)

Recraft V3 is nearly perfect for short text rendering — its HuggingFace ELO 1172 score is primarily driven by extremely high accuracy on these tasks. ERNIE-Image also performs excellently, with LongTextBench 0.9733 proving its strong text recognition capability.

Based on community feedback and official benchmarks:

  • Recraft V3: Short text accuracy approaching 99%, with precise font, size, and position control
  • ERNIE-Image: Short text accuracy approximately 90-95%, with bilingual Chinese-English text rendering support

Long Text (10+ Words/Sentences)

This is Recraft V3's core selling point. Their official claim: "the only model in the world that can generate images with long texts." Recraft V3 can render entire paragraphs of text in a single image while maintaining readability and layout consistency.

ERNIE-Image also performs reasonably well with long text, relying on the powerful instruction-following capability of its 8B parameters. The LongTextBench 0.9733 score covers various lengths from phrases to sentences.

Key difference: Recraft V3's text rendering is "precise control" (users specify position and size), while ERNIE-Image's is "auto-layout" (the model decides text placement).

Multilingual Text Support

Language Recraft V3 ERNIE-Image
English ✅ Excellent ✅ Excellent
Chinese ⚠️ Limited ✅ Native support
Japanese ⚠️ Limited ✅ Supported
Spanish ✅ Supported ✅ Supported
Arabic ❌ Untested ⚠️ Limited

ERNIE-Image's text rendering has significant advantages with CJK (Chinese/Japanese/Korean) characters — a capability most English-trained image generation models cannot match.

Cost and Accessibility Comparison

Recraft V3

  • Free tier: 50 credits/day
  • Basic plan: €10/month (1,000 credits)
  • API access: RESTful API available
  • Deployment: Web and API only, no local deployment option
  • Vector graphics: ✅ Native support (SVG output)

ERNIE-Image

  • Local deployment: Completely free, Apache 2.0 license
  • GPU requirements: BF16 ~16GB+ VRAM, FP8 ~8GB+, GGUF ~12GB+
  • Free cloud: Google Colab free T4 GPU works
  • Third-party APIs: SiliconFlow (~¥0.11/image), WaveSpeed, Civitai, etc.
  • Commercial use: ✅ Apache 2.0, unrestricted

Cost Comparison (Generating 1,000 Text-Inclusive Images)

Option Estimated Cost
Recraft V3 Basic Plan €10-30 (depends on complexity)
ERNIE-Image Local GPU ¥0 (existing GPU) + electricity
ERNIE-Image SiliconFlow ≈ ¥110
ERNIE-Image Colab Free ¥0 (T4 GPU)

Conclusion: For heavy users, ERNIE-Image's cost advantage is significant. For occasional designers, Recraft V3's free tier might suffice.

Real-World Use Case Comparison

Use Case 1: Poster Design

  • Recraft V3 advantage: Precise text positioning and sizing, ideal for professional poster design
  • ERNIE-Image advantage: Native Chinese poster support, no prompt translation needed

Use Case 2: E-commerce Product Photos

  • Recraft V3 advantage: Precise brand text rendering on product packaging
  • ERNIE-Image advantage: Local deployment protects business privacy, extremely low batch generation cost

Use Case 3: Social Media Content

  • Recraft V3 advantage: Vector output, directly usable in design software
  • ERNIE-Image advantage: Direct Chinese social media (Xiaohongshu, Weibo) content generation

Use Case 4: Academic Research / Education

  • Recraft V3 disadvantage: Closed source, cannot reproduce or improve
  • ERNIE-Image advantage: Open source, usable for research, education, fine-tuning

Recommendation Matrix

Your Need Recommendation
Precise text positioning and layout control Recraft V3
Chinese text rendering ERNIE-Image
Budget-conscious or high-volume use ERNIE-Image
Local deployment (data privacy/offline) ERNIE-Image
Vector graphic output Recraft V3
Academic research or secondary development ERNIE-Image
Occasional use, want plug-and-play Recraft V3
Open, auditable model ERNIE-Image

Conclusion

Recraft V3 and ERNIE-Image represent two extremes of text rendering: ultimate precision vs open freedom.

Recraft V3's advantage lies in its professional design tool positioning — precise text positioning, vector output, one-stop editing capabilities. For professional designers, it's a "plug-and-play" text rendering solution.

ERNIE-Image's advantage is its open-source nature and cost-effectiveness. An 8B parameter compact model, Apache 2.0 licensed, supporting local deployment, free Colab operation, and third-party APIs as low as ¥0.11/image. More importantly, it natively supports CJK text rendering — something most English-trained models can't do.

If you need a "tool" to quickly generate text-inclusive images, Recraft V3 is currently the best choice.

If you need a "platform" to build your own text rendering workflow, ERNIE-Image is currently the best choice.

In the 2026 AI image generation landscape, these two models are not competitors but complementary tools — serving different user groups and use cases.


Sources: Recraft V3 technical blog, ERNIE-Image technical report (arXiv 2605.25347), HuggingFace Leaderboard, community testing feedback.

ERNIE-Image Team