ERNIE-Image vs Boogu-Image-0.1: 8B Precision vs 10B All-Rounder — The June 2026 Open-Source Text-to-Image Face-Off

Jun 27, 2026

ERNIE-Image vs Boogu-Image-0.1: 8B Precision vs 10B All-Rounder — The June 2026 Open-Source Text-to-Image Face-Off

Summary: Baidu's ERNIE-Image (8B DiT) has been a top-tier open-source text-to-image model since its April 2026 release. In late June 2026, the Boogu team launched Boogu-Image-0.1 — a 10B-parameter unified image generation and editing model family. This article provides a head-to-head comparison across five dimensions: architecture, performance, editing capabilities, deployment costs, and real-world usability.


A New Variable in the Open-Source Text-to-Image Landscape

The open-source image generation space is undergoing rapid iteration. From FLUX.2 to ERNIE-Image, to HunyuanImage 3.0, each new model release is redefining what "open-source deployable" means.

In June 2026, a new competitor entered the arena — Boogu-Image-0.1.

The Boogu team released three variants simultaneously on HuggingFace: Base (high-quality text-to-image), Turbo (4-step ultra-fast generation), and Edit (image editing). The Edit variant is particularly notable — in the current open-source ecosystem, truly usable image editing models are scarce, and ERNIE-Image's official editing model remains in the "coming soon" state.

This article puts ERNIE-Image (8B) and Boogu-Image-0.1 (10B) side by side for a comprehensive comparison.


1. Architecture and Parameter Comparison

ERNIE-Image: 8B DiT + 3B PE

ERNIE-Image uses a single-stream Diffusion Transformer architecture with an 8B core diffusion model, paired with a 3B-parameter Prompt Enhancer (PE). The PE rewrites short user prompts into richer, structured descriptions before passing them to the diffusion model.

Two official variants:

  • ERNIE-Image (SFT): ~50 inference steps, emphasizing instruction fidelity and general-purpose capability
  • ERNIE-Image-Turbo: DMD + RL optimized, ~8 steps for high aesthetic quality

Key design: ERNIE-Image uses FLUX.2 VAE as its variational autoencoder, meaning it shares a latent space with the FLUX family — improving LoRA compatibility.

Boogu-Image-0.1: 10B Unified Architecture

Boogu-Image-0.1 uses a unified architecture where Base/Turbo/Edit share the same 10B parameter base, differentiated by training strategies:

  • Base: Joint training, 25~50 steps, focused on text-to-image quality
  • Turbo: Decoupled DMD distillation, only 4 inference steps
  • Edit: Joint training, 25~50 steps, supports text-instruction-to-image (TI2I) editing

Key design: Boogu's Edit variant natively supports image editing — including object addition/removal, attribute modification, background replacement, and style transfer. This is a capability not yet covered in the ERNIE-Image ecosystem.

Parameter Scale Differences

Dimension ERNIE-Image Boogu-Image-0.1
Diffusion Model 8B DiT 10B
Auxiliary Model 3B PE (optional) None
Total Parameters ~11B (with PE) 10B
Inference Steps 8~50 4~50
Minimum Steps 8 (Turbo) 4 (Turbo)
Image Editing ❌ No official Edit ✅ Edit variant
License Apache-2.0 Apache-2.0

2. Performance Benchmark Comparison

GENEval Instruction Following

ERNIE-Image performs strongly on the GENEval benchmark:

Model Single Object Two Object Counting Colors Position Attribute Binding Overall
ERNIE-Image (w/o PE) 1.0000 0.9596 0.7781 0.9282 0.8550 0.7925 0.8856
ERNIE-Image (w/ PE) 0.9906 0.9596 0.8187 0.8830 0.8625 0.7225 0.8728
ERNIE-Image-Turbo 1.0000 0.9621 0.7906 0.9202 0.7975 0.7300 0.8667

Boogu-Image-0.1 has not yet published official GENEval data, but scored 53.58 on the Qwen-Image-Bench, claiming it "outperforms models 2-8x its size."

Text Rendering

ERNIE-Image's core differentiating capability is precise text rendering. On LongTextBench:

  • ERNIE-Image (w/ PE): 0.9733 average (EN: 0.9804, ZH: 0.9661)

Boogu-Image-0.1 also claims bilingual (CN/EN) text rendering support, and its Edit variant provides fine-grained in-image text editing (add/remove/replace characters, adjust fonts/weights/colors/layouts). However, Boogu has not yet published LongTextBench comparison data.

Summary: ERNIE-Image has thoroughly validated benchmark data for text rendering, while Boogu's official data is incomplete. However, for text modification within image editing scenarios, Boogu's Edit variant is ahead.


3. Editing Capabilities: Boogu's Core Advantage

This is the most critical difference in this comparison.

Boogu-Image-0.1-Edit

Boogu's Edit variant natively supports Text-Instruction-to-Image (TI2I) editing tasks:

  • Object addition/removal: Add or remove specified objects in an image
  • Attribute modification: Change colors, materials, and styles of objects
  • Background replacement: Replace background scenes while preserving the subject
  • Style transfer: Convert images to different artistic styles

The official documentation specifically highlights the Edit model's texture preservation advantages — when handling complex textures like concrete, graffiti, and skin, Boogu-Edit preserves more detail than competitors (such as Qwen-Image-Edit).

Hardware requirements: The Edit variant needs 2550 inference steps, CFG recommended 2.05.0. On 12GB VRAM, CPU Offload is needed; on 24GB VRAM, it runs directly.

ERNIE-Image's Editing Status

ERNIE-Image currently has no officially released editing model. The community uses these workarounds:

  • Generic Inpainting workflows in ComfyUI (using masks)
  • FLUX Kontext hybrid pipeline (discussed in detail in EI-102)
  • Bernini-R + FLUX Kontext hybrid approach

These solutions are inconsistent in quality and require additional model weights and complex pipeline configurations.

Conclusion: If your core need is image editing, Boogu-Image-0.1-Edit is currently the most straightforward choice in the open-source ecosystem. The community has been waiting months for ERNIE-Image's official editing model.


4. Deployment Costs and Hardware Requirements

VRAM Requirements Comparison

Model BF16 VRAM FP8 VRAM 12GB Runnable 24GB Runnable
ERNIE-Image (8B) ~16GB ~8-9GB ✅ (FP8+Offload) ✅ Easy
ERNIE-Image-Turbo ~16GB ~8-9GB ✅ (FP8+Offload) ✅ Easy
Boogu Base (10B) ~20GB ~10-12GB ⚠️ (FP8+CPU Offload) ✅
Boogu Turbo (10B) ~20GB ~10-12GB ⚠️ (FP8+CPU Offload) ✅
Boogu Edit (10B) ~20GB ~10-12GB ⚠️ (FP8+CPU Offload) ✅

Key point: Boogu's 10B parameters are ~25% larger than ERNIE-Image's 8B, requiring ~4GB more VRAM in BF16. For 24GB cards (RTX 3090/4090), both run fine. For 12-16GB cards, Boogu requires more aggressive optimization.

Inference Speed

Model Steps Single 1024² time (RTX 4090 estimate)
ERNIE-Image-Turbo 8 ~15-20 seconds
Boogu Turbo 4 ~8-12 seconds
ERNIE-Image (SFT) 50 ~60-80 seconds
Boogu Base 25~50 ~30-80 seconds

Boogu Turbo's 4-step design gives it a clear advantage in speed-critical scenarios, though fewer steps typically mean reduced detail fidelity.

API Deployment Options

Platform ERNIE-Image Boogu-Image-0.1
FAL.AI ✅ Supported ❌ Not yet
Atlas Cloud ✅ Supported ❌ Not yet
WaveSpeed AI ✅ Supported ❌ Not yet
Civitai Orchestration ✅ Supported ❌ Not yet
Local Deployment ✅ Diffusers ✅ Diffusers
ComfyUI ✅ Official nodes ✅ Comfy-Org nodes

Key difference: ERNIE-Image is already integrated with multiple cloud API platforms, enabling zero-local-GPU deployment. Boogu currently only supports local deployment (the team explicitly states they provide no paid API).


5. Community Ecosystem Comparison

ERNIE-Image Community

As of June 2026:

  • HuggingFace: 2.67K+ followers, multiple variant models
  • Civitai: Dozens of community LoRAs and workflows
  • Reddit: Active discussion in r/StableDiffusion and r/ComfyUI
  • YouTube: Multiple ComfyUI tutorials and comparison videos
  • Quantized versions: Community-maintained NVFP4/FP8/INT8 versions

Boogu-Image-0.1 Community

As of June 2026:

  • HuggingFace: Recently released, growing followers
  • Reddit: Emerging discussion in r/LocalLLaMA and r/StableDiffusion
  • YouTube: Benji's AI Playground tutorial (5,593 views)
  • ComfyUI: Comfy-Org provides official model files
  • Quantized versions: Official fp8 variants available

Conclusion: ERNIE-Image's community ecosystem is clearly ahead. Boogu, as a new model, is rapidly building its community.


6. Practical Use Case Recommendations

Choose ERNIE-Image When:

  1. Precise text rendering needed: Posters, comics, infographics, multilingual layouts
  2. Commercial deployment: Need API integration (FAL/Atlas/WaveSpeed/Civitai)
  3. Community resource dependent: Need mature LoRA ecosystem and workflow templates
  4. High instruction-following requirements: Complex multi-constraint prompts
  5. Limited VRAM: 8B model is more friendly to 12-16GB GPUs

Choose Boogu-Image-0.1 When:

  1. Image editing needed: Native Edit variant for object addition/removal, style transfer
  2. Ultra-fast generation: Turbo variant with 4-step inference for rapid iteration
  3. Unified model preference: Base/Turbo/Edit share architecture, no model switching
  4. Texture preservation: Edit model excels with complex textures
  5. Fully local deployment: No cloud API needed, prioritizing data privacy

Scenarios Where Neither Is Ideal

  • Ultra-high-res output (4K+): Both require additional upscaling workflows
  • Character consistency: Both need additional LoRA or IP-Adapter
  • Video generation: Both focus on images, need pairing with video models (e.g., Wan 2.6)

7. Summary

Dimension ERNIE-Image (8B) Boogu-Image-0.1 (10B) Winner
Text Rendering ✅ 0.9733 LongTextBench ⚠️ Incomplete data ERNIE-Image
Image Editing ❌ No official Edit ✅ Edit variant Boogu
Fast Generation 8 steps (Turbo) 4 steps (Turbo) Boogu
Community Ecosystem ✅ Mature 🌱 Growing ERNIE-Image
API Deployment ✅ Multi-platform ❌ Local only ERNIE-Image
VRAM Friendly 8B 10B ERNIE-Image
License Apache-2.0 Apache-2.0 Tie
Texture Preservation Standard ✅ Edit variant excels Boogu

Final recommendation: If you need a ready-to-use text-to-image model with strong text rendering and a mature community ecosystem, ERNIE-Image remains the best choice for mid-2026. If you need image editing capabilities or ultra-fast generation, Boogu-Image-0.1 is worth trying. Both use Apache-2.0 licensing, meaning you can freely combine them — use ERNIE-Image for base image generation, then Boogu-Edit for post-processing. This is entirely feasible in ComfyUI.

Update reminder: ERNIE-Image's official editing model is still in development and will directly change this comparison landscape once released. Stay tuned to Baidu's official updates.


Data in this article is sourced from HuggingFace official model cards, GitHub repositories, and arXiv technical reports. Boogu-Image-0.1's official team explicitly states they provide no paid API or commercial services. Any paid product under the name "Boogu-Image" is unofficial.

ERNIE-Image Team