ERNIE-Image × Hunyuan3D 2.0: Complete Workflow from Text-to-Image to Textured 3D Models
Combine Baidu's ERNIE-Image with Tencent's Hunyuan3D 2.0 to build an end-to-end open-source pipeline: from text descriptions to high-quality textured 3D models. ERNIE-Image generates high-fidelity condition images, Hunyuan3D handles geometry generation and texture synthesis — both models are open-source under Apache 2.0 and can be orchestrated in ComfyUI.
Why a 2D → 3D Workflow?
AI image generation in 2026 is already mature. ERNIE-Image achieves SOTA-level text rendering and complex instruction following with just 8B parameters. But for game development, e-commerce 3D showcases, and VR/AR content creation, static 2D images are no longer enough — you need rotatable, interactive 3D assets.
Traditional 3D modeling requires professional Maya/Blender skills with a steep learning curve. Hunyuan3D 2.0 changes the game: it can automatically generate textured 3D mesh models from a single condition image. And ERNIE-Image happens to be one of the best tools for generating those high-quality condition images.
The combination advantage:
- ERNIE-Image: Precise text rendering + structured layout → condition images with clear text/logos
- Hunyuan3D-DiT: 2.6B parameter Flow Diffusion Transformer → accurate geometry from condition images
- Hunyuan3D-Paint: Geometric priors + diffusion → high-resolution texture maps
Hunyuan3D 2.0 Architecture Overview
Hunyuan3D 2.0 consists of two core components:
Hunyuan3D-DiT (Shape Generation)
Built on a Flow-based Diffusion Transformer with 2.6B parameters. Its core task is to extract 3D geometric information from condition images (or multi-view images) and generate untextured mesh models.
Key features:
- Supports single-image input (single-view reconstruction) and multi-view input
- Generated geometry aligns closely with condition image contours and structure
- Output formats:
.objor.glb
Hunyuan3D-Paint (Texture Synthesis)
The texture synthesis model leverages geometric priors and diffusion capabilities to generate high-resolution, vibrant texture maps for either generated or hand-crafted mesh models.
Key features:
- Texture resolution up to 2048×2048
- Automatic UV unwrapping
- Output texture maps directly usable in game engines (Unity/Unreal Engine)
Complete 5-Step Workflow: Text to 3D
Step 1: Generate Condition Image with ERNIE-Image
This is the most critical step in the entire pipeline. The quality of the condition image directly determines the geometric accuracy of the 3D model.
Best practices:
- Use ERNIE-Image (50-step SFT version) rather than Turbo — final output needs maximum quality
- Guidance scale:
guidance_scale=4.0 - Resolution:
1024×1024 - Enable PE (Prompt Enhancer) for richer details
Prompt example:
A detailed 3D-rendered product of a ceramic coffee mug, white background,
front view, clean lighting, no shadows, product photography style,
high resolution
Why ERNIE-Image is particularly well-suited?
ERNIE-Image's text rendering capability means you can directly embed brand logos or product names into condition images, and these texts will appear clearly in the 3D texture — something other open-source models struggle with.
Step 2: Background Removal (Optional but Recommended)
Hunyuan3D performs better with solid-color background condition images. Use ComfyUI's RemBG node or BiRefNet for background removal.
# BiRefNet quick background removal
from birefnet import BiRefNet
model = BiRefNet.from_pretrained("bi-refnet")
transparent_image = model.remove_background(ernie_output)
Step 3: Hunyuan3D-DiT Geometry Generation
Feed the cleaned condition image into Hunyuan3D-DiT:
from hunyuan3d import Hunyuan3DDiT
Initialize model
shape_model = Hunyuan3DDiT.from_pretrained("Tencent/Hunyuan3D-2.0-DiT")
Generate geometry from condition image
mesh = shape_model.generate(
condition_image=cleaned_condition,
num_steps=50,
guidance_scale=7.5
)
Export OBJ/GLB
mesh.export("output/model.obj")
mesh.export_glb("output/model.glb")
Step 4: Hunyuan3D-Paint Texture Synthesis
from hunyuan3d import Hunyuan3DPaint
Initialize texture model
paint_model = Hunyuan3DPaint.from_pretrained("Tencent/Hunyuan3D-2.0-Paint")
Generate texture maps
textured_mesh = paint_model.generate(
mesh=mesh,
reference_image=cleaned_condition,
texture_resolution=(2048, 2048)
)
Export textured model
textured_mesh.export_glb("output/textured_model.glb")
Step 5: Orchestrate in ComfyUI
ComfyUI provides a visual node-based workflow that chains the entire pipeline:
[ERNIE-Image Node] → [RemBG Node] → [Hunyuan3D-DiT Node] → [Hunyuan3D-Paint Node] → [GLB Export Node]
ComfyUI node installation:
# Update ComfyUI to latest version (native Hunyuan3D support)
cd ComfyUI
git pull
Or use ComfyUI-Hunyuan3DWrapper plugin
cd ComfyUI/custom_nodes
git clone https://github.com/kijai/ComfyUI-Hunyuan3DWrapper.git
cd ComfyUI-Hunyuan3DWrapper
pip install -r requirements.txt
Hardware Requirements & Performance
VRAM Requirements
| Step | Model | BF16 VRAM | FP8 VRAM | GGUF VRAM |
|---|---|---|---|---|
| Condition Image | ERNIE-Image | ~16 GB | ~10 GB | ~8 GB |
| Geometry | Hunyuan3D-DiT | ~20 GB | ~12 GB | Not supported |
| Texture | Hunyuan3D-Paint | ~14 GB | ~8 GB | Not supported |
Recommended configurations:
- Minimum: RTX 3090 (24GB) — FP8 quantization, step-by-step execution
- Recommended: RTX 4090 (24GB) — Full pipeline in FP8
- Ideal: A100 80GB — All models loaded simultaneously
Generation Time Estimates
On RTX 4090:
| Step | Time (BF16) | Time (FP8) |
|---|---|---|
| ERNIE-Image Turbo (8 steps) | ~5 sec | ~3 sec |
| ERNIE-Image SFT (50 steps) | ~25 sec | ~15 sec |
| Hunyuan3D-DiT Geometry | ~30 sec | ~18 sec |
| Hunyuan3D-Paint Texture | ~40 sec | ~25 sec |
| Total | ~100 sec | ~61 sec |
With ERNIE-Image Turbo + FP8 quantization, the total time drops to under 1 minute.
Practical Use Cases
E-commerce Product Showcase
- Generate product front-view with ERNIE-Image (including brand text)
- Hunyuan3D generates 360° rotatable 3D model
- Export GLB for Shopify/Taobao 3D product pages
Advantage: No physical photography needed — hundreds of 3D product showcases per day.
Game Asset Rapid Prototyping
- Design phase: ERNIE-Image generates character/prop concept art
- Hunyuan3D quickly generates engine-importable 3D models
- Art team refines on top of generated base
Advantage: Reduces concept-to-3D-prototype cycle from days to minutes.
Educational 3D Content
- ERNIE-Image generates textbook illustrations (with text annotations)
- Hunyuan3D generates 3D teaching models
- Embed into VR/AR educational applications
FAQ
Q: Text in 3D models is blurry. What to do?
ERNIE-Image's text rendering is strong in 2D, but Hunyuan3D's texture synthesis may blur text details. Suggestions:
- Use larger font sizes in the ERNIE-Image prompt
- Set texture resolution to 2048×2048 instead of default 1024×1024
- For critical text, consider adding it in a 3D modeling tool post-generation
Q: Single-view vs Multi-view — which is better?
- Single-view: Faster, suitable for simple objects (products, vessels)
- Multi-view: Generate 4-6 angle views with ERNIE-Image; geometric accuracy improves significantly for complex objects (characters, machinery)
Q: Can I run this on a 16GB GPU?
Yes, but you'll need:
- ERNIE-Image Turbo NVFP4 quantization (~4.8GB)
- Execute geometry and texture steps sequentially, freeing VRAM between
- CPU offloading (speed will drop noticeably)
Summary
The ERNIE-Image × Hunyuan3D 2.0 combination represents the best practice for open-source 3D content generation in 2026:
- Fully open-source: Apache 2.0 license, worry-free commercial use
- Fully automated: Text descriptions to 3D assets, end-to-end
- High quality: ERNIE-Image text rendering + Hunyuan3D geometric precision
- Deployable: ComfyUI visual orchestration, SGLang production deployment
As both models continue to iterate, this workflow's quality will only improve. For teams needing bulk 3D content, this is a tech stack worth investing in.