ERNIE-Image × Hunyuan3D 2.0: Complete Workflow from Text-to-Image to Textured 3D Models

Jun 2, 2026

ERNIE-Image × Hunyuan3D 2.0: Complete Workflow from Text-to-Image to Textured 3D Models

Combine Baidu's ERNIE-Image with Tencent's Hunyuan3D 2.0 to build an end-to-end open-source pipeline: from text descriptions to high-quality textured 3D models. ERNIE-Image generates high-fidelity condition images, Hunyuan3D handles geometry generation and texture synthesis — both models are open-source under Apache 2.0 and can be orchestrated in ComfyUI.

Why a 2D → 3D Workflow?

AI image generation in 2026 is already mature. ERNIE-Image achieves SOTA-level text rendering and complex instruction following with just 8B parameters. But for game development, e-commerce 3D showcases, and VR/AR content creation, static 2D images are no longer enough — you need rotatable, interactive 3D assets.

Traditional 3D modeling requires professional Maya/Blender skills with a steep learning curve. Hunyuan3D 2.0 changes the game: it can automatically generate textured 3D mesh models from a single condition image. And ERNIE-Image happens to be one of the best tools for generating those high-quality condition images.

The combination advantage:

  • ERNIE-Image: Precise text rendering + structured layout → condition images with clear text/logos
  • Hunyuan3D-DiT: 2.6B parameter Flow Diffusion Transformer → accurate geometry from condition images
  • Hunyuan3D-Paint: Geometric priors + diffusion → high-resolution texture maps

Hunyuan3D 2.0 Architecture Overview

Hunyuan3D 2.0 consists of two core components:

Hunyuan3D-DiT (Shape Generation)

Built on a Flow-based Diffusion Transformer with 2.6B parameters. Its core task is to extract 3D geometric information from condition images (or multi-view images) and generate untextured mesh models.

Key features:

  • Supports single-image input (single-view reconstruction) and multi-view input
  • Generated geometry aligns closely with condition image contours and structure
  • Output formats: .obj or .glb

Hunyuan3D-Paint (Texture Synthesis)

The texture synthesis model leverages geometric priors and diffusion capabilities to generate high-resolution, vibrant texture maps for either generated or hand-crafted mesh models.

Key features:

  • Texture resolution up to 2048×2048
  • Automatic UV unwrapping
  • Output texture maps directly usable in game engines (Unity/Unreal Engine)

Complete 5-Step Workflow: Text to 3D

Step 1: Generate Condition Image with ERNIE-Image

This is the most critical step in the entire pipeline. The quality of the condition image directly determines the geometric accuracy of the 3D model.

Best practices:

  • Use ERNIE-Image (50-step SFT version) rather than Turbo — final output needs maximum quality
  • Guidance scale: guidance_scale=4.0
  • Resolution: 1024×1024
  • Enable PE (Prompt Enhancer) for richer details

Prompt example:

A detailed 3D-rendered product of a ceramic coffee mug, white background,
front view, clean lighting, no shadows, product photography style,
high resolution

Why ERNIE-Image is particularly well-suited?
ERNIE-Image's text rendering capability means you can directly embed brand logos or product names into condition images, and these texts will appear clearly in the 3D texture — something other open-source models struggle with.

Step 2: Background Removal (Optional but Recommended)

Hunyuan3D performs better with solid-color background condition images. Use ComfyUI's RemBG node or BiRefNet for background removal.

# BiRefNet quick background removal
from birefnet import BiRefNet
model = BiRefNet.from_pretrained("bi-refnet")
transparent_image = model.remove_background(ernie_output)

Step 3: Hunyuan3D-DiT Geometry Generation

Feed the cleaned condition image into Hunyuan3D-DiT:

from hunyuan3d import Hunyuan3DDiT

Initialize model

shape_model = Hunyuan3DDiT.from_pretrained("Tencent/Hunyuan3D-2.0-DiT")

Generate geometry from condition image

mesh = shape_model.generate(
condition_image=cleaned_condition,
num_steps=50,
guidance_scale=7.5
)

Export OBJ/GLB

mesh.export("output/model.obj")
mesh.export_glb("output/model.glb")

Step 4: Hunyuan3D-Paint Texture Synthesis

from hunyuan3d import Hunyuan3DPaint

Initialize texture model

paint_model = Hunyuan3DPaint.from_pretrained("Tencent/Hunyuan3D-2.0-Paint")

Generate texture maps

textured_mesh = paint_model.generate(
mesh=mesh,
reference_image=cleaned_condition,
texture_resolution=(2048, 2048)
)

Export textured model

textured_mesh.export_glb("output/textured_model.glb")

Step 5: Orchestrate in ComfyUI

ComfyUI provides a visual node-based workflow that chains the entire pipeline:

[ERNIE-Image Node] → [RemBG Node] → [Hunyuan3D-DiT Node] → [Hunyuan3D-Paint Node] → [GLB Export Node]

ComfyUI node installation:

# Update ComfyUI to latest version (native Hunyuan3D support)
cd ComfyUI
git pull

Or use ComfyUI-Hunyuan3DWrapper plugin

cd ComfyUI/custom_nodes
git clone https://github.com/kijai/ComfyUI-Hunyuan3DWrapper.git
cd ComfyUI-Hunyuan3DWrapper
pip install -r requirements.txt

Hardware Requirements & Performance

VRAM Requirements

Step Model BF16 VRAM FP8 VRAM GGUF VRAM
Condition Image ERNIE-Image ~16 GB ~10 GB ~8 GB
Geometry Hunyuan3D-DiT ~20 GB ~12 GB Not supported
Texture Hunyuan3D-Paint ~14 GB ~8 GB Not supported

Recommended configurations:

  • Minimum: RTX 3090 (24GB) — FP8 quantization, step-by-step execution
  • Recommended: RTX 4090 (24GB) — Full pipeline in FP8
  • Ideal: A100 80GB — All models loaded simultaneously

Generation Time Estimates

On RTX 4090:

Step Time (BF16) Time (FP8)
ERNIE-Image Turbo (8 steps) ~5 sec ~3 sec
ERNIE-Image SFT (50 steps) ~25 sec ~15 sec
Hunyuan3D-DiT Geometry ~30 sec ~18 sec
Hunyuan3D-Paint Texture ~40 sec ~25 sec
Total ~100 sec ~61 sec

With ERNIE-Image Turbo + FP8 quantization, the total time drops to under 1 minute.

Practical Use Cases

E-commerce Product Showcase

  1. Generate product front-view with ERNIE-Image (including brand text)
  2. Hunyuan3D generates 360° rotatable 3D model
  3. Export GLB for Shopify/Taobao 3D product pages

Advantage: No physical photography needed — hundreds of 3D product showcases per day.

Game Asset Rapid Prototyping

  1. Design phase: ERNIE-Image generates character/prop concept art
  2. Hunyuan3D quickly generates engine-importable 3D models
  3. Art team refines on top of generated base

Advantage: Reduces concept-to-3D-prototype cycle from days to minutes.

Educational 3D Content

  1. ERNIE-Image generates textbook illustrations (with text annotations)
  2. Hunyuan3D generates 3D teaching models
  3. Embed into VR/AR educational applications

FAQ

Q: Text in 3D models is blurry. What to do?

ERNIE-Image's text rendering is strong in 2D, but Hunyuan3D's texture synthesis may blur text details. Suggestions:

  • Use larger font sizes in the ERNIE-Image prompt
  • Set texture resolution to 2048×2048 instead of default 1024×1024
  • For critical text, consider adding it in a 3D modeling tool post-generation

Q: Single-view vs Multi-view — which is better?

  • Single-view: Faster, suitable for simple objects (products, vessels)
  • Multi-view: Generate 4-6 angle views with ERNIE-Image; geometric accuracy improves significantly for complex objects (characters, machinery)

Q: Can I run this on a 16GB GPU?

Yes, but you'll need:

  1. ERNIE-Image Turbo NVFP4 quantization (~4.8GB)
  2. Execute geometry and texture steps sequentially, freeing VRAM between
  3. CPU offloading (speed will drop noticeably)

Summary

The ERNIE-Image × Hunyuan3D 2.0 combination represents the best practice for open-source 3D content generation in 2026:

  • Fully open-source: Apache 2.0 license, worry-free commercial use
  • Fully automated: Text descriptions to 3D assets, end-to-end
  • High quality: ERNIE-Image text rendering + Hunyuan3D geometric precision
  • Deployable: ComfyUI visual orchestration, SGLang production deployment

As both models continue to iterate, this workflow's quality will only improve. For teams needing bulk 3D content, this is a tech stack worth investing in.

ERNIE-Image Team