ERNIE-Image Community Quantized Models Deep Dive: FP8 / INT8 / NVFP4 Benchmarks on RTX 5090 / 3090 / 3060

Jul 24, 2026

ERNIE-Image Community Quantized Models Deep Dive: FP8 / INT8 / NVFP4 Benchmarks on RTX 5090 / 3090 / 3060

ERNIE-Image's 8B model requires roughly 16GB of VRAM at BF16 precision — a significant barrier for many users. Community developer Bedovyy has released quantized versions of both ERNIE-Image and ERNIE-Image-Turbo on HuggingFace, covering three formats (NVFP4, FP8, INT8) with comprehensive benchmarks across three GPU generations: RTX 5090 (Blackwell), RTX 3090 (Ampere), and RTX 3060 (Ampere).

This article breaks down the real-world performance data to help you find the optimal quantization setup.

Understanding the Quantization Formats

  • BF16 (Baseline): 16-bit float, no quantization, ~16GB, lossless quality
  • FP8 (e4m3): 8-bit float, ~8GB file size, minimal quality loss
  • INT8 (rowwise): 8-bit integer quantization, ~8GB, significant speed gains, acceptable quality
  • NVFP4: NVIDIA Blackwell-specific 4-bit float format, ~4.78GB, hardware-accelerated on RTX 50 series only

ERNIE-Image-Turbo Benchmarks

Turbo requires just 8 inference steps. How much faster does quantization make it?

GPU Quantization it/s Time vs BF16
RTX 5090 BF16 2.09 4.87s 100%
FP8 3.69 3.32s 147%
INT8 4.31 3.05s 160%
NVFP4 5.09 2.72s 179%
RTX 3090 BF16 0.88 12.42s 100%
FP8 0.84 12.73s 98% (regression)
INT8 1.66 7.04s 176%
NVFP4 0.83 12.71s 98% (regression)
RTX 3060 BF16 0.26 43.02s 100%
FP8 0.39 28.66s 150%
INT8 0.82 14.43s 298%
NVFP4 0.39 28.72s 150%

ERNIE-Image Base Benchmarks

Base requires 50 inference steps, making quantization gains even more dramatic:

GPU Quantization it/s Time vs BF16
RTX 5090 BF16 1.08 20.08s 100%
FP8 1.97 11.67s 172%
INT8 2.14 10.89s 184%
NVFP4 2.56 9.35s 215%
RTX 3090 BF16 0.40 53.33s 100%
INT8 0.79 28.08s 190%
FP8 0.39 54.71s 97% (regression)
NVFP4 0.38 55.20s 97% (regression)
RTX 3060 BF16 0.11 201.41s 100%
FP8 0.17 130.48s 154%
INT8 0.35 62.42s 323%
NVFP4 0.17 130.87s 154%

Four Key Takeaways

1. NVFP4 is RTX 50 Series Exclusive

NVFP4 delivers a 79% speed boost on Turbo and 115% on Base — but only on RTX 5090 (Blackwell architecture). On RTX 3090 (Ampere) and RTX 3060 (Ampere), NVFP4 actually performs slightly worse than BF16.

Bottom line: NVFP4 is only for RTX 50 series users. If you're on the 30 series, stay away from NVFP4.

2. INT8 is the Universal Winner

int8rowwise quantization shows significant speedups on every tested GPU with zero regressions:

  • RTX 5090: 60-84% faster
  • RTX 3090: 76-90% faster
  • RTX 3060: 198-223% faster (nearly 3x!)

For 30 series users, INT8 is the only quantization worth considering.

3. FP8 is a Decent Second Choice — But Watch Out on 3090

FP8 delivers 50-70% speedup on RTX 5090 and RTX 3060. But on RTX 3090, it actually shows a 3% regression with the Base model — suggesting Ampere's large-core architecture doesn't handle FP8 well.

4. Base Model Gains from INT8 Outweigh Turbo Gains

Interesting finding: on RTX 3060, switching Base from BF16 to INT8 cuts generation time from 201 seconds to 62 seconds — saving 139 seconds per image. The absolute time savings are much larger than with Turbo, even though the percentage speedup is similar.

This means: if you need Base quality (50-step inference) but are limited by a mid-range GPU, INT8 quantization is the single most impactful optimization you can make.

Recommended Configurations by GPU

RTX 5090

Model: ERNIE-Image-Turbo
Quantization: NVFP4 (best) or INT8 (second)
Speed: 2.72s per image (8 steps)

NVFP4 makes Turbo generation near-real-time on the 5090. Even Base at 50 steps completes in 9.35 seconds.

RTX 3090 / 3080

Model: ERNIE-Image-Turbo
Quantization: INT8 (only choice)
Speed: 7.04s per image (8 steps)

Skip NVFP4 and FP8. INT8 works well on the 3090, cutting Turbo from 12.4s to 7s.

RTX 3060 / 4060

Model: ERNIE-Image-Turbo or Base
Quantization: INT8 (strongly recommended)
Speed: 14.43s (Turbo 8-step) / 62.42s (Base 50-step)

INT8 brings Turbo on the 3060 from 43 seconds down to 14 seconds — transforming it from "unbearable" to "acceptable." Base at 62 seconds is also practical.

How to Use Community Quantized Models

First, download the quantized model files:

# From HuggingFace
git lfs install
git clone https://huggingface.co/Bedovyy/ERNIE-Image-Quantized

Or download specific formats directly

NVFP4: ernie-image-turbo-nvfp4.safetensors (~4.78GB)

FP8: ernie-image-turbo-fp8.safetensors (~8.22GB)

INT8: ernie-image-turbo-int8.safetensors (~8.22GB)

Place in your ComfyUI models directory:

ComfyUI/models/diffusion_models/ernie-image-turbo-int8.safetensors

To quantize models yourself, use the open-source comfy-dit-quantizer:

git clone https://github.com/bedovyy/comfy-dit-quantizer
cd comfy-dit-quantizer
pip install -r requirements.txt
python quantize.py --model ernie-image-turbo.safetensors --config config.json

Quality Considerations

Community quantization typically preserves diagonal layers (adaLN and self_attention.norm) at full precision while quantizing MLP and attention projection layers. Based on Bedovyy's sample comparisons, FP8 and INT8 show virtually no perceptible quality difference from BF16. NVFP4 has minor detail loss in some cases, but at 4.78GB (one-third of BF16), it's a worthwhile tradeoff for most applications.

If you're extremely quality-sensitive (e.g., commercial print), use at least FP8. INT8 occasionally shows subtle quality loss in complex text rendering scenarios, but is essentially imperceptible for standard image generation.

Summary

The ERNIE-Image community quantization ecosystem has matured significantly. Bedovyy's four-format quantization suite covers everything from flagship RTX 5090 to entry-level RTX 3060, giving users at every budget level a viable deployment path.

Core recommendations:

  • RTX 50 series → NVFP4 (fastest)
  • RTX 30 series → INT8 (only effective choice)
  • Low-end GPUs → INT8 + Turbo (3x speedup, from unusable to usable)
  • Quality priority → FP8 (50%+ speedup with near-lossless quality on most GPUs)

ERNIE-Image Team