ERNIE-Image Editing Model Latest Updates: Bernini-R + FLUX Kontext Hybrid Pipeline in Practice

Jun 25, 2026

ERNIE-Image Editing Model Latest Updates: Bernini-R + FLUX Kontext Hybrid Pipeline in Practice

ERNIE-Image's official editing model continues to delay. ByteDance's Bernini-R brings a unified video editing framework. Combined with FLUX Kontext, build a complete open-source image + video editing pipeline.

In April 2026, ERNIE-Image 8B text-to-image model was open-sourced, and the community erupted. But one question has remained unanswered: when will the official editing model launch?

From Reddit's "same indefinite delay as Z-Image Edit?" discussion to the community alternative search documented in EI-096, the absence of an editing model has been the biggest pain point in the ERNIE-Image ecosystem.

In June 2026, the landscape shifted. ByteDance released Bernini-R — a unified video editing framework based on the Wan 2.2 architecture, with native ComfyUI support and Apache 2.0 open-source licensing. Meanwhile, FLUX Kontext continues to dominate the image editing space.

This article demonstrates how to chain Bernini-R, FLUX Kontext, and ERNIE-Image into a complete open-source image + video editing pipeline.

Bernini-R: ByteDance's Unified Editing Framework

Architecture Highlights

Bernini-R's core innovation is its unified architecture — one model handling multiple editing tasks:

  • Text-to-Video Generation
  • Video Editing
  • Reference-Guided Video Editing (RV2V)
  • Multi-Reference-to-Video Generation

Technical Specifications

Feature Specification
Base Architecture Wan 2.2 DiT (dual-expert Transformer)
High-Noise Transformer Coarse-grained editing
Low-Noise Transformer Fine-grained editing
Semantic Planner Qwen2.5-VL-7B-Instruct
Position Encoding SA3D RoPE (segment-aware 3D rotary position embeddings)
Precision Versions FP16 (24GB+ VRAM) / FP8 (16GB+ VRAM)
License Apache 2.0
ComfyUI Support Official native nodes

The Dual Transformer Architecture

Bernini-R's dual Transformer design is inspired by staged denoising theory in diffusion models:

  • High-Noise Transformer: Handles large-scale structural changes and subject replacement (e.g., character clothing swap, scene transitions)
  • Low-Noise Transformer: Handles fine detail optimization and texture refinement

This design lets Bernini-R complete a full coarse-to-fine editing process in a single inference flow, without running multiple independent models as traditional approaches require.

FLUX Kontext: The Image Editing Benchmark

Core Capabilities

FLUX Kontext (12B parameters) remains the benchmark for open-source image editing:

  • Precise local redrawing: Mask-based inpainting quality far exceeds most alternatives
  • Text-guided editing: Editing with natural language instructions + visual masks
  • Outpainting expansion: Style-consistent canvas expansion
  • ComfyUI official support: June 2026 ComfyUI update fixed memory management issues

Hardware Requirements

Precision VRAM Use Case
BF16 24GB+ Best quality
FP8 16GB+ Balanced
GGUF Q4 12GB+ Low-end

Hybrid Pipeline in Practice: ERNIE-Image + FLUX Kontext + Bernini-R

Complete Workflow Architecture

[ERNIE-Image Text-to-Image] → High-quality static images
     ↓
[FLUX Kontext Image Editing] → Refined images (inpainting, outpainting, object replacement)
     ↓
[Bernini-R Video Editing] → Dynamic video (character animation, scene transitions, video interpolation)
     ↓
[Output: Edited video]

Scenario 1: E-commerce Product Showcase Video

1. ERNIE-Image generates products in different scenes
2. FLUX Kontext removes backgrounds, adds brand elements
3. Bernini-R converts refined images to showcase videos

Scenario 2: Character Animation Creation

1. ERNIE-Image generates character design images
2. FLUX Kontext adjusts character details (clothing, expressions)
3. Bernini-R generates character animation via reference guidance

Scenario 3: Scene Transition and Style Transfer

1. ERNIE-Image generates original scene
2. FLUX Kontext local redrawing (season change, time transition)
3. Bernini-R generates smooth transition video

ComfyUI Configuration Guide

Bernini-R Node Setup

# Download Bernini-R model weights
# HF: https://huggingface.co/ByteDance/Bernini-R
# ComfyUI version: https://huggingface.co/Comfy-Org/Bernini-R

File placement:

models/unet/bernini-r-high-noise.safetensors

models/unet/bernini-r-low-noise.safetensors

models/vae/bernini-vae.safetensors

FLUX Kontext Node Setup

# ComfyUI natively supports FLUX Kontext
# No additional custom nodes needed

Model files:

models/unet/flux-kontext-dev.safetensors

models/vae/flux-vae.safetensors

Memory Management

When chaining multiple models in the same ComfyUI workflow, memory management is critical:

{
  "memory_management": {
    "unload_models_after_use": true,
    "clear_cache_between_nodes": true,
    "max_vram": "24GB"
  }
}

Key tip: Add a Clear Cache node between FLUX Kontext and Bernini-R to free up VRAM from the previous model.

Performance Comparison: Editing Approaches

Approach Image Editing Video Editing VRAM Ease of Use Open Source
FLUX Kontext ⭐⭐⭐⭐⭐ ❌ 24GB+ High ✅
Bernini-R ⭐⭐⭐⭐ ⭐⭐⭐⭐⭐ 16GB+ High ✅
FLUX Kontext + Bernini-R ⭐⭐⭐⭐⭐ ⭐⭐⭐⭐⭐ 24GB+ Medium ✅
Midjourney Edit ⭐⭐⭐⭐⭐ ❌ API High ❌
Runway Gen-3 ⭐⭐⭐⭐ ⭐⭐⭐⭐ API High ❌

The FLUX Kontext + Bernini-R combination achieves the best balance between editing capability and open-source controllability.

Progress Comparison with EI-096

In EI-096 (published 2026-06-16), we documented the community anxiety over the editing model delay. Alternative options were limited:

  • FLUX Kontext was the only mature open-source image editing solution
  • FLUX.2-Klein 9B served as a lightweight supplement
  • Community expectations for ERNIE's official editing model gradually turned to disappointment

Now (2026-06-25), Bernini-R fills the video editing gap:

  • Bernini-R released: ByteDance open-sourced in June, native ComfyUI support
  • Community validation: YouTube tutorials 8500+ views, active Reddit discussion
  • GGUF quantized version: Community-released GGUF version lowers the barrier further
  • Lightx2v LoRAs: Community LoRA acceleration solutions are mature

FAQ

Q: Why two models (FLUX Kontext + Bernini-R)?
A: FLUX Kontext has higher quality for pure image editing (inpainting, outpainting), while Bernini-R excels at video editing. They complement rather than replace each other.

Q: Can Bernini-R do image editing?
A: Yes, but quality doesn't match specialized FLUX Kontext. For image-only editing, FLUX Kontext alone suffices.

Q: Will ERNIE-Image's official editing model still be released?
A: Baidu hasn't provided a clear timeline. The community has shifted to the Bernini-R + FLUX Kontext hybrid as a practical alternative.

Q: Is the quality loss significant for the FP8 version?
A: According to community testing, Bernini-R FP8 has minimal visual quality difference from FP16 (~5% detail loss), but VRAM requirements drop by approximately 50%.

Summary

The editing capability puzzle in the ERNIE-Image ecosystem is being completed. FLUX Kontext handles image refinement, Bernini-R handles video editing, and ERNIE-Image serves as the high-quality static image starting point — the trio forms a complete open-source creative pipeline.

While ERNIE-Image's official editing model is still pending, the community has found viable alternatives. For content creators, today's open-source toolchain is already powerful enough: 8B text-to-image → 12B image editing → 14B video editing — the entire workflow is open-source, self-deployable, and fine-tunable.

ERNIE-Image Team