ERNIE-Image × WAN 2.7 Image-to-Video Complete Workflow: From First/Last Frame Control to Native Audio

Jul 1, 2026

ERNIE-Image × WAN 2.7 Image-to-Video Complete Workflow: From First/Last Frame Control to Native Audio

TL;DR: WAN 2.7 is Alibaba's latest AI video model, adding first/last frame control, 9-grid I2V, voice cloning, and native audio over version 2.6. This article details the complete ERNIE-Image + WAN 2.7 workflow in ComfyUI, covering the end-to-end pipeline from high-quality image generation to audio-synced video.


WAN 2.7 Core Upgrades

In late March 2026, Alibaba's Tongyi Lab released WAN 2.7 — the latest in the Wan video generation series. If 2.6 answered "can it move?", 2.7 answers "can it be precisely controlled?"

5 Core Upgrades:

Feature WAN 2.6 WAN 2.7 Improvement
First/Last Frame Control Separate checkpoint Integrated into main model ~30% precision improvement
Multi-frame I2V Single reference 9-grid reference New capability
Voice Cloning None Subject + voice reference cloning New capability
Instruction Editing Not supported Supported New capability
Native Audio Manual overlay Synchronized generation 80-90% accuracy

Architecture: Full Attention DiT

WAN 2.7 uses a Diffusion Transformer (DiT) with Full Attention architecture. Unlike frame-by-frame processing, it handles spatial and temporal relationships across the entire video sequence simultaneously. This means significantly less character drift and physical inconsistency in 15-second clips.

"WAN 2.7's real shift isn't flashier visuals. It's control. You can finally tell the model exactly where the video starts, where it ends, and what the character must look like — and it actually listens." — SeaArt AI Review


ERNIE-Image × WAN 2.7 Base Workflow

Why ERNIE-Image as the Starting Point?

ERNIE-Image, as an 8B DiT text-to-image model, is an ideal precursor for image-to-video in three dimensions:

  1. High-quality static images: Leading open-source performance in character consistency, scene detail, and lighting
  2. Accurate text rendering: LongTextBench score of 0.9733 — generated images contain accurate text
  3. Structured layout capability: Posters, comics, multi-panel layouts — providing clear starting points for video generation

Workflow A: ERNIE-Image → WAN 2.7 I2V (Basic)

Text Prompt → ERNIE-Image (1024×1024) → WAN 2.7 I2V (1080p, 15s) → Output Video

ComfyUI Node Connection:

  1. ERNIE-Image Node (ErnieImagePipeline)

    • Input: Text prompt
    • Output: 1024×1024 high-quality image
    • Recommendation: Base version 50 steps / Turbo version 8 steps
  2. WAN 2.7 I2V Node (ComfyUI Partner Nodes)

    • Input: ERNIE-Image output image
    • Parameters: 15s duration, 1080p resolution
    • Output: Video with native audio

Complete Workflow Download: ComfyUI officially provides WAN 2.7 Image to Video Workflow templates

Workflow B: First/Last Frame Control (Advanced)

Prompt A → ERNIE-Image (Start Frame) → \
                               → WAN 2.7 (First/Last Frame) → Coherent Video
Prompt B → ERNIE-Image (End Frame)   → /

This is WAN 2.7's killer feature. ERNIE-Image generates the start and end frames, WAN 2.7 generates the coherent transition.

Practical Example:

  • Start frame: Rider on a Bajaj scooter approaching from 30 feet
  • End frame: Rider at arm's length, hand raised, speaking to camera
  • WAN 2.7 generates the coherent approach sequence

Technical Tips:

  • Keep the same character appearance in both frames (ERNIE-Image excels at character consistency)
  • Use the same aspect ratio (16:9 recommended)
  • Don't position the character too extremely in either frame

9-Grid Image-to-Video: From Static to Narrative

What is 9-Grid I2V?

WAN 2.7's new 9-Grid Image-to-Video feature converts a 3×3 grid of reference images into a single coherent video with smooth transitions.

ERNIE-Image 9-Grid Pipeline:

  1. Generate 9 scene images with ERNIE-Image (same character, different poses/angles)
  2. Use ComfyUI's Grid Image node to merge into a 3×3 grid
  3. Input to WAN 2.7 9-Grid I2V node
  4. Output: Coherent narrative video

Technical Requirements

  • Keep the same aspect ratio across all panels
  • Reading order: left-to-right, top-to-bottom
  • Minimum 512px on short edge per panel
  • ERNIE-Image Turbo (8 steps) is ideal for quickly generating 9 reference images

Voice Cloning and Native Audio

Subject & Voice Reference Cloning

WAN 2.7 supports uploading a character image + short audio clip to simultaneously replicate visual appearance and vocal characteristics.

Workflow:

  1. ERNIE-Image generates character reference image
  2. Record/extract voice reference (5-10 second clip)
  3. WAN 2.7 locks both visual appearance and voice
  4. Generates video with synchronized dialogue

Native Audio Quality

Scenario Accuracy Notes
Engine sound tracking speed 90%+ Tunnel echo auto-generated
Ambient sound 85%+ Complex scenes 80-90%
Character dialogue 70-80% >150 WPM causes drift
Background music 80%+ May need manual ducking during dialogue

Practical Advice: Native audio reaches 80-90% accuracy in complex scenes. A manual audio sync pass before delivery is recommended.


ERNIE-Image × WAN 2.7 Cost Analysis

Local Deployment (Recommended)

Component Hardware Requirement Cost
ERNIE-Image Base 24GB VRAM One-time hardware investment
ERNIE-Image Turbo 16GB VRAM (FP8) Lower VRAM requirement
WAN 2.7 TBD (expected 24-48GB) Local when open weights release

Cloud API

Plan Cost per Video Use Case
WAN 2.7 Free Tier ~$0 (15 credits) Testing
WAN 2.7 Starter $0.40-0.60/video Personal creation
WAN 2.7 Plus $0.40-0.60/video 4-6 videos/month
60-second brand video $6-13 Commercial projects

Complete ComfyUI Workflow Configuration

Node List

  1. ErnieImagePipeline — ERNIE-Image image generation
  2. WAN 2.7 I2V — Image to video conversion
  3. First/Last Frame Control — Frame anchoring
  4. Grid Image Merge — 9-grid assembly
  5. Audio Sync — Audio synchronization (if needed)

Official Template Downloads

ComfyUI officially releases these WAN 2.7 workflow templates:

  • Image to Video Workflow
  • Text to Video Workflow
  • Reference to Video Workflow
  • Video Edit Workflow

Key Parameters

Parameter ERNIE-Image Base ERNIE-Image Turbo WAN 2.7 I2V
Steps 50 8 N/A
Guidance Scale 4.0 1.0 N/A
Resolution 1024×1024 1024×1024 1080p
Duration N/A N/A 5/10/15s
Use PE True True N/A

Practical Case: 60-Second Brand Promo Video

Step Breakdown

  1. ERNIE-Image product photos (Turbo, 8 steps) — 4 product scene images
  2. WAN 2.7 I2V — 15-second video clip per image
  3. First/Last Frame Control — Character consistency across clips
  4. Native Audio — Auto-generated ambient sound and background music
  5. Post-production — Manual audio sync, simple editing

Cost Estimate

  • ERNIE-Image generating 4 images: Local ≈ $0
  • WAN 2.7 generating 4 videos: ~48 credits ≈ $2.40
  • Total: ~$2.40 (cloud) or $0 (fully local)

FAQ

When will WAN 2.7 go open source?

Based on the Wan series release pattern, expect open weights in mid-to-late Q2 2026. Monitor Wan-Video GitHub.

Minimum hardware for ERNIE-Image + WAN 2.7?

ERNIE-Image Turbo FP8 runs on 16GB VRAM. WAN 2.7 local deployment TBD, expected 24-48GB VRAM.

Maximum video resolution?

WAN 2.7 supports up to 1080p output, 15-second duration.


Summary

The ERNIE-Image × WAN 2.7 combination completely breaks the boundary between static image generation and video creation. ERNIE-Image provides high-quality, structured image starting points, while WAN 2.7 delivers precise frame control, native audio, and narrative capabilities.

Key Advantages:

  • ERNIE-Image: 8B parameters, Apache 2.0 open source, strongest text rendering
  • WAN 2.7: First/last frame control, 9-grid I2V, native audio, ComfyUI support
  • Combined: End-to-end open-source pipeline from text to audio-synced high-quality video

As WAN 2.7 weights become open source and the ComfyUI node ecosystem matures, this combination will become the go-to solution for open-source AI video creation.


This article is based on SeaArt AI review, ComfyUI official blog, ERNIE-Image Hub (ernie-image.org), and the ERNIE-Image technical report. All technical data as of July 1, 2026.

ERNIE-Image Team