ERNIE-Image × WAN 2.7 Image-to-Video Complete Workflow: From First/Last Frame Control to Native Audio
TL;DR: WAN 2.7 is Alibaba's latest AI video model, adding first/last frame control, 9-grid I2V, voice cloning, and native audio over version 2.6. This article details the complete ERNIE-Image + WAN 2.7 workflow in ComfyUI, covering the end-to-end pipeline from high-quality image generation to audio-synced video.
WAN 2.7 Core Upgrades
In late March 2026, Alibaba's Tongyi Lab released WAN 2.7 — the latest in the Wan video generation series. If 2.6 answered "can it move?", 2.7 answers "can it be precisely controlled?"
5 Core Upgrades:
| Feature | WAN 2.6 | WAN 2.7 | Improvement |
|---|---|---|---|
| First/Last Frame Control | Separate checkpoint | Integrated into main model | ~30% precision improvement |
| Multi-frame I2V | Single reference | 9-grid reference | New capability |
| Voice Cloning | None | Subject + voice reference cloning | New capability |
| Instruction Editing | Not supported | Supported | New capability |
| Native Audio | Manual overlay | Synchronized generation | 80-90% accuracy |
Architecture: Full Attention DiT
WAN 2.7 uses a Diffusion Transformer (DiT) with Full Attention architecture. Unlike frame-by-frame processing, it handles spatial and temporal relationships across the entire video sequence simultaneously. This means significantly less character drift and physical inconsistency in 15-second clips.
"WAN 2.7's real shift isn't flashier visuals. It's control. You can finally tell the model exactly where the video starts, where it ends, and what the character must look like — and it actually listens." — SeaArt AI Review
ERNIE-Image × WAN 2.7 Base Workflow
Why ERNIE-Image as the Starting Point?
ERNIE-Image, as an 8B DiT text-to-image model, is an ideal precursor for image-to-video in three dimensions:
- High-quality static images: Leading open-source performance in character consistency, scene detail, and lighting
- Accurate text rendering: LongTextBench score of 0.9733 — generated images contain accurate text
- Structured layout capability: Posters, comics, multi-panel layouts — providing clear starting points for video generation
Workflow A: ERNIE-Image → WAN 2.7 I2V (Basic)
Text Prompt → ERNIE-Image (1024×1024) → WAN 2.7 I2V (1080p, 15s) → Output Video
ComfyUI Node Connection:
ERNIE-Image Node (ErnieImagePipeline)
- Input: Text prompt
- Output: 1024×1024 high-quality image
- Recommendation: Base version 50 steps / Turbo version 8 steps
WAN 2.7 I2V Node (ComfyUI Partner Nodes)
- Input: ERNIE-Image output image
- Parameters: 15s duration, 1080p resolution
- Output: Video with native audio
Complete Workflow Download: ComfyUI officially provides WAN 2.7 Image to Video Workflow templates
Workflow B: First/Last Frame Control (Advanced)
Prompt A → ERNIE-Image (Start Frame) → \
→ WAN 2.7 (First/Last Frame) → Coherent Video
Prompt B → ERNIE-Image (End Frame) → /
This is WAN 2.7's killer feature. ERNIE-Image generates the start and end frames, WAN 2.7 generates the coherent transition.
Practical Example:
- Start frame: Rider on a Bajaj scooter approaching from 30 feet
- End frame: Rider at arm's length, hand raised, speaking to camera
- WAN 2.7 generates the coherent approach sequence
Technical Tips:
- Keep the same character appearance in both frames (ERNIE-Image excels at character consistency)
- Use the same aspect ratio (16:9 recommended)
- Don't position the character too extremely in either frame
9-Grid Image-to-Video: From Static to Narrative
What is 9-Grid I2V?
WAN 2.7's new 9-Grid Image-to-Video feature converts a 3×3 grid of reference images into a single coherent video with smooth transitions.
ERNIE-Image 9-Grid Pipeline:
- Generate 9 scene images with ERNIE-Image (same character, different poses/angles)
- Use ComfyUI's Grid Image node to merge into a 3×3 grid
- Input to WAN 2.7 9-Grid I2V node
- Output: Coherent narrative video
Technical Requirements
- Keep the same aspect ratio across all panels
- Reading order: left-to-right, top-to-bottom
- Minimum 512px on short edge per panel
- ERNIE-Image Turbo (8 steps) is ideal for quickly generating 9 reference images
Voice Cloning and Native Audio
Subject & Voice Reference Cloning
WAN 2.7 supports uploading a character image + short audio clip to simultaneously replicate visual appearance and vocal characteristics.
Workflow:
- ERNIE-Image generates character reference image
- Record/extract voice reference (5-10 second clip)
- WAN 2.7 locks both visual appearance and voice
- Generates video with synchronized dialogue
Native Audio Quality
| Scenario | Accuracy | Notes |
|---|---|---|
| Engine sound tracking speed | 90%+ | Tunnel echo auto-generated |
| Ambient sound | 85%+ | Complex scenes 80-90% |
| Character dialogue | 70-80% | >150 WPM causes drift |
| Background music | 80%+ | May need manual ducking during dialogue |
Practical Advice: Native audio reaches 80-90% accuracy in complex scenes. A manual audio sync pass before delivery is recommended.
ERNIE-Image × WAN 2.7 Cost Analysis
Local Deployment (Recommended)
| Component | Hardware Requirement | Cost |
|---|---|---|
| ERNIE-Image Base | 24GB VRAM | One-time hardware investment |
| ERNIE-Image Turbo | 16GB VRAM (FP8) | Lower VRAM requirement |
| WAN 2.7 | TBD (expected 24-48GB) | Local when open weights release |
Cloud API
| Plan | Cost per Video | Use Case |
|---|---|---|
| WAN 2.7 Free Tier | ~$0 (15 credits) | Testing |
| WAN 2.7 Starter | $0.40-0.60/video | Personal creation |
| WAN 2.7 Plus | $0.40-0.60/video | 4-6 videos/month |
| 60-second brand video | $6-13 | Commercial projects |
Complete ComfyUI Workflow Configuration
Node List
- ErnieImagePipeline — ERNIE-Image image generation
- WAN 2.7 I2V — Image to video conversion
- First/Last Frame Control — Frame anchoring
- Grid Image Merge — 9-grid assembly
- Audio Sync — Audio synchronization (if needed)
Official Template Downloads
ComfyUI officially releases these WAN 2.7 workflow templates:
- Image to Video Workflow
- Text to Video Workflow
- Reference to Video Workflow
- Video Edit Workflow
Key Parameters
| Parameter | ERNIE-Image Base | ERNIE-Image Turbo | WAN 2.7 I2V |
|---|---|---|---|
| Steps | 50 | 8 | N/A |
| Guidance Scale | 4.0 | 1.0 | N/A |
| Resolution | 1024×1024 | 1024×1024 | 1080p |
| Duration | N/A | N/A | 5/10/15s |
| Use PE | True | True | N/A |
Practical Case: 60-Second Brand Promo Video
Step Breakdown
- ERNIE-Image product photos (Turbo, 8 steps) — 4 product scene images
- WAN 2.7 I2V — 15-second video clip per image
- First/Last Frame Control — Character consistency across clips
- Native Audio — Auto-generated ambient sound and background music
- Post-production — Manual audio sync, simple editing
Cost Estimate
- ERNIE-Image generating 4 images: Local ≈ $0
- WAN 2.7 generating 4 videos: ~48 credits ≈ $2.40
- Total: ~$2.40 (cloud) or $0 (fully local)
FAQ
When will WAN 2.7 go open source?
Based on the Wan series release pattern, expect open weights in mid-to-late Q2 2026. Monitor Wan-Video GitHub.
Minimum hardware for ERNIE-Image + WAN 2.7?
ERNIE-Image Turbo FP8 runs on 16GB VRAM. WAN 2.7 local deployment TBD, expected 24-48GB VRAM.
Maximum video resolution?
WAN 2.7 supports up to 1080p output, 15-second duration.
Summary
The ERNIE-Image × WAN 2.7 combination completely breaks the boundary between static image generation and video creation. ERNIE-Image provides high-quality, structured image starting points, while WAN 2.7 delivers precise frame control, native audio, and narrative capabilities.
Key Advantages:
- ERNIE-Image: 8B parameters, Apache 2.0 open source, strongest text rendering
- WAN 2.7: First/last frame control, 9-grid I2V, native audio, ComfyUI support
- Combined: End-to-end open-source pipeline from text to audio-synced high-quality video
As WAN 2.7 weights become open source and the ComfyUI node ecosystem matures, this combination will become the go-to solution for open-source AI video creation.
This article is based on SeaArt AI review, ComfyUI official blog, ERNIE-Image Hub (ernie-image.org), and the ERNIE-Image technical report. All technical data as of July 1, 2026.