AlibabaAlibaba·🎬 Video Generation

Wan 3.0 Reference

anonymized
Try on Venice.ai ↗
Quick reference
Wan 3.0 Reference — TLDR
  • 🏢 Alibaba's Wan 3.0 video family, in reference-to-video mode
  • 🆕 Newest Wan generation here; follows Wan 2.7 Reference
  • 📏 Alibaba describes 30-second video generation in a single pass
  • 🌐 Positioned as generating video "from any input"
  • 🎯 Reference material anchors subject and scene consistency across frames
  • 👁️ Photorealistic output with strong subject coherence between frames
  • 💬 Sibling endpoints cover text-to-video and image-to-video separately
  • ⚡ Released August 2026 by Alibaba
💰 Pricing
$0.060 – $3.30
per generation
📅 On Venice since
Aug 1, 2026
23 days ago
Provider

Alibaba Group is a Chinese multinational technology company founded in 1999 and headquartered in Hangzhou, Zhejiang. Originally built around e-commerce and cloud computing, Alibaba has become one of the most prolific contributors to open-weight AI research,…

Read full profile →
68 models on Venice
29 video · 22 text · 7 image · 6 inpaint · 2 embedding · 2 tts
Since Jan 11, 2025

About this model

Wan 3.0 Reference is the reference-to-video entry point into Alibaba's Wan video generation family, released in August 2026. Rather than describing a scene purely in words, you supply reference material — images or other assets that define a character, product or setting — and the model animates around them, aiming to keep that subject recognisable from frame to frame.

The headline generational change over Wan 2.7 Reference is clip length and input breadth: Alibaba presents Wan 3.0 as generating 30 seconds of video from any input in a single pass. That framing shifts the family away from short novelty clips toward longer takes, where a subject has to stay believable for the duration rather than for a few seconds.

Within this catalog, the variant sits alongside the same-generation Wan 3.0 text-to-video and Wan 3.0 image-to-video endpoints, which expose the other input paths of the same release. Earlier lineage remains listed too, including Wan 2.7 and Wan 2.2 Enhanced, if you need previous-generation behaviour.

Choose Wan 3.0 Reference when identity lock matters — recurring characters, branded objects, or continuity across shots — and a purely prompt-driven generation would drift. For scenes with no fixed subject, the text-to-video sibling is the simpler starting point.

This About section is AI-generated from public sources (Claude Opus 5), with no human editing. It may contain inaccuracies — verify critical details against the sources listed above.

Data sources: Venice API · HuggingFace · Wikipedia — enrichment updated 12h ago