AlibabaAlibaba·🎬 Video Generation

Wan 3.0

anonymized
Try on Venice.ai ↗
Quick reference
Wan 3.0 — TLDR
  • 🆕 Newest generation of Alibaba's Wan video model family
  • 📏 Generates clips of up to 30 seconds in one pass
  • 🌐 Accepts text, images, audio, video and documents as input
  • 👁️ Image-to-video variant animates a supplied still frame from a prompt
  • 🎯 Emphasis on subject and scene coherence across frames
  • 🏢 Built by Alibaba, delivered through Alibaba Cloud Model Studio
  • ⚡ Text-to-video and reference-to-video siblings ship in the same generation
  • 🔧 Follows Wan 2.7's multimodal image-to-video interface
💰 Pricing
$0.060 – $3.30
per generation
📅 On Venice since
Aug 1, 2026
23 days ago
Provider

Alibaba Group is a Chinese multinational technology company founded in 1999 and headquartered in Hangzhou, Zhejiang. Originally built around e-commerce and cloud computing, Alibaba has become one of the most prolific contributors to open-weight AI research,…

Read full profile →
68 models on Venice
29 video · 22 text · 7 image · 6 inpaint · 2 embedding · 2 tts
Since Jan 11, 2025

About this model

Wan 3.0 is the image-to-video entry point into the newest generation of Alibaba's Wan visual-generation family, presented by Alibaba Cloud as its latest Wan video generation and served through Model Studio. On this catalog it sits alongside its same-generation siblings Wan 3.0 text-to-video and Wan 3.0 Reference, all dated to the same release. You supply a starting frame plus a prompt, and the model produces photorealistic motion with attention to keeping subjects consistent from the first frame onward.

The clearest generational change versus Wan 2.7 is duration and input breadth. Wan 2.7's image-to-video model already accepted multimodal input — text, images, audio and video — across first-frame-to-video, first-and-last-frame-to-video and video continuation tasks. Alibaba describes Wan 3.0 as generating up to 30 seconds of video from any input, adding documents to the list of accepted material.

That difference is practical rather than cosmetic. The earlier image-to-video APIs covering versions 2.1 through 2.6 — the era of Wan 2.6 — were framed around producing a short, smooth clip from a first-frame image and a text prompt. Longer single-pass output reduces the need to stitch several short clips together to sustain continuous camera movement or a single unbroken take.

Related Wan variants on this catalog include Wan 2.2 Enhanced and the editing-oriented Wan 2.7 Edit.

This About section is AI-generated from public sources (Claude Opus 5), with no human editing. It may contain inaccuracies — verify critical details against the sources listed above.

Data sources: Venice API · HuggingFace · Wikipedia — enrichment updated 12h ago