The AI Video Model That Just Hit #1 — With One Big Catch About "Open Weights"
Wan 3.0 Review 2026: Alibaba's new model leads Artificial Analysis in two video categories with 30-second, native-audio generation from text, images, documents, or a webpage URL. The "open source" claim making the rounds online, though, doesn't hold up under a closer look.
Wan 3.0 landed on Artificial Analysis's leaderboards in August 2026 and immediately took the top spot in text-to-video with audio — a genuinely notable result in a field crowded with well-funded competitors from Google, ByteDance, and OpenAI's former Sora team. What makes it more than just another leaderboard entry is the scope of what it accepts as input: text, an image, a reference video, up to 20 mixed reference assets, and — unusually for this category — a PDF, spreadsheet, presentation, or public webpage URL, which the model can use as creative source material for a 30-second audiovisual sequence.
The story that's spread alongside the benchmark result is that Wan 3.0 is open-weight — a natural assumption, since earlier Wan models had a real open-source reputation. That part of the story doesn't hold up under a closer look at Alibaba's own repository, and it's worth untangling before anything else in this review, because it changes who Wan 3.0 is actually for.
What Wan 3.0 Actually Does
Multimodal input, one model
Text, image, first-frame, first-and-last-frame control, reference video, and reference audio all feed into the same model, alongside document and webpage input — up to 20 mixed reference assets per generation, depending on input mode.
30-second native-audio generation
Officially documented at 2 to 30 seconds, at 480P, 720P, or 1080P, 30fps, with dialogue, background music, and environmental sound effects generated alongside the video — audio is on by default and can be disabled.
Documents and webpages as source material
The model accepts PDF, DOC/DOCX, XLS/XLSX, PPT/PPTX, TXT, Markdown, and public webpage URLs as reference input — a genuinely unusual capability that goes well beyond a standard text-to-video model's scope.
Editing and extension
Beyond fresh generation, Wan 3.0 supports editing existing video and extending a clip's length, positioning it as a working tool across a production pipeline rather than a one-shot generator.
Pricing
| Resolution | Official list rate | 30-second cost |
|---|---|---|
| 480P | $0.068/second | ≈$2.04 |
| 720P | $0.14/second | ≈$4.20 |
| 1080P | $0.28/second | ≈$8.40 |
Those are Alibaba's standard international rates through Model Studio, checked directly against official documentation. A temporary 30% promotional discount is live through September 24, 2026, cutting the 720P 30-second cost to roughly $2.10 — worth knowing about if you're pricing a project this week, but not a rate to expect as the ongoing baseline once the promotion ends. There's no unlimited free API tier; Model Studio offers access to try the interface, but generation itself is metered.
Is It Actually #1? What Artificial Analysis Really Shows
Yes, with real qualification worth understanding rather than skipping past. Artificial Analysis currently ranks Wan 3.0 first in text-to-video with audio, at an Elo of 1243 based on 5,771 blind user-preference votes — but its confidence interval (±9) overlaps with second-place Gemini Omni Flash (1238, ±6), which is why Artificial Analysis itself places both in a shared 1-2 rank range rather than declaring a clean, uncontested winner. The methodology is blind pairwise preference, not a lab benchmark measuring objective quality — real, useful signal, but a different kind of claim than "objectively the best."
The more interesting, more honest story is that Wan 3.0 doesn't lead everywhere. It's also first in video editing with audio (1190 Elo) — but it sits sixth in image-to-video with audio (1178 Elo), well behind the category leaders there. That's worth knowing before assuming "#1" means uniformly best: Wan's strength is concentrated in from-scratch generation and editing, not every video-related task Artificial Analysis tracks.
The Open-Weights Controversy
This is the detail most worth getting right before publishing anything about Wan 3.0. Earlier Wan models had a genuine open-weight reputation, and there is an official-looking GitHub repository — AlibabaCloud-Official/Wan3.0 — carrying an Apache-2.0 license and README language describing it as open-source. That repository, as of this review, contains only a README and license file: no model weights, no releases, and just 11 stars and 2 forks, which is not the footprint of an actively-used open-weight release.
More tellingly, Artificial Analysis's own open-weight category leader for text-to-video with audio is MiniMax H3, not Wan 3.0 — independent analysis describes Wan 3.0 as API-only, with Wan 2.2 remaining the last Wan generation with actually downloadable weights. The honest summary: there's a real Apache-2.0 Wan3.0 repository, but no verified evidence that Wan 3.0's model weights are publicly downloadable for self-hosting. Anyone specifically interested in Wan for self-hosting reasons should not assume that's currently possible.
Pros & Cons
✓ Strengths
- ✅ Genuinely leads two major Artificial Analysis categories — text-to-video with audio and video editing with audio
- ✅ 30-second generation with native dialogue, music, and sound effects in one pass is a real capability jump over shorter-clip competitors
- ✅ Document and webpage input is a distinctive, practically useful capability most video models don't offer
- ✅ Official, transparent per-second pricing with no opaque quote-only tier
✗ Weaknesses
- ❌ Despite an Apache-2.0-licensed GitHub repository, no downloadable model weights are currently available — not the self-hostable model some coverage implies
- ❌ Ranks sixth in image-to-video with audio, well behind category leaders there
- ❌ Early independent testing reports inconsistent character consistency, imperfect on-screen text, and occasional painterly-looking backgrounds in complex scenes
- ❌ Very new (August 2026 release) with a still-small independent review pool
Wan 3.0 vs MiniMax H3 vs Gemini Omni Flash vs Seedance
| Model | Text-to-video Elo (with audio) | Open weights | Best for |
|---|---|---|---|
| Wan 3.0 | 1243 | No (API only) | Longest single-pass generation, document/webpage input |
| Gemini Omni Flash | 1238 | No | Google multimodal ecosystem integration |
| MiniMax H3 | 1225 | Yes | Teams that specifically need self-hostable, downloadable weights |
| Seedance 2.0 (720p) | 1220 | No | Cinematic, reference-driven creative workflows |
The Elo gaps between these four are narrow enough that confidence intervals overlap across most of them — this is a genuinely competitive top tier, not a single runaway leader. The clearer differentiator for most readers is access model rather than raw score: MiniMax H3 is the pick if downloadable, self-hostable weights are the actual requirement, since it's the one Artificial Analysis itself credits as leading the open-weight category. One independent platform test reported Wan 3.0 generating a 30-second 720p clip with audio at meaningfully lower platform-credit cost than Seedance 2.5 — a real data point, though it reflects one provider's pricing rather than a universal comparison across every platform hosting these models.
Who It's For
Choose Wan 3.0 if: you need longer single-pass generations with native audio, want to turn documents or reference webpages into video content, and are comfortable working through an API rather than a polished creative-suite interface.
Look elsewhere if: self-hosting or downloadable model weights are the actual reason you're interested in an "open" video model (MiniMax H3 fits that need instead), your priority is image-to-video specifically, or you want a mature, full-featured creative production platform rather than an API-first model.
Expert Editorial Opinion
The genuinely strong part of this story is the benchmark result itself. Leading two Artificial Analysis categories within weeks of release, in a field this competitive, isn't a fluke — and the document-and-webpage input capability is a real, distinctive feature rather than a marketing footnote, giving Wan 3.0 an editorial and advertising use case most pure text-to-video tools don't cover as directly.
The open-weights confusion is worth dwelling on, because it's a useful case study in how a reputation can outlive the product that earned it. Wan's earlier generations built real goodwill in the open-source AI community, and that goodwill appears to be carrying forward onto Wan 3.0 in public conversation even though the evidence — an Apache-2.0 repository with no actual model files, and Artificial Analysis crediting a different model as the open-weight category leader — doesn't support it. That's not necessarily a deliberate deception on Alibaba's part; it may simply be a licensing choice on a repository that hasn't caught up to the marketing narrative around it. Either way, a reader specifically drawn to Wan because of its open-source history should know that Wan 3.0 itself doesn't currently continue that pattern.
The #6 ranking in image-to-video is the detail most likely to get lost in "Wan 3.0 hits #1" coverage, and it's worth keeping in view: this is a model with a genuinely strong, narrow lead — text-to-video and editing, specifically — not a uniform leader across every video-generation task. That's a more useful way to evaluate it than either the hype framing or a dismissive one.
Final Verdict
Wan 3.0 earns real credit for a genuine benchmark-leading result in two categories and a distinctive, practically useful multimodal input range, including document and webpage support most competitors don't offer. The score reflects real limitations rather than hype: a materially lower ranking in image-to-video, an early and still-small independent review pool, reduced flexibility from being API-only despite an Apache-2.0-branded repository, and a positioning gap between its "open" reputation and its actual access model that a careful buyer needs to know about upfront.
| Dimension | Weight | Score | Why |
|---|---|---|---|
| Technical quality | 30% | 8.5/10 | Genuinely #1 in two Artificial Analysis categories with a real capability jump (30-second, native audio); early reports note inconsistency in complex scenes and dialogue |
| Price-to-value | 25% | 7.5/10 | Transparent per-second list pricing is reasonable for the capability offered; the temporary 30% discount shouldn't be read as the ongoing baseline |
| Maturity & documentation | 20% | 7.0/10 | Docs are current (updated September 17, 2026) but the model itself is barely a month old with a still-small independent review pool |
| Ceiling & flexibility | 15% | 6.0/10 | API-only access with no verified self-hosting option limits the ceiling compared to genuinely open-weight alternatives |
| Honesty of positioning | 10% | 5.5/10 | An Apache-2.0-licensed repository with no downloadable weights, while Artificial Analysis credits a different model as the open-weight leader, creates real, avoidable confusion about what "open" means here |
Weighted total: (8.5×0.30) + (7.5×0.25) + (7.0×0.20) + (6.0×0.15) + (5.5×0.10) = 7.30/10 — Score band: 7.0–7.9, "Competent but compromised: right for a specific use case, not the self-hosting story circulating about it."
❓ Frequently Asked Questions
🔗 Related ToolRadar Reviews
Official documentation and current pricing: Alibaba Cloud Model Studio. Given how new this model is and the temporary promotional pricing in effect, verify current rates and check for updated open-weight status directly before building a workflow around it.

Comments
Post a Comment