Press ESC or click to close
Latest
Loading latest reviews…

The AI Video Model That Hit #1 — With One Big Catch

Mahmoud Salamoun · September 20, 2026 · 5 min read
The AI Video Model That Hit #1 — With One Big Catch
AI Video Generation Multimodal Video Model

The AI Video Model That Just Hit #1 — With One Big Catch About "Open Weights"

Wan 3.0 Review 2026: Alibaba's new model leads Artificial Analysis in two video categories with 30-second, native-audio generation from text, images, documents, or a webpage URL. The "open source" claim making the rounds online, though, doesn't hold up under a closer look.

7.3/ 10 · 10 min read
📋 Technical Desk Review — built from Alibaba's official documentation and pricing pages, Artificial Analysis benchmark data, and independently verified third-party testing. No hands-on generation claimed. Last verified: September 18, 2026
📋 Table of Contents
  1. What Wan 3.0 Actually Does
  2. Pricing
  3. Is It Actually #1? What Artificial Analysis Really Shows
  4. The Open-Weights Controversy
  5. Pros & Cons
  6. Wan 3.0 vs MiniMax H3 vs Gemini Omni Flash vs Seedance
  7. Who It's For
  8. FAQ

Wan 3.0 landed on Artificial Analysis's leaderboards in August 2026 and immediately took the top spot in text-to-video with audio — a genuinely notable result in a field crowded with well-funded competitors from Google, ByteDance, and OpenAI's former Sora team. What makes it more than just another leaderboard entry is the scope of what it accepts as input: text, an image, a reference video, up to 20 mixed reference assets, and — unusually for this category — a PDF, spreadsheet, presentation, or public webpage URL, which the model can use as creative source material for a 30-second audiovisual sequence.

The story that's spread alongside the benchmark result is that Wan 3.0 is open-weight — a natural assumption, since earlier Wan models had a real open-source reputation. That part of the story doesn't hold up under a closer look at Alibaba's own repository, and it's worth untangling before anything else in this review, because it changes who Wan 3.0 is actually for.

MS
Mahmoud Salamoun
Founder, ToolRadar · Reviewed September 18, 2026
Independent AI tools reviewer with a background in marketing and content, and hands-on daily experience directing AI tools like Gemini and ChatGPT for real work. This review is based on official Alibaba Cloud documentation, Artificial Analysis benchmark data, and verified third-party testing — not a claim of hands-on generation with Wan 3.0 unless stated otherwise.

What Wan 3.0 Actually Does

01

Multimodal input, one model

Text, image, first-frame, first-and-last-frame control, reference video, and reference audio all feed into the same model, alongside document and webpage input — up to 20 mixed reference assets per generation, depending on input mode.

02

30-second native-audio generation

Officially documented at 2 to 30 seconds, at 480P, 720P, or 1080P, 30fps, with dialogue, background music, and environmental sound effects generated alongside the video — audio is on by default and can be disabled.

03

Documents and webpages as source material

The model accepts PDF, DOC/DOCX, XLS/XLSX, PPT/PPTX, TXT, Markdown, and public webpage URLs as reference input — a genuinely unusual capability that goes well beyond a standard text-to-video model's scope.

04

Editing and extension

Beyond fresh generation, Wan 3.0 supports editing existing video and extending a clip's length, positioning it as a working tool across a production pipeline rather than a one-shot generator.

Pricing

ResolutionOfficial list rate30-second cost
480P$0.068/second≈$2.04
720P$0.14/second≈$4.20
1080P$0.28/second≈$8.40

Those are Alibaba's standard international rates through Model Studio, checked directly against official documentation. A temporary 30% promotional discount is live through September 24, 2026, cutting the 720P 30-second cost to roughly $2.10 — worth knowing about if you're pricing a project this week, but not a rate to expect as the ongoing baseline once the promotion ends. There's no unlimited free API tier; Model Studio offers access to try the interface, but generation itself is metered.

💡 Worth knowing: API-generated video URLs stay valid for only 24 hours, so output needs to be downloaded promptly rather than treated as permanent storage. Also worth knowing: the Model Studio console doesn't let you disable the watermark, but the API does support a watermark parameter that can be set to false.

Is It Actually #1? What Artificial Analysis Really Shows

Yes, with real qualification worth understanding rather than skipping past. Artificial Analysis currently ranks Wan 3.0 first in text-to-video with audio, at an Elo of 1243 based on 5,771 blind user-preference votes — but its confidence interval (±9) overlaps with second-place Gemini Omni Flash (1238, ±6), which is why Artificial Analysis itself places both in a shared 1-2 rank range rather than declaring a clean, uncontested winner. The methodology is blind pairwise preference, not a lab benchmark measuring objective quality — real, useful signal, but a different kind of claim than "objectively the best."

The more interesting, more honest story is that Wan 3.0 doesn't lead everywhere. It's also first in video editing with audio (1190 Elo) — but it sits sixth in image-to-video with audio (1178 Elo), well behind the category leaders there. That's worth knowing before assuming "#1" means uniformly best: Wan's strength is concentrated in from-scratch generation and editing, not every video-related task Artificial Analysis tracks.

The Open-Weights Controversy

This is the detail most worth getting right before publishing anything about Wan 3.0. Earlier Wan models had a genuine open-weight reputation, and there is an official-looking GitHub repository — AlibabaCloud-Official/Wan3.0 — carrying an Apache-2.0 license and README language describing it as open-source. That repository, as of this review, contains only a README and license file: no model weights, no releases, and just 11 stars and 2 forks, which is not the footprint of an actively-used open-weight release.

More tellingly, Artificial Analysis's own open-weight category leader for text-to-video with audio is MiniMax H3, not Wan 3.0 — independent analysis describes Wan 3.0 as API-only, with Wan 2.2 remaining the last Wan generation with actually downloadable weights. The honest summary: there's a real Apache-2.0 Wan3.0 repository, but no verified evidence that Wan 3.0's model weights are publicly downloadable for self-hosting. Anyone specifically interested in Wan for self-hosting reasons should not assume that's currently possible.

Pros & Cons

✓ Strengths

  • ✅ Genuinely leads two major Artificial Analysis categories — text-to-video with audio and video editing with audio
  • ✅ 30-second generation with native dialogue, music, and sound effects in one pass is a real capability jump over shorter-clip competitors
  • ✅ Document and webpage input is a distinctive, practically useful capability most video models don't offer
  • ✅ Official, transparent per-second pricing with no opaque quote-only tier

✗ Weaknesses

  • ❌ Despite an Apache-2.0-licensed GitHub repository, no downloadable model weights are currently available — not the self-hostable model some coverage implies
  • ❌ Ranks sixth in image-to-video with audio, well behind category leaders there
  • ❌ Early independent testing reports inconsistent character consistency, imperfect on-screen text, and occasional painterly-looking backgrounds in complex scenes
  • ❌ Very new (August 2026 release) with a still-small independent review pool

Wan 3.0 vs MiniMax H3 vs Gemini Omni Flash vs Seedance

ModelText-to-video Elo (with audio)Open weightsBest for
Wan 3.01243No (API only)Longest single-pass generation, document/webpage input
Gemini Omni Flash1238NoGoogle multimodal ecosystem integration
MiniMax H31225YesTeams that specifically need self-hostable, downloadable weights
Seedance 2.0 (720p)1220NoCinematic, reference-driven creative workflows

The Elo gaps between these four are narrow enough that confidence intervals overlap across most of them — this is a genuinely competitive top tier, not a single runaway leader. The clearer differentiator for most readers is access model rather than raw score: MiniMax H3 is the pick if downloadable, self-hostable weights are the actual requirement, since it's the one Artificial Analysis itself credits as leading the open-weight category. One independent platform test reported Wan 3.0 generating a 30-second 720p clip with audio at meaningfully lower platform-credit cost than Seedance 2.5 — a real data point, though it reflects one provider's pricing rather than a universal comparison across every platform hosting these models.

Who It's For

Choose Wan 3.0 if: you need longer single-pass generations with native audio, want to turn documents or reference webpages into video content, and are comfortable working through an API rather than a polished creative-suite interface.

Look elsewhere if: self-hosting or downloadable model weights are the actual reason you're interested in an "open" video model (MiniMax H3 fits that need instead), your priority is image-to-video specifically, or you want a mature, full-featured creative production platform rather than an API-first model.

Expert Editorial Opinion

The genuinely strong part of this story is the benchmark result itself. Leading two Artificial Analysis categories within weeks of release, in a field this competitive, isn't a fluke — and the document-and-webpage input capability is a real, distinctive feature rather than a marketing footnote, giving Wan 3.0 an editorial and advertising use case most pure text-to-video tools don't cover as directly.

The open-weights confusion is worth dwelling on, because it's a useful case study in how a reputation can outlive the product that earned it. Wan's earlier generations built real goodwill in the open-source AI community, and that goodwill appears to be carrying forward onto Wan 3.0 in public conversation even though the evidence — an Apache-2.0 repository with no actual model files, and Artificial Analysis crediting a different model as the open-weight category leader — doesn't support it. That's not necessarily a deliberate deception on Alibaba's part; it may simply be a licensing choice on a repository that hasn't caught up to the marketing narrative around it. Either way, a reader specifically drawn to Wan because of its open-source history should know that Wan 3.0 itself doesn't currently continue that pattern.

The #6 ranking in image-to-video is the detail most likely to get lost in "Wan 3.0 hits #1" coverage, and it's worth keeping in view: this is a model with a genuinely strong, narrow lead — text-to-video and editing, specifically — not a uniform leader across every video-generation task. That's a more useful way to evaluate it than either the hype framing or a dismissive one.

Final Verdict

ToolRadar Performance Score
7.3 / 10

Wan 3.0 earns real credit for a genuine benchmark-leading result in two categories and a distinctive, practically useful multimodal input range, including document and webpage support most competitors don't offer. The score reflects real limitations rather than hype: a materially lower ranking in image-to-video, an early and still-small independent review pool, reduced flexibility from being API-only despite an Apache-2.0-branded repository, and a positioning gap between its "open" reputation and its actual access model that a careful buyer needs to know about upfront.

DimensionWeightScoreWhy
Technical quality30%8.5/10Genuinely #1 in two Artificial Analysis categories with a real capability jump (30-second, native audio); early reports note inconsistency in complex scenes and dialogue
Price-to-value25%7.5/10Transparent per-second list pricing is reasonable for the capability offered; the temporary 30% discount shouldn't be read as the ongoing baseline
Maturity & documentation20%7.0/10Docs are current (updated September 17, 2026) but the model itself is barely a month old with a still-small independent review pool
Ceiling & flexibility15%6.0/10API-only access with no verified self-hosting option limits the ceiling compared to genuinely open-weight alternatives
Honesty of positioning10%5.5/10An Apache-2.0-licensed repository with no downloadable weights, while Artificial Analysis credits a different model as the open-weight leader, creates real, avoidable confusion about what "open" means here

Weighted total: (8.5×0.30) + (7.5×0.25) + (7.0×0.20) + (6.0×0.15) + (5.5×0.10) = 7.30/10 — Score band: 7.0–7.9, "Competent but compromised: right for a specific use case, not the self-hosting story circulating about it."

❓ Frequently Asked Questions

Not currently verified. Alibaba maintains an Apache-2.0-licensed GitHub repository for Wan3.0, but as of this review it contains no downloadable model weights or releases. Artificial Analysis credits MiniMax H3, not Wan 3.0, as the leading open-weight text-to-video model — Wan 2.2 remains the last Wan generation with actually downloadable weights.
It leads Artificial Analysis's text-to-video-with-audio and video-editing-with-audio leaderboards, but its confidence interval overlaps with second-place Gemini Omni Flash, and Artificial Analysis itself places both in a shared 1-2 rank range. It also ranks sixth in image-to-video with audio, so "#1 at everything" isn't accurate.
Official list pricing is $0.068/second at 480P, $0.14/second at 720P, and $0.28/second at 1080P — roughly $4.20 for a 30-second 720P clip. A temporary 30% discount is live through September 24, 2026; there's no unlimited free API tier.
Yes, audio — including dialogue, background music, and environmental sound effects — is generated by default alongside the video, and can be disabled if not needed.
Yes. Alibaba's documentation confirms support for PDF, DOC/DOCX, XLS/XLSX, PPT/PPTX, TXT, Markdown, and public webpage URLs as reference input the model can draw on when generating a video — a genuinely unusual capability for this category.

Official documentation and current pricing: Alibaba Cloud Model Studio. Given how new this model is and the temporary promotional pricing in effect, verify current rates and check for updated open-weight status directly before building a workflow around it.

Share this review
Mahmoud Salamoun
Written by
Mahmoud Salamoun
Independent AI tools reviewer based in the Middle East. I test and rate AI tools so you don't have to — no sponsorships, no bias, just honest analysis.
Rate this review
(-/5)

Comments