Alibaba rolls out Wan 3.0 with 30-second document-to-video generation
Alibaba has moved Wan 3.0 beyond its earlier beta, bringing 30-second video generation, native audio, multimodal references, and document-to-video workflows to its creation platforms.
Alibaba officially rolled out Wan 3.0 on August 24, extending a public beta that began earlier in the month. The video model can generate clips lasting up to 30 seconds at 1080p, combine visual and audio references, and turn documents, spreadsheets, presentations, or web pages into source material for a video.
The rollout matters because it moves document-to-video generation into a category-leading commercial model rather than treating it as a separate extraction and production workflow. Access and maturity still vary by platform: Alibaba’s Chinese creation services list the model as available, while the international Model Studio page continues to describe API access as an approval-based preview.
Key takeaways
- Wan 3.0 generates videos lasting up to 30 seconds in one pass, with optional native audio and resolutions from 480p to 1080p.
- Inputs can include text, images, video, audio, PDFs, presentations, spreadsheets, Markdown files, and web pages.
- Alibaba has expanded the model across its Wan and Qwen creation products after an August 6 public beta.
- International API users may still need approval, and Alibaba has not announced downloadable Wan 3.0 weights.
Wan 3.0 widens the input pipeline
Earlier video generators generally start from a text prompt, a still image, or a small set of visual references. Wan 3.0 adds documents and web pages to that pipeline. Alibaba’s product material describes workflows that use reports, tables, plans, and presentation decks as creative references alongside conventional media.
That does not mean the model reliably converts every factual document into a finished explainer. It means a creator can supply the source material directly, reducing the need to manually extract its themes, visual elements, and structure before generation. Factual claims, charts, captions, and rendered text still need review against the original file.
Alibaba also says the model can coordinate characters, environments, dialogue, camera direction, and sound from a shared set of references. Official demonstrations include multilingual performances, product-focused social videos, and continuous action sequences. These are vendor-selected examples, not independent measurements of consistency across ordinary prompts.
Access differs by channel
Reuters reported that Alibaba officially rolled out Wan 3.0 on August 24 after operating it in public beta since August 6. Chinese coverage lists access through Alibaba Cloud Model Studio, Wan’s website, Qwen creation products, and several partner platforms.
The international Model Studio release page is more restrictive. It presents an application form for the wan3.0-video API and says full availability is still coming. Developers should therefore check the status attached to their account and region instead of assuming that a visible model page guarantees immediate API access.
This is a hosted release. Despite older Wan generations being associated with open-weight distribution, Alibaba’s current Wan 3.0 materials do not provide downloadable production weights or a self-hosting license.
Pricing and operational limits
Alibaba’s international page lists generation at $0.05 per second for 480p, $0.10 for 720p, and $0.20 for 1080p. At those rates, a 30-second result costs about $1.50, $3, or $6 before retries and discarded generations.
Chinese pricing and temporary promotions use separate rates. ITHome reported a 30% discount on Alibaba Cloud Bailian and Qwen AI platforms from August 24 through September 23. Teams should confirm the applicable billing surface, currency, promotion period, and whether failed or moderated jobs consume credits.
The strongest practical limit is evaluation cost. Long clips are more expensive to regenerate, and errors near the end of a 30-second sequence can waste an otherwise usable result. A sensible test uses short, low-resolution drafts before committing to 1080p output.
Why the release changes video workflows
Thirty seconds is long enough to carry a short product explanation, campaign scene, or narrative beat without stitching several separately generated clips. Document inputs also create a direct path from business material to video drafts, which could affect marketing, training, tourism, and internal communications workflows.
That convenience raises familiar provenance and accuracy risks. Generated people, events, products, charts, and places must be identified as synthetic where viewers could mistake them for evidence. Organizations should also check whether uploaded documents contain confidential material and review Alibaba’s regional processing, retention, and usage terms before submitting them.
The next confirmation point is broad, unapproved API availability outside Alibaba’s Chinese platforms. Independent testing will also need to establish whether Wan 3.0’s long-form consistency and document grounding hold up beyond its official showcase.
Source check
- Alibaba Cloud Model Studio’s Wan 3.0 release page documents the 30-second limit, multimodal inputs, resolutions, pricing, and approval-based API access.
- Alibaba Cloud’s international launch article confirms the earlier public beta and document-input workflow.
- Reuters coverage independently confirms the August 24 rollout and the August 6 beta date.
- The Next Web describes the launch as widened access following the earlier beta.
Try the related loot
Rent Out Your Idle GPU on Vast.ai—The 75% Revenue Share Comes With Real Work
