aiminute. ← All AI news
New Models AI Minute Newsroom 2026-08-24

Alibaba's new video model takes your slide deck as the input and hands back thirty seconds of film

Alibaba's new video model takes your slide deck as the input and hands back thirty seconds of film

Alibaba Cloud brought Wan3.0 out of beta on Monday 24 August, announcing it on WeChat. What it is leading with is the input side rather than the output: alongside text, images, audio and video, the model accepts documents, spreadsheets, slide decks, PDFs and web pages given by URL, and builds a video sequence from them. Alibaba calls the feature Omni-Reference. Clips run up to 30 seconds, double the 15-second ceiling of Wan2.7-Video, and a single unified model now replaces the previous split line-up. It also proposes a length based on your prompt and can extend a clip you already have. Wan3.0 has been in public beta since 6 August, and Alibaba says it has been used since then in short-drama and film production, advertising and marketing, tourism promotion and music videos. It is served through Alibaba Cloud Model Studio and Qwen Cloud as wan3.0-video; the weights have not been published and API access is not yet fully open. The launch landed on the same day as Alibaba's $10 billion share placement — the largest primary follow-on ever by a Hong Kong-listed company — raised to pay for AI.

Why it mattersNearly every video model you can name starts from a sentence or a picture. Pointing one at a PDF or a quarterly deck changes who it is for. The person making a training video, a product explainer or a tourism spot already has the content; it is just in the wrong medium, and inventing the images was never the hard part. Thirty seconds also matters more than it sounds: it is roughly where a clip stops being a demo and starts being something you could put in front of a customer. The catch is the familiar one with Chinese frontier releases this year — the capability is announced, the weights are not, and access runs through a cloud you have to sign in to.
#Video

✓ Verified · 3 sources

WhatsApp X Telegram
Read in the app — free, in 9 languages

Related stories

Alibaba trained a model on a rack of 100 real phones — it now finishes 92 out of 100 tasks on an actual handset
2026-08-23
An eight-billion-parameter model that reads pictures and draws them, at 4K, with an Apache licence — and the community had it running in two days
2026-08-22
Google's open models have been downloaded a billion times, and outsiders have built 100,000 versions of them
2026-08-22
DeepSeek's cheap workhorse can now see — and on agent tasks that need eyes it says it is close to Anthropic's best
2026-08-21
Adobe will now generate the music, the voiceover and the door slam — and it says the licence covers you
2026-08-21