
Alibaba's Wan3.0 generates 30-second videos from text, images, and documents
Alibaba's Wan3.0 is in beta, doubling Wan2.5's ceiling to 30-second videos and accepting up to 10 images, 5 videos, 5 audio clips, plus PDFs, web pages, and PowerPoint files in a single prompt. Alibaba claims tighter consistency on characters, props, and spatial layouts — the specific fix for the facial and interface drift that makes generated clips unusable — at $0.05 to $0.28 per second via wan.video, Alibaba Cloud Model Studio, and the Qwen Cloud API. The launch lands the same week Alibaba reported a 75% year-over-year profit drop from AI spending and announced the largest share sale by a Hong Kong-listed company, marking video generation as a volume business it intends to buy share in.
Source: the-decoder.com ↗
AI-generated videos often suffer from visual drift and distortion, especially in faces and user interfaces. Wan3.0 aims to fix that by keeping details from reference material like characters, props, and spatial layouts more consistent.
Why this matters
- → Solves the visual drift problem that makes most AI video unusable for production
- → Multimodal input (text, images, video, audio, PDFs) in one prompt cuts production friction
- → Signals video generation shifting from novelty to volume business—Alibaba betting billions despite 75% profit