Wan: Open and Advanced Large-Scale Video Generative Models
Ang WangBaole AiBin WenChaojie MaoChen-Wei XieDi ChenFeiwu YuHaiming ZhaoJianxiao YangJianyuan Zeng
Presents an open-source suite of diffusion transformer video foundation models scaling up to 14 billion parameters that outperforms leading commercial systems while providing a lightweight variant that runs on consumer GPUs with under 8.2 GB of VRAM.
Recent advancements in artificial intelligence have accelerated the development of video generation models, yet a significant divide remains between high-performing proprietary commercial solutions and open-source alternatives. Existing open-source systems often suffer from inferior visual quality, limited functionality beyond basic text-to-video generation, and heavy computational demands that exclude organizations with modest hardware resources. The article introduces and evaluates Wan, an open-source suite of large-scale video foundation models designed to overcome these performance, versatility, and efficiency barriers.
To establish a highly capable and accessible framework, the developers built Wan upon a diffusion transformer architecture optimized with flow matching and a custom spatio-temporal variational autoencoder. The models were trained on billions of carefully filtered images and videos totaling trillions of tokens, using dense captioning pipelines and progressive multi-stage training curricula. The suite provides a flagship 14-billion-parameter model designed for maximum generation quality and a lightweight 1.3-billion-parameter model tailored for low-resource environments, alongside specialized adaptations for image-to-video, unified editing, personalization, camera control, real-time streaming, and synchronized audio generation.
Empirical evaluations across standardized benchmarks and human assessments indicate that Wan sets a new standard for open video generative models. The 14-billion-parameter model achieved an aggregate score of 86.22% on the public VBench leaderboard, outperforming leading commercial offerings such as OpenAI's Sora (84.28%) and Runway Gen-3 (82.32%), while winning the majority of direct human preference comparisons across visual quality, motion quality, and text alignment. Additionally, the custom autoencoder achieved video reconstruction speeds 2.5 times faster than current state-of-the-art baselines, and the 1.3-billion-parameter model demonstrated that high-quality generation is feasible on consumer-grade hardware, requiring only 8.19 gigabytes of memory while outperforming several larger open-source competitors.
These findings indicate that organizations no longer need to depend exclusively on closed, costly proprietary APIs for professional-grade video synthesis. The release of open model weights, training pipelines, and evaluation tooling significantly lowers development costs, accelerates production timelines, and mitigates vendor lock-in risks for enterprises and research institutions. Furthermore, the architecture’s bilingual visual text rendering and downstream versatility allow teams to consolidate multiple media workflows into a unified foundation.
Decision-makers and development teams should evaluate adopting or piloting Wan for automated media generation, content localization, and creative prototyping workflows. For resource-constrained or edge deployments, teams should leverage the 1.3-billion model with recommended quantization techniques, reserving the 14-billion model for high-fidelity production pipelines. Future initiatives should focus on addressing remaining model limitations, specifically improving fine-grained visual details during extreme motion dynamics, reducing raw inference latency for the largest variant, and fine-tuning domain-specific knowledge for specialized applications such as medicine and education.
- Paper: Photorealistic Video Generation with Diffusion Models, Agrim Gupta et al. (2024). W.A.L.T establishes foundational techniques for using diffusion transformers with spatiotemporal compression and windowed attention for photorealistic video generation, which Wan directly builds upon.
- Paper: Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets, Andreas Blattmann et al. (2023). Stable Video Diffusion details critical principles for scalable data curation and multi-stage pre-training for latent video diffusion models that Wan refines and scales up.
- Paper: VBench: Comprehensive Benchmark Suite for Video Generative Models, Ziqi Huang et al. (2023). VBench defines the multi-dimensional benchmark suite and automated video evaluation metrics adopted by Wan to systematically assess temporal and visual generative quality.
- Paper: Scaling Rectified Flow Transformers for High-Resolution Image Synthesis, Patrick Esser et al. (2024). This paper establishes scalable multimodal diffusion transformer architectures (MM-DiT) and flow-matching training paradigms fundamental to large-scale generative models like Wan.
- Paper: Video Diffusion Models, Jonathan Ho et al. (2022). Video Diffusion Models introduced the foundational space-time factorization and joint image-video training recipes that underpin modern video generative systems.
- Paper: Align Your Latents: High-Resolution Video Synthesis with Latent Diffusion Models, Andreas Blattmann et al. (2023). Align Your Latents demonstrates how to extend latent image diffusion architectures to temporal video synthesis via interleaved temporal layers.
- Paper: Lumiere: A Space-Time Diffusion Model for Video Generation, Omer Bar-Tal et al. (2024). Lumiere presents space-time U-Net architectures and holistic temporal representations that inform modern end-to-end video foundation modeling.
- Paper: Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning, Rohit Girdhar et al. (2024). Emu Video illustrates explicit image-to-video factorization paradigms that Wan incorporates across its multi-task downstream video generation capabilities.
- Paper: Advancing Open-Source World Models, Robbyant Team et al. (2026). LingBot-World advances open-source video foundation models like Wan into interactive, action-conditioned physical world simulators.
- Paper: Qwen-Image-VAE-2.0 Technical Report, Zekai Zhang et al. (2026). Qwen-Image-VAE-2.0 extends visual compression research by introducing higher-compression autoencoders designed to further optimize downstream generative models.
- Paper: Qwen3.5-Omni Technical Report, Qwen Team (2026). Qwen3.5-Omni integrates generative and comprehension capabilities across video, speech, and text into a unified omnimodal interactive agent architecture.
