Wan: Open and Advanced Large-Scale Video Generative Models

Ang WangBaole AiBin WenChaojie MaoChen-Wei XieDi ChenFeiwu YuHaiming ZhaoJianxiao YangJianyuan Zeng

article2025arXiv2,532 citations

Presents an open-source suite of diffusion transformer video foundation models scaling up to 14 billion parameters that outperforms leading commercial systems while providing a lightweight variant that runs on consumer GPUs with under 8.2 GB of VRAM.

Listen

Recent advancements in artificial intelligence have accelerated the development of video generation models, yet a significant divide remains between high-performing proprietary commercial solutions and open-source alternatives. Existing open-source systems often suffer from inferior visual quality, limited functionality beyond basic text-to-video generation, and heavy computational demands that exclude organizations with modest hardware resources. The article introduces and evaluates Wan, an open-source suite of large-scale video foundation models designed to overcome these performance, versatility, and efficiency barriers.

To establish a highly capable and accessible framework, the developers built Wan upon a diffusion transformer architecture optimized with flow matching and a custom spatio-temporal variational autoencoder. The models were trained on billions of carefully filtered images and videos totaling trillions of tokens, using dense captioning pipelines and progressive multi-stage training curricula. The suite provides a flagship 14-billion-parameter model designed for maximum generation quality and a lightweight 1.3-billion-parameter model tailored for low-resource environments, alongside specialized adaptations for image-to-video, unified editing, personalization, camera control, real-time streaming, and synchronized audio generation.

Empirical evaluations across standardized benchmarks and human assessments indicate that Wan sets a new standard for open video generative models. The 14-billion-parameter model achieved an aggregate score of 86.22% on the public VBench leaderboard, outperforming leading commercial offerings such as OpenAI's Sora (84.28%) and Runway Gen-3 (82.32%), while winning the majority of direct human preference comparisons across visual quality, motion quality, and text alignment. Additionally, the custom autoencoder achieved video reconstruction speeds 2.5 times faster than current state-of-the-art baselines, and the 1.3-billion-parameter model demonstrated that high-quality generation is feasible on consumer-grade hardware, requiring only 8.19 gigabytes of memory while outperforming several larger open-source competitors.

These findings indicate that organizations no longer need to depend exclusively on closed, costly proprietary APIs for professional-grade video synthesis. The release of open model weights, training pipelines, and evaluation tooling significantly lowers development costs, accelerates production timelines, and mitigates vendor lock-in risks for enterprises and research institutions. Furthermore, the architecture’s bilingual visual text rendering and downstream versatility allow teams to consolidate multiple media workflows into a unified foundation.

Decision-makers and development teams should evaluate adopting or piloting Wan for automated media generation, content localization, and creative prototyping workflows. For resource-constrained or edge deployments, teams should leverage the 1.3-billion model with recommended quantization techniques, reserving the 14-billion model for high-fidelity production pipelines. Future initiatives should focus on addressing remaining model limitations, specifically improving fine-grained visual details during extreme motion dynamics, reducing raw inference latency for the largest variant, and fine-tuning domain-specific knowledge for specialized applications such as medicine and education.

  • Paper: Advancing Open-Source World Models, Robbyant Team et al. (2026). LingBot-World advances open-source video foundation models like Wan into interactive, action-conditioned physical world simulators.
  • Paper: Qwen-Image-VAE-2.0 Technical Report, Zekai Zhang et al. (2026). Qwen-Image-VAE-2.0 extends visual compression research by introducing higher-compression autoencoders designed to further optimize downstream generative models.
  • Paper: Qwen3.5-Omni Technical Report, Qwen Team (2026). Qwen3.5-Omni integrates generative and comprehension capabilities across video, speech, and text into a unified omnimodal interactive agent architecture.
Cover for Wan: Open and Advanced Large-Scale Video Generative Models

Abstract

This report presents Wan, a comprehensive and open suite of video foundation models designed to push the boundaries of video generation. Built upon the mainstream diffusion transformer paradigm, Wan achieves significant advancements in generative capabilities through a series of innovations, including our novel VAE, scalable pre-training strategies, large-scale data curation, and automated evaluation metrics. These contributions collectively enhance the model's performance and versatility. Specifically, Wan is characterized by four key features: Leading Performance: The 14B model of Wan, trained on a vast dataset comprising billions of images and videos, demonstrates the scaling laws of video generation with respect to both data and model size. It consistently outperforms the existing open-source models as well as state-of-the-art commercial solutions across multiple internal and external benchmarks, demonstrating a clear and significant performance superiority. Comprehensiveness: Wan offers two capable models, i.e., 1.3B and 14B parameters, for efficiency and effectiveness respectively. It also covers multiple downstream applications, including image-to-video, instruction-guided video editing, and personal video generation, encompassing up to eight tasks. Consumer-Grade Efficiency: The 1.3B model demonstrates exceptional resource efficiency, requiring only 8.19 GB VRAM, making it compatible with a wide range of consumer-grade GPUs. Openness: We open-source the entire series of Wan, including source code and all models, with the goal of fostering the growth of the video generation community. This openness seeks to significantly expand the creative possibilities of video production in the industry and provide academia with high-quality video foundation models. All the code and models are available at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Data Processing Pipeline
  • 3.1 Pre-training Data
  • 3.2 Post-training Data
  • 3.3 Dense Video Caption
  • 3.3.1 Open Source Dataset
  • 3.3.2 In-house Dataset
  • 3.3.3 Model Design
  • 3.3.4 Evaluation
  • 4 Model Design and Acceleration
  • 4.1 Spatio-temporal Variational Autoencoder
  • 4.1.1 Model Design
  • 4.1.2 Training
  • 4.1.3 Efficient Inference
  • 4.1.4 Evaluation
  • 4.2 Model Training
  • 4.2.1 Video Diffusion Transformer
  • 4.2.2 Pre-training
  • 4.2.3 Post-training
  • 4.3 Model Scaling and Training Efficiency
  • 4.3.1 Workload Analysis
  • 4.3.2 Parallelism Strategy
  • 4.3.3 Memory Optimization
  • 4.3.4 Cluster Reliability
  • 4.4 Inference
  • 4.4.1 Parallel Strategy
  • 4.4.2 Diffusion Cache
  • 4.4.3 Quantization
  • 4.5 Prompt Alignment
  • 4.6 Benchmarks
  • 4.7 Evaluation
  • 4.7.1 Metrics and Results
  • 4.7.2 Ablation Study
  • 5 Extended Applications
  • 5.1 Image-to-Video Generation
  • 5.1.1 Model Design
  • 5.1.2 Dataset
  • 5.1.3 Evaluation
  • 5.2 Unified Video Editing
  • 5.2.1 Model Design
  • 5.2.2 Datasets and Implementation
  • 5.2.3 Evaluation
  • 5.3 Text-to-Image Generation
  • 5.4 Video Personalization
  • 5.4.1 Model Design
  • 5.4.2 Dataset
  • 5.4.3 Evaluation
  • 5.5 Camera Motion Controllability
  • 5.6 Real-time Video Generation
  • 5.6.1 Method
  • 5.6.2 Streaming Video Generation
  • 5.6.3 Consistency Model Distillation
  • 5.7 Audio Generation
  • 5.7.1 Model Design
  • 5.7.2 Evaluation
  • 6 Limitation and Conclusion
  • 7 Contributors
  • References

Citation

MLA
Wan, T., et al. “Wan: Open and Advanced Large-Scale Video Generative Models”. arXiv, 2025, http://arxiv.org/abs/2503.20314v2.
APA
Wan, T., Wang, A., Ai, B., Wen, B., Mao, C., Xie, C.-W., Chen, D., Yu, F., Zhao, H., Yang, J., Zeng, J., Wang, J., Zhang, J., Zhou, J., Wang, J., Chen, J., Zhu, K., Zhao, K., Yan, K., … Liu, Z. (2025). Wan: Open and Advanced Large-Scale Video Generative Models. arXiv. http://arxiv.org/abs/2503.20314v2
Chicago
Wan, T., A. Wang, B. Ai, et al. 2025. “Wan: Open and Advanced Large-Scale Video Generative Models”. arXiv. http://arxiv.org/abs/2503.20314v2.
Harvard
Wan, T. et al. (2025) “Wan: Open and Advanced Large-Scale Video Generative Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2503.20314v2.
Vancouver
1. Wan T, Wang A, Ai B, et al (2025) Wan: Open and Advanced Large-Scale Video Generative Models. arXiv

BibTeX

@article{wan2025wan,
  title = {Wan: Open and Advanced Large-Scale Video Generative Models},
  author = {Wan, Team and Wang, Ang and Ai, Baole and Wen, Bin and Mao, Chaojie and Xie, Chen-Wei and Chen, Di and Yu, Feiwu and Zhao, Haiming and Yang, Jianxiao and Zeng, Jianyuan and Wang, Jiayu and Zhang, Jingfeng and Zhou, Jingren and Wang, Jinkai and Chen, Jixuan and Zhu, Kai and Zhao, Kang and Yan, Keyu and Huang, Lianghua and Feng, Mengyang and Zhang, Ningyi and Li, Pandeng and Wu, Pingyu and Chu, Ruihang and Feng, Ruili and Zhang, Shiwei and Sun, Siyang and Fang, Tao and Wang, Tianxing and Gui, Tianyi and Weng, Tingyu and Shen, Tong and Lin, Wei and Wang, Wei and Wang, Wei and Zhou, Wenmeng and Wang, Wente and Shen, Wenting and Yu, Wenyuan and Shi, Xianzhong and Huang, Xiaoming and Xu, Xin and Kou, Yan and Lv, Yangyu and Li, Yifei and Liu, Yijing and Wang, Yiming and Zhang, Yingya and Huang, Yitong and Li, Yong and Wu, You and Liu, Yu and Pan, Yulin and Zheng, Yun and Hong, Yuntao and Shi, Yupeng and Feng, Yutong and Jiang, Zeyinzi and Han, Zhen and Wu, Zhi-Fan and Liu, Ziyu},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2503.20314v2},
  eprint = {2503.20314}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors