VideoPoet: A Large Language Model for Zero-Shot Video Generation
Dan KondratyukLijun YuXiuye GuJosé LezamaJonathan HuangGrant SchindlerRachel HornungVighnesh BirodkarJimmy YanMing-Chang Chiu
Demonstrates that a unified decoder-only language model trained on discrete multimodal tokens can match or outperform diffusion approaches across diverse video generation and editing tasks without task-specific architectural changes.
Recent advances in generative artificial intelligence have enabled automated video creation, but the field relies almost entirely on diffusion models that require complex, separate modules or modifications to handle different tasks. In contrast, large language model architectures have shown immense versatility across language, speech, and robotics, yet their application to high-quality video generation has remained largely understudied. The article introduces VideoPoet, a unified decoder-only language model framework designed to evaluate and demonstrate that language models can perform diverse, high-fidelity video generation tasks within a single architecture.
To achieve this, the approach tokenizes multimodal inputs—including text embeddings, discrete visual tokens, and audio tokens—into a shared vocabulary of approximately 300,000 codes. The system leverages a two-stage pretraining and task-adaptation protocol, training an 8-billion-parameter model on roughly two trillion tokens across one billion image-text pairs and 270 million videos. This process incorporates alternating gradient descent across multiple tasks, such as text-to-video, image-to-video, video future prediction, inpainting, outpainting, and stylization. A custom spatial super-resolution transformer then upsamples base outputs to higher visual resolutions.
The article demonstrates several significant findings. First, the unified model achieves state-of-the-art results across standard zero-shot video generation benchmarks, attaining strong performance on datasets such as MSR-VTT and UCF-101. Second, side-by-side human evaluations indicate that the system is highly competitive with leading video diffusion models, earning clear preference for motion realism and interestingness. Third, scaling model capacity to 8 billion parameters substantially enhances temporal consistency, prompt fidelity, spatial reasoning, and counting. Finally, the framework demonstrates flexible task chaining, smoothly combining capabilities like animating a static image into a video and subsequent video stylization without generative degradation.
These findings prove that large language model architectures represent a viable, high-performance alternative to diffusion methods for video generation. Adopting language models allows organizations to leverage mature training infrastructure, hardware optimizations, and unified multitask scaling without maintaining disparate specialized models. The primary trade-off involves computational cost during scaling and runtime. Next steps include exploring additional acceleration techniques for inference, expanding model capabilities to direct text generation, and implementing governance strategies such as digital watermarking to mitigate ethical risks and deceptive misuse. Limitations include visual fidelity upper bounds set by discrete tokenizers, challenges rendering fine-grained details during rapid motion, and baseline aesthetic differences from excluding certain copyrighted datasets.
- Paper: Phenaki: Variable Length Video Generation From Open Domain Textual Description, Ruben Villegas et al. (2023). Phenaki introduced autoregressive token-based video generation from text sequences using spatiotemporal visual tokenizers, directly establishing the foundation for VideoPoet's decoder-only visual token modeling.
- Paper: VideoGPT: Video Generation using VQ-VAE and Transformers, Wilson Yan et al. (2021). VideoGPT established the core paradigm of compressing video into discrete VQ-VAE tokens and generating them sequentially using transformer language models.
- Paper: Scaling Autoregressive Models for Content-Rich Text-to-Image Generation, Jiahui Yu et al. (2022). Parti demonstrated that scaling autoregressive sequence-to-sequence language models over discrete visual tokens achieves competitive generative quality compared to diffusion models, inspiring VideoPoet's LLM-driven video synthesis.
- Paper: Zero-Shot Text-to-Image Generation, Aditya Ramesh et al. (2021). This work pioneered zero-shot visual generation using large autoregressive transformers over discrete image tokens, establishing the core framework extended to video in VideoPoet.
- Paper: Video Diffusion Models, Jonathan Ho et al. (2022). Provides the foundational video generation benchmarks, conditioning setups, and spatiotemporal modeling concepts that VideoPoet positions its LLM-based architecture against.
- Paper: Imagen Video: High Definition Video Generation with Diffusion Models, Jonathan Ho et al. (2022). Establishes standard text-to-video cascaded architectures and spatial-temporal super-resolution pipelines that VideoPoet adapts into its custom super-resolution transformer.
- Paper: Make-A-Video: Text-to-Video Generation without Text-Video Data, Uriel Singer et al. (2023). Make-A-Video introduces critical zero-shot text-to-video evaluation methodologies and task adaptation strategies that VideoPoet directly benchmarks against.
- Paper: ViViT: A Video Vision Transformer, Anurag Arnab et al. (2021). Introduces spatiotemporal tubelet tokenization and factorized transformer attention mechanisms essential for processing video as discrete sequence tokens in transformer architectures.
- Paper: Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization, Yang Jin et al. (2024). Video-LaVIT extends unified video-language pretraining by decoupling discrete visual keyframes from explicit motion tokens within a next-token prediction language model framework.
- Paper: Show-o: One Single Transformer to Unify Multimodal Understanding and Generation, Jinheng Xie et al. (2025). Show-o further develops unified autoregressive transformer architectures by jointly integrating multimodal understanding with generative token modeling in a single network.
- Paper: Position: Video as the New Language for Real-World Decision Making, Sherry Yang et al. (2024). Extends the concept of autoregressive video token generation into a universal foundation model and decision-making interface for robotics, simulation, and real-world planning.
- Paper: VideoScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video Generation, Xuan He et al. (2024). VideoScore applies fine-grained multimodal vision-language models to systematically evaluate generative video quality, consistency, and prompt fidelity on benchmarks produced by models like VideoPoet.
- Paper: Generative Multimodal Models are In-Context Learners, Quan Sun et al. (2024). Emu2 builds on unified generative sequence modeling by scaling autoregressive multimodal generation and visual in-context learning across interleaved images, videos, and text.
- Paper: Wan: Open and Advanced Large-Scale Video Generative Models, Ang Wang et al. (2025). Wan expands upon large-scale multimodal foundation models for unified video generation, introducing scaled spatio-temporal representations and advanced generative flow architectures.
