Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis
Willi MenapaceAliaksandr SiarohinIvan SkorokhodovEkaterina DeynekaTsai-Shien ChenAnil KagYuwei FangAleksei StoliarElisa RicciJian Ren
Presents Snap Video, a video-first diffusion model that replaces standard U-Nets with a scalable spatiotemporal transformer architecture, achieving up to 4.5 times faster inference alongside state-of-the-art motion quality and temporal consistency in text-to-video generation.
Creating realistic, high-quality video directly from text prompts has become a critical frontier in generative artificial intelligence. Most current video generation methods simply adapt image-based diffusion models by adding temporal layers to standard neural network backbones known as U-Nets. However, video content exhibits heavy spatial and temporal repetition across consecutive frames. Treating video as a sequence of separate images requires redundant computation for every frame, limits model scalability, and frequently generates static images or severe motion artifacts rather than coherent, dynamic motion.
The main objective of the article is to introduce and evaluate Snap Video, a scalable text-to-video architecture designed specifically to address the computational bottlenecks and motion quality limitations of traditional image-derived video models.
The authors designed a two-stage cascaded generation system using a transformer architecture based on Far-reaching Interleaved Transformers (FIT) scaled up to 3.9 billion parameters. This model learns a compact, compressed representation of video data and performs joint spatial and temporal calculations simultaneously rather than sequentially. The authors also reformulated the mathematical diffusion framework (EDM) by introducing an input scaling factor to preserve optimal signal-to-noise ratios across video frames and treated static images as infinite-framerate videos during joint training. The approach was trained on an internal dataset comprising 1.265 million images and 238,000 hours of captioned video, and evaluated on standard benchmarks alongside blinded human user studies.
The evaluation revealed several key findings. First, the proposed transformer architecture achieved a 3.31-fold speedup in training time and a 4.49-fold speedup during inference compared to standard U-Net architectures, allowing the model to scale efficiently into billions of parameters. Second, the model set state-of-the-art visual quality benchmarks on standard zero-shot video datasets, UCF-101 and MSR-VTT. Third, in human preference studies on dynamic scenes, participants favored the proposed model over leading public and proprietary systems for prompt alignment in 80% to 81% of cases, for motion quantity in 88% to 96% of cases, and for motion quality in 70% to 79% of cases, while matching top commercial models in photorealism.
These findings demonstrate that treating video as an inherently distinct, compressible modality yields substantial operational and quality advantages. By drastically cutting inference latency and training overhead while scaling parameter capacity, the architecture lowers compute infrastructure costs and shortens generation turnaround times. It also solves the persistent industry challenge of generating believable, large-scale motion without flickering or temporal degradation.
For technical and product leaders evaluating text-to-video capabilities, the primary recommendation is to pivot away from separable U-Net pipelines toward compressed transformer-based architectures and video-calibrated noise schedules for high-resolution synthesis. Organizations looking to adopt the approach should consider implementing cascaded architectures that divide motion synthesis from high-resolution detail generation.
A primary limitation of the study is its reliance on a proprietary training dataset of 238,000 video hours, which includes synthetic captions generated by an internal model. While the quantitative benchmark gains and user evaluations provide strong confidence in the architecture's scalability and motion fidelity, external deployment across open-domain prompts should account for potential data-dependent edge cases and the computational demands of multi-stage inference.
- Paper: Video Diffusion Models, Jonathan Ho et al. (2022). It introduces foundational space-time factorized 3D U-Nets and joint image-video training for diffusion models, establishing the standard paradigm that Snap Video seeks to replace with scalable transformers.
- Paper: Align Your Latents: High-Resolution Video Synthesis with Latent Diffusion Models, Andreas Blattmann et al. (2023). It establishes the canonical approach of inserting temporal layers into latent image diffusion models, which Snap Video directly contrasts against its unified compressed transformer architecture.
- Paper: Imagen Video: High Definition Video Generation with Diffusion Models, Jonathan Ho et al. (2022). It demonstrates high-definition cascaded video diffusion architectures, providing direct foundational context for Snap Video's multi-stage cascaded synthesis pipeline.
- Paper: ViViT: A Video Vision Transformer, Anurag Arnab et al. (2021). It establishes foundational spatial-temporal factorized transformer architectures on video token sequences that precede and inform modern video transformer backbones.
- Paper: Make-A-Video: Text-to-Video Generation without Text-Video Data, Uriel Singer et al. (2023). It provides a key benchmark baseline for leveraging pretrained text-to-image models with factorized temporal layers without needing paired video datasets.
- Paper: Phenaki: Variable Length Video Generation From Open Domain Textual Description, Ruben Villegas et al. (2023). It introduces transformer-based video tokenization and variable-length text-to-video generation, providing valuable background on non-U-Net video synthesis.
- Paper: Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets, Andreas Blattmann et al. (2023). It details large-scale dataset filtering and training stages for latent video diffusion models, contextualizing the training data recipes utilized in modern video foundation models.
- Paper: HunyuanVideo: A Systematic Framework For Large Video Generative Models, Weijie Kong et al. (2024). It scales diffusion transformer architectures for open-source video generation to over 13 billion parameters, directly expanding upon the large-scale transformer video synthesis explored in Snap Video.
- Paper: Wan: Open and Advanced Large-Scale Video Generative Models, Ang Wang et al. (2025). It generalizes large-scale video diffusion transformers and custom 3D VAEs to a comprehensive suite supporting editing, multi-modal control, and real-time generation.
- Paper: CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer, Zhuoyi Yang et al. (2025). It continues the evolution of transformer-based video diffusion models by introducing expert cross-modal transformer blocks paired with 3D spatiotemporal video compression.
- Paper: Tora: Trajectory-oriented Diffusion Transformer for Video Generation, Zhenghao Zhang et al. (2025). It extends video diffusion transformer architectures to incorporate trajectory-based fine-grained motion control.
- Paper: StreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from Text, Roberto Henschel et al. (2025). It addresses long-form autoregressive video extension and memory preservation, moving beyond the short-clip generation setting of base video models.
- Paper: T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation, Kaiyue Sun 0001 et al. (2025). It introduces a comprehensive compositional benchmark to rigorously evaluate advanced text-to-video systems like Snap Video across multi-object dynamic actions and attribute bindings.
- Paper: MotionClone: Training-Free Motion Cloning for Controllable Video Generation, Pengyang Ling et al. (2025). It develops a training-free motion cloning approach using temporal attention guidance, offering downstream motion control capabilities for video diffusion models.
- Paper: Position: Video as the New Language for Real-World Decision Making, Sherry Yang et al. (2024). It expands the conceptual horizon of video generation architectures by proposing generative video models as foundation simulators and decision-making agents.
