Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis

Willi MenapaceAliaksandr SiarohinIvan SkorokhodovEkaterina DeynekaTsai-Shien ChenAnil KagYuwei FangAleksei StoliarElisa RicciJian Ren

article2024CVPR122 citations

Presents Snap Video, a video-first diffusion model that replaces standard U-Nets with a scalable spatiotemporal transformer architecture, achieving up to 4.5 times faster inference alongside state-of-the-art motion quality and temporal consistency in text-to-video generation.

Listen

Creating realistic, high-quality video directly from text prompts has become a critical frontier in generative artificial intelligence. Most current video generation methods simply adapt image-based diffusion models by adding temporal layers to standard neural network backbones known as U-Nets. However, video content exhibits heavy spatial and temporal repetition across consecutive frames. Treating video as a sequence of separate images requires redundant computation for every frame, limits model scalability, and frequently generates static images or severe motion artifacts rather than coherent, dynamic motion.

The main objective of the article is to introduce and evaluate Snap Video, a scalable text-to-video architecture designed specifically to address the computational bottlenecks and motion quality limitations of traditional image-derived video models.

The authors designed a two-stage cascaded generation system using a transformer architecture based on Far-reaching Interleaved Transformers (FIT) scaled up to 3.9 billion parameters. This model learns a compact, compressed representation of video data and performs joint spatial and temporal calculations simultaneously rather than sequentially. The authors also reformulated the mathematical diffusion framework (EDM) by introducing an input scaling factor to preserve optimal signal-to-noise ratios across video frames and treated static images as infinite-framerate videos during joint training. The approach was trained on an internal dataset comprising 1.265 million images and 238,000 hours of captioned video, and evaluated on standard benchmarks alongside blinded human user studies.

The evaluation revealed several key findings. First, the proposed transformer architecture achieved a 3.31-fold speedup in training time and a 4.49-fold speedup during inference compared to standard U-Net architectures, allowing the model to scale efficiently into billions of parameters. Second, the model set state-of-the-art visual quality benchmarks on standard zero-shot video datasets, UCF-101 and MSR-VTT. Third, in human preference studies on dynamic scenes, participants favored the proposed model over leading public and proprietary systems for prompt alignment in 80% to 81% of cases, for motion quantity in 88% to 96% of cases, and for motion quality in 70% to 79% of cases, while matching top commercial models in photorealism.

These findings demonstrate that treating video as an inherently distinct, compressible modality yields substantial operational and quality advantages. By drastically cutting inference latency and training overhead while scaling parameter capacity, the architecture lowers compute infrastructure costs and shortens generation turnaround times. It also solves the persistent industry challenge of generating believable, large-scale motion without flickering or temporal degradation.

For technical and product leaders evaluating text-to-video capabilities, the primary recommendation is to pivot away from separable U-Net pipelines toward compressed transformer-based architectures and video-calibrated noise schedules for high-resolution synthesis. Organizations looking to adopt the approach should consider implementing cascaded architectures that divide motion synthesis from high-resolution detail generation.

A primary limitation of the study is its reliance on a proprietary training dataset of 238,000 video hours, which includes synthetic captions generated by an internal model. While the quantitative benchmark gains and user evaluations provide strong confidence in the architecture's scalability and motion fidelity, external deployment across open-domain prompts should account for potential data-dependent edge cases and the computational demands of multi-stage inference.

arXiv: 2402.14797
Cover for Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis

Abstract

Contemporary models for generating images show remarkable quality and versatility. Swayed by these advantages, the research community repurposes them to generate videos. Since video content is highly redundant, we argue that naively bringing advances of image models to the video generation domain reduces motion fidelity, visual quality and impairs scalability. In this work, we build Snap Video, a video-first model that systematically addresses these challenges. To do that, we first extend the EDM framework to take into account spatially and temporally redundant pixels and naturally support video generation. Second, we show that a U-Net—a workhorse behind image generation—scales poorly when generating videos, requiring significant computational overhead. Hence, we propose a new transformer-based architecture that trains 3.31 times faster than U-Nets (and is ∼4.5 faster at inference). This allows us to efficiently train a text-to-video model with billions of parameters for the first time, reach state-of-the-art results on a number of benchmarks, and generate videos with substantially higher quality, temporal consistency, and motion complexity. The user studies showed that our model was favored by a large margin over the most recent methods.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Introduction to EDM
  • 3.2. EDM for High-Resolution Video Generation
  • 3.3. Image-Video Modality Matching
  • 3.4. Scalable Video Generator
  • 3.5. Training
  • 3.6. Inference
  • 4. Evaluation
  • 4.1. Datasets
  • 4.2. Evaluation Protocol
  • 4.3. Ablations
  • 4.4. Quantitative Evaluation
  • 4.5. Qualitative Evaluation
  • 5. Conclusions
  • References

Knowls

  1. Knowl 1 — Snap Video’s video-first cascaded generator

    model/method

    Snap Video is a text-to-video diffusion system designed to model videos jointly in space and time rather than adapting an image generator frame by frame. It uses a two-stage cascade: a first-stage model generates 36×6436\times64 pixel videos focused on scene structure and motion, and a second-stage model upsamples them to 288×512288\times512 pixels while specializing in high-frequency details. The cascade avoids temporal flicker that can arise from latent autoencoders and allows the two models to allocate capacity to different aspects of video synthesis. Text conditioning is supplied by a T5-11B encoder, while noise level, frame rate, and original resolution are also provided as conditioning information.

  2. Knowl 2 — SNR-corrected EDM for high-resolution videos

    model/method

    The paper modifies the variance-exploding EDM diffusion process to account for redundancy across video frames and spatial upsampling. For a clean video xx, noise standard deviation σ\sigma, standard Gaussian noise ϵ\epsilon, temporal length TT, and spatial upsampling factor ss, the forward process is defined as

    xσ=xσin+σϵ,σin=sT.x_{\sigma}=\frac{x}{\sigma_{\mathrm{in}}}+\sigma\epsilon,\qquad \sigma_{\mathrm{in}}=s\sqrt{T}.

    Averaging corresponding pixels over a T×s×sT\times s\times s block reduces the noise variance by Ts2Ts^2 and therefore increases the effective signal-to-noise ratio by Ts2Ts^2 when no input scaling is used. Scaling the signal by σin=sT\sigma_{\mathrm{in}}=s\sqrt{T} restores the signal-to-noise ratio expected by the original EDM formulation. For the paper’s 16-frame videos, this gives σin=4\sigma_{\mathrm{in}}=4 for the first-stage model and σin=32\sigma_{\mathrm{in}}=32 for the second-stage model. The reverse sampler compensates for the scaled signal so that generated samples have their original magnitude.

  3. Knowl 3 — Reformulated EDM preconditioning and loss

    equation

    Let FθF_{\theta} be the neural network, DθD_{\theta} its denoiser, xσx_{\sigma} a noisy video, σ\sigma the noise standard deviation, cinc_{\mathrm{in}}, coutc_{\mathrm{out}}, and cskipc_{\mathrm{skip}} scalar preconditioning functions, and FtgtF_{\mathrm{tgt}} and cnrmc_{\mathrm{nrm}} the training target and target normalization factor. Snap Video retains the EDM denoiser parameterization

    Dθ(xσ)=cout(σ)Fθ ⁣(cin(σ)xσ)+cskip(σ)xσ,D_{\theta}(x_{\sigma})=c_{\mathrm{out}}(\sigma)F_{\theta}\!\left(c_{\mathrm{in}}(\sigma)x_{\sigma}\right)+c_{\mathrm{skip}}(\sigma)x_{\sigma},

    and trains the network with

    L(Fθ)=Eσ,x,ϵ[w(σ)∥Fθ ⁣(cin(σ)xσ)−cnrm(σ)Ftgt∥22].\mathcal{L}(F_{\theta})=\mathbb{E}_{\sigma,x,\epsilon}\left[w(\sigma)\left\|F_{\theta}\!\left(c_{\mathrm{in}}(\sigma)x_{\sigma}\right)-c_{\mathrm{nrm}}(\sigma)F_{\mathrm{tgt}}\right\|_2^2\right].

    The redesigned preconditioning functions normalize the scaled noisy input and target, preserve the original EDM loss weighting, keep the training target from acquiring an unstable small-noise term, and reduce exactly to EDM when σin=1\sigma_{\mathrm{in}}=1. Setting the EDM data-scale parameter to σdata=1\sigma_{\mathrm{data}}=1 makes the resulting target and loss weighting coincide with the commonly used vv-prediction formulation.

  4. Knowl 4 — Joint spatiotemporal FIT architecture

    model/method

    Snap Video replaces the frame-wise U-Net computation used by many video diffusion systems with a scaled Far-reaching Interleaved Transformer (FIT). A noisy video is patchified using patches spanning the spatial dimensions only; temporal patching is avoided because temporal patches larger than one frame were found to weaken motion modeling. Patch tokens are grouped so that each group spans all video frames, enabling every group-level operation to access the complete temporal extent. In each FIT block, conditioning tokens are read first, patch information is read into a larger set of learnable latent tokens through groupwise cross-attention, self-attention is performed in the latent space, and the result is written back to patch tokens through groupwise cross-attention. Feed-forward modules replace expensive local patch-token self-attention layers. The updated patch tokens are projected back to pixels, and latent-token self-conditioning preserves the compressed representation across diffusion steps. Joint computation in this learned compressed space allows the model to scale to tens of thousands of input patches and billions of parameters while modeling spatial and temporal dependencies together.

  5. Knowl 5 — Unified image-video training through infinite frame rate

    model/method

    To train jointly on the much larger image corpus and the smaller video corpus without using separate diffusion processes, Snap Video represents an image as a video with infinitely high frame rate. The training procedure varies the effective frame rate so that image and video examples occupy a continuous modality range. This avoids assigning different input-scaling rules to images and videos and makes image examples contribute to learning temporal denoising behavior rather than merely teaching an image-only model. In the paper’s ablations, treating images as infinite-frame-rate videos consistently improves the FID of the generated videos.

  6. Knowl 6 — Training and sampling configuration

    experimental setup

    Snap Video is optimized with LAMB using learning rate 5×10−35\times10^{-3}, cosine learning-rate decay, and a total batch size of 2048 videos plus 2048 images. The first-stage model is trained for 550,000 steps. The second-stage upsampler is initialized from the first-stage weights and fine-tuned on high-resolution videos for 370,000 iterations. At inference, samples begin from Gaussian noise and are generated with the deterministic EDM sampler, using 256 sampling steps for the first stage and 40 for the second stage. Classifier-free guidance is used for text alignment, and dynamic thresholding together with oscillating guidance was found to improve quality.

  7. Knowl 7 — FIT scaling improves quality and efficiency

    data/table

    The following comparison uses the first-stage setting at 64×3664\times36 pixels on the authors’ internal dataset. FID and FVD are lower-is-better; CLIPSIM is higher-is-better. The reported training and inference timing values are in milliseconds per video per GPU, so lower values indicate faster execution. The 500M-parameter FIT is both faster and better on FID and CLIPSIM than the 284M-parameter U-Net, while the 3.9B-parameter FIT further improves all quality metrics.

    Could not parse LaTeX table

    Relative to the 284M U-Net, the 500M FIT trains 3.31×3.31\times faster and performs inference 4.49×4.49\times faster. Increasing FIT capacity from 500M to 3.9B parameters improves FID from 3.07 to 2.51, FVD from 27.79 to 12.31, and CLIPSIM from 0.2459 to 0.2579; the 3.9B model requires only a 1.24×1.24\times longer inference time than the 284M U-Net.

  8. Knowl 8 — Ablations validate the video diffusion design

    empirical result

    Using the 500M FIT at 64×3664\times36 pixels, the paper evaluates the original EDM configuration, the scaled diffusion process, different σdata\sigma_{\mathrm{data}} values, different input scalings, and whether images are treated as infinite-frame-rate videos. The reported results are:

    Could not parse LaTeX table

    The authors report that the proposed framework improves over original EDM on all three metrics, that σdata=1\sigma_{\mathrm{data}}=1 is beneficial, that using an input scaling below the value required by the video dimensions harms the intended diffusion behavior, and that treating images as infinite-frame-rate videos improves FID. The ablation supports using the SNR-matched scaling together with the unified image-video training rule rather than applying an image diffusion configuration directly to videos.

  9. Knowl 9 — Training corpus and zero-shot evaluation protocol

    experimental setup

    Snap Video is trained on an internal captioned corpus containing 1.265 million images and 238,000 hours of video. Because many videos lack high-quality captions, a video-captioning model supplies synthetic captions for part of the video data. Evaluation uses datasets not observed during training: UCF-101 contains 13,320 videos from 101 action classes, and MSR-VTT contains 10,000 web videos with 20 captions each; its test split contains 2,990 videos and 59,800 captions. For all zero-shot comparisons, the model generates 16-frame videos at 24 frames per second and 512×288512\times288 resolution, with additional evaluation at the commonly used 288×288288\times288 square resolution. UCF-101 evaluation generates 10,000 videos with the original class distribution and reports FVD and Inception Score. MSR-VTT evaluation generates one sample for each of the 59,800 test captions and reports CLIP-FID and CLIPSIM.

  10. Knowl 10 — Zero-shot benchmark performance

    data/table

    Snap Video is compared with published text-to-video systems on UCF-101 and MSR-VTT. Lower FVD and FID are better, while higher Inception Score and CLIPSIM are better. On UCF-101, the native 512×288512\times288 Snap Video output obtains the best reported FVD and FID among the listed methods, while both Snap Video resolutions obtain an Inception Score of 38.89. On MSR-VTT, Snap Video obtains the lowest reported CLIP-FID and FVD, although its CLIPSIM is below several methods using CLIP-based text representations.

    Could not parse LaTeX table
    Could not parse LaTeX table
  11. Knowl 11 — User preference for motion and text alignment

    empirical result

    A paired user study compares Snap Video with Gen-2, PikaLab, and Floor33 on 65 prompts describing dynamic scenes. Each method uses its default settings, and five users vote on each paired comparison. The values below are the percentages of votes favoring Snap Video.

    Could not parse LaTeX table

    Snap Video is preferred by a large margin for text-video alignment and for both motion measures. Its photorealism is reported as comparable to Gen-2 and better than PikaLab and Floor33. Qualitative comparisons additionally show that competing systems sometimes produce weakly moving dynamic images, flickering, or motion artifacts, whereas the joint spatiotemporal model more often produces vivid, temporally coherent motion.

Coverage note — No substantial contributed material was omitted; appendix-only hyperparameters and additional qualitative samples were not separately decomposed.

References

  1. 1.Pika lab discord server. https://www.pika.art/. Accessed: 2023-11-01. 2, 7, 8, 3, 6
  2. 2.Jie An, Songyang Zhang, Harry Yang, Sonal Gupta, Jia-Bin Huang, Jiebo Luo, and Xi Yin. Latent-shift: Latent diffusion with temporal shift for efficient text-to-video generation. arXiv, 2023. 2, 3, 7
  3. 3.Yogesh Balaji, Seungjun Nah, Xun Huang, Arash Vahdat, Jiaming Song, Qinsheng Zhang, Karsten Kreis, Miika Aittala, Timo Aila, Samuli Laine, Bryan Catanzaro, Tero Karras, and Ming-Yu Liu. ediff-i: Text-to-image diffusion models with an ensemble of expert denoisers. ArXiv, 2022. 2, 3
  4. 4.Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis. Align your latents: High-resolution video synthesis with latent diffusion models. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 1, 2, 3, 4, 5, 6, 7, 8, 15
  5. 5.Tim Brooks, Janne Hellsten, Miika Aittala, Ting-Chun Wang, Timo Aila, Jaakko Lehtinen, Ming-Yu Liu, Alexei A Efros, and Tero Karras. Generating long videos of dynamic scenes. In Advances in Neural Information Processing Systems (NeurIPS), 2022. 2
  6. 6.Tim Brooks, Aleksander Holynski, and Alexei A. Efros. Instructpix2pix: Learning to follow image editing instructions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 1
  7. 7.Ting Chen. On the importance of noise scheduling for diffusion models. arXiv, 2023. 2, 3, 4
  8. 8.Ting Chen and Lala Li. Fit: Far-reaching interleaved transformers. arXiv, 2023. 2, 3, 5, 8, 1
  9. 9.Aidan Clark, Jeff Donahue, and Karen Simonyan. Efficient video generation on complex datasets. arXiv, 2019. 2
  10. 10.Prafulla Dhariwal and Alexander Quinn Nichol. Diffusion models beat GANs on image synthesis. In Advances in Neural Information Processing Systems (NeurIPS), 2021. 6, 7
  11. 11.Patrick Esser, Johnathan Chiu, Parmida Atighehchian, Jonathan Granskog, and Anastasis Germanidis. Structure and content-guided video synthesis with diffusion models. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 2, 7, 8, 3, 6
  12. 12.Songwei Ge, Thomas Hayes, Harry Yang, Xi Yin, Guan Pang, David Jacobs, Jia-Bin Huang, and Devi Parikh. Long video generation with time-agnostic vqgan and timesensitive transformer. In Proceedings of the European Conference of Computer Vision (ECCV), 2022. 2
  13. 13.Songwei Ge, Seungjun Nah, Guilin Liu, Tyler Poon, Andrew Tao, Bryan Catanzaro, David Jacobs, Jia-Bin Huang, Ming-Yu Liu, and Yogesh Balaji. Preserve your own correlation: A noise prior for video diffusion models. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2023. 1, 2, 3, 4, 5, 6, 7, 8, 15
  14. 14.Jiatao Gu, Shuangfei Zhai, Yizhe Zhang, Miguel Angel Bautista, and Josh Susskind. f-dm: A multi-stage diffusion model via progressive signal transformation. International Conference on Learning Representations (ICLR), 2023. 3
  15. 15.Jiatao Gu, Shuangfei Zhai, Yizhe Zhang, Josh Susskind, and Navdeep Jaitly. Matryoshka diffusion models. arXiv, 2023. 3
  16. 16.Yuwei Guo, Ceyuan Yang, Anyi Rao, Yaohui Wang, Yu Qiao, Dahua Lin, and Bo Dai. Animatediff: Animate your personalized text-to-image diffusion models without specific tuning. arXiv, 2023. 2
  17. 17.Yingqing He, Tianyu Yang, Yong Zhang, Ying Shan, and Qifeng Chen. Latent video diffusion models for high-fidelity long video generation. arXiv, 2023. 2, 3, 4, 7, 8, 6
  18. 18.Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. In Advances in Neural Information Processing Systems (NeurIPS), 2017. 6
  19. 19.Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance. arXiv, 2022. 6, 2
  20. 20.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. In Advances in Neural Information Processing Systems (NeurIPS), 2020. 3
  21. 21.Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P. Kingma, Ben Poole, Mohammad Norouzi, David J. Fleet, and Tim Salimans. Imagen video: High definition video generation with diffusion models. arXiv, 2022. 1, 2, 3, 4, 5, 6, 8, 15
  22. 22.Jonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan, Mohammad Norouzi, and David J. Fleet. Video diffusion models. In ICLR Workshop on Deep Generative Models for Highly Structured Data, 2022. 2, 4, 3
  23. 23.Wenyi Hong, Ming Ding, Wendi Zheng, Xinghan Liu, and Jie Tang. Cogvideo: Large-scale pretraining for text-to-video generation via transformers. arXiv, 2022. 2, 7
  24. 24.Emiel Hoogeboom, Jonathan Heek, and Tim Salimans. Simple diffusion: End-to-end diffusion for high resolution images. In Proceedings of the 40th International Conference on Machine Learning (ICML), 2023. 3
  25. 25.Tero Karras, Miika Aittala, Timo Aila, and Samuli Laine. Elucidating the design space of diffusion-based generative models. In Advances in Neural Information Processing Systems (NeurIPS), 2022. 2, 3, 4, 6, 8
  26. 26.Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv, 2015. 2
  27. 27.Tuomas Kynkäanniemi, Tero Karras, Miika Aittala, Timo ¨ Aila, and Jaakko Lehtinen. The role of imagenet classes in frechet inception distance. In ´ International Conference on Learning Representations (ICLR), 2023. 7
  28. 28.Alex X. Lee, Richard Zhang, Frederik Ebert, P. Abbeel, Chelsea Finn, and S. Levine. Stochastic adversarial video prediction. arXiv, abs/1804.01523, 2018. 2
  29. 29.Yitong Li, Martin Min, Dinghan Shen, David Carlson, and Lawrence Carin. Video generation from text. Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), 2018. 2
  30. 30.Yanyu Li, Huan Wang, Qing Jin, Ju Hu, Pavlo Chemerys, Yun Fu, Yanzhi Wang, Sergey Tulyakov, and Jian Ren. Snapfusion: Text-to-image diffusion model on mobile devices within two seconds. arXiv, 2023. 1
  31. 31.Yaofang Liu, Xiaodong Cun, Xuebo Liu, Xintao Wang, Yong Zhang, Haoxin Chen, Yang Liu, Tieyong Zeng, Raymond Chan, and Ying Shan. Evalcrafter: Benchmarking and evaluating large video generation models. arXiv, 2023. 8, 3, 7
  32. 32.Z. Luo, D. Chen, Y. Zhang, Y. Huang, L. Wang, Y. Shen, D. Zhao, J. Zhou, and T. Tan. Videofusion: Decomposed diffusion models for high-quality video generation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 2, 6, 7
  33. 33.Siwei Ma, Xinfeng Zhang, Chuanmin Jia, Zhenghui Zhao, Shiqi Wang, and Shanshe Wang. Image and video compression with neural networks: A review. IEEE Transactions on Circuits and Systems for Video Technology, 2019. 2, 5
  34. 34.Gaurav Mittal, Tanya Marwah, and Vineeth N. Balasubramanian. Sync-draw: Automatic video generation using deep recurrent attentive architectures. In Proceedings of the 25th ACM International Conference on Multimedia, 2017. 2
  35. 35.Guillaume Le Moing, Jean Ponce, and Cordelia Schmid. CCVS: Context-aware controllable video synthesis. In Advances in Neural Information Processing Systems (NeurIPS), 2021. 2
  36. 36.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zach DeVito, Martin Raison, ¨ Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. PyTorch: An Imperative Style, High-Performance Deep Learning Library. 2019. 1
  37. 37.Chenyang QI, Xiaodong Cun, Yong Zhang, Chenyang Lei, Xintao Wang, Ying Shan, and Qifeng Chen. Fatezero: Fusing attentions for zero-shot text-based video editing. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2023. 1
  38. 38.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning (ICML), 2021. 7
  39. 39.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research (JMLR), 2022. 6, 7, 15
  40. 40.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bjorn Ommer. High-resolution image syn- ¨ thesis with latent diffusion models. arXiv, 2021. 1, 2, 3
  41. 41.Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-net: Convolutional networks for biomedical image segmentation. In Medical Image Computing and Computer-Assisted Intervention (MICCAI), 2015. 2, 4, 8, 1
  42. 42.Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 1
  43. 43.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L. Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, Seyedeh Sara Mahdavi, Raphael Gontijo Lopes, Tim Salimans, Jonathan Ho, David J. Fleet, and Mohammad Norouzi. Photorealistic textto-image diffusion models with deep language understanding. arXiv, 2022. 1, 2, 3, 6, 7
  44. 44.Masaki Saito, Shunta Saito, Masanori Koyama, and Sosuke Kobayashi. Train sparsely, generate densely: Memory-efficient unsupervised training of high-resolution temporal gan, 2020. 2
  45. 45.Tim Salimans and Jonathan Ho. Progressive distillation for fast sampling of diffusion models. In International Conference on Learning Representations (ICLR), 2022. 4, 7, 12
  46. 46.Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen. Improved techniques for training gans. In Advances in Neural Information Processing Systems (NeurIPS), 2016. 7
  47. 47.Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. Mostgan-v: Video generation with temporal motion styles. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 2
  48. 48.Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, Devi Parikh, Sonal Gupta, and Yaniv Taigman. Make-a-video: Text-to-video generation without text-video data. arXiv, 2022. 1, 2, 3, 4, 5, 6, 7, 8
  49. 49.Ivan Skorokhodov, Sergey Tulyakov, and Mohamed Elhoseiny. Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 2
  50. 50.Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. In International Conference on Machine Learning (ICML), 2015. 3
  51. 51.Yang Song and Stefano Ermon. Generative modeling by estimating gradients of the data distribution. In Advances in Neural Information Processing Systems (NeurIPS), 2019. 3
  52. 52.Yang Song and Stefano Ermon. Improved techniques for training score-based generative models. In Advances in Neural Information Processing Systems (NeurIPS), 2020. 3
  53. 53.Yang Song, Conor Durkan, Iain Murray, and Stefano Ermon. Maximum likelihood training of score-based diffusion models, 2021. 3
  54. 54.Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equations. In International Conference on Learning Representations (ICLR), 2021. 3
  55. 55.Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah. UCF101: A dataset of 101 human actions classes from videos in the wild. arXiv, 2012. 2, 6, 7, 8, 3, 11
  56. 56.Nitish Srivastava, Elman Mansimov, and Ruslan Salakhudinov. Unsupervised learning of video representations using lstms. In International Conference on Machine Learning (ICML), 2015. 2
  57. 57.Jiayan Teng, Wendi Zheng, Ming Ding, Wenyi Hong, Jianqiao Wangni, Zhuoyi Yang, and Jie Tang. Relay diffusion: Unifying diffusion process across resolutions for image synthesis. arXiv, 2023. 3
  58. 58.Yu Tian, Jian Ren, Menglei Chai, Kyle Olszewski, Xi Peng, Dimitris N Metaxas, and Sergey Tulyakov. A good image generator is what you need for high-resolution video synthesis. In International Conference on Learning Representations (ICLR), 2021. 2
  59. 59.Sergey Tulyakov, Ming-Yu Liu, Xiaodong Yang, and Jan Kautz. Mocogan: Decomposing motion and content for video generation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018. 2
  60. 60.Thomas Unterthiner, Sjoerd van Steenkiste, Karol Kurach, Raphael Marinier, Marcin Michalski, and Sylvain Gelly. Towards accurate generative models of video: A new metric & challenges. arXiv, 2018. 6, 7
  61. 61.Ruben Villegas, Mohammad Babaeizadeh, Pieter-Jan Kindermans, Hernan Moraldo, Han Zhang, Mohammad Taghi Saffar, Santiago Castro, Julius Kunze, and D. Erhan. Phenaki: Variable length video generation from open domain textual description. In International Conference on Learning Representations (ICLR), 2023. 2
  62. 62.Wenjing Wang, Huan Yang, Zixi Tuo, Huiguo He, Junchen Zhu, Jianlong Fu, and Jiaying Liu. Videofactory: Swap attention in spatiotemporal diffusions for text-to-video generation. arXiv, 2023. 1, 2, 5, 6, 7, 15
  63. 63.Chenfei Wu, Lun Huang, Qianxi Zhang, Binyang Li, Lei Ji, Fan Yang, Guillermo Sapiro, and Nan Duan. Godiva: Generating open-domain videos from natural descriptions. ArXiv, 2021. 2, 6, 7
  64. 64.Chenfei Wu, Jian Liang, Lei Ji, Fan Yang, Yuejian Fang, Daxin Jiang, and Nan Duan. Nuwa: Visual synthesis pre- ¨ training for neural visual world creation. In Proceedings of the European Conference of Computer Vision (ECCV), 2022. 2, 7
  65. 65.Jun Xu, Tao Mei, Ting Yao, and Yong Rui. Msr-vtt: A large video description dataset for bridging video and language. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016. 2, 6, 7, 8
  66. 66.Wilson Yan, Yunzhi Zhang, Pieter Abbeel, and Aravind Srinivas. Videogpt: Video generation using vq-vae and transformers. arXiv, 2021. 2
  67. 67.Sheng-Siang Yin, Chenfei Wu, Huan Yang, Jianfeng Wang, Xiaodong Wang, Minheng Ni, Zhengyuan Yang, Linjie Li, Shuguang Liu, Fan Yang, Jianlong Fu, Gong Ming, Lijuan Wang, Zicheng Liu, Houqiang Li, and Nan Duan. Nuwa-xl: Diffusion over diffusion for extremely long video generation. In Annual Meeting of the Association for Computational Linguistics, 2023. 2
  68. 68.Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh. Large batch optimization for deep learning: Training bert in 76 minutes. In International Conference on Learning Representations (ICLR), 2020. 6, 1
  69. 69.Jiahui Yu, Yuanzhong Xu, Jing Yu Koh, Thang Luong, Gunjan Baid, Zirui Wang, Vijay Vasudevan, Alexander Ku, Yinfei Yang, Burcu Karagol Ayan, Ben Hutchinson, Wei Han, Zarana Parekh, Xin Li, Han Zhang, Jason Baldridge, and Yonghui Wu. Scaling autoregressive models for content-rich text-to-image generation. Transactions on Machine Learning Research, 2022. 1
  70. 70.Lijun Yu, Yong Cheng, Kihyuk Sohn, Jose Lezama, Han ´ Zhang, Huiwen Chang, Alexander G. Hauptmann, MingHsuan Yang, Yuan Hao, Irfan Essa, and Lu Jiang. Magvit: Masked generative video transformer. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 2
  71. 71.Sihyun Yu, Jihoon Tack, Sangwoo Mo, Hyunsu Kim, Junho Kim, Jung-Woo Ha, and Jinwoo Shin. Generating videos with dynamics-aware implicit generative adversarial networks. In International Conference on Learning Representations (ICLR), 2022. 2
  72. 72.Daquan Zhou, Weimin Wang, Hanshu Yan, Weiwei Lv, Yizhe Zhu, and Jiashi Feng. Magicvideo: Efficient video generation with latent diffusion models. arXiv, 2023. 2, 3, 6, 7

Citation

MLA
Menapace, W., et al. “Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis”. arXiv, 2024, http://arxiv.org/abs/2402.14797v1.
APA
Menapace, W., Siarohin, A., Skorokhodov, I., Deyneka, E., Chen, T.-S., Kag, A., Fang, Y., Stoliar, A., Ricci, E., Ren, J., & Tulyakov, S. (2024). Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis. arXiv. http://arxiv.org/abs/2402.14797v1
Chicago
Menapace, W., A. Siarohin, I. Skorokhodov, et al. 2024. “Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis”. arXiv. http://arxiv.org/abs/2402.14797v1.
Harvard
Menapace, W. et al. (2024) “Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.14797v1.
Vancouver
1. Menapace W, Siarohin A, Skorokhodov I, et al (2024) Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis. arXiv

BibTeX

@article{menapace2024snap,
  title = {Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis},
  author = {Menapace, Willi and Siarohin, Aliaksandr and Skorokhodov, Ivan and Deyneka, Ekaterina and Chen, Tsai-Shien and Kag, Anil and Fang, Yuwei and Stoliar, Aleksei and Ricci, Elisa and Ren, Jian and Tulyakov, Sergey},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.14797v1},
  eprint = {2402.14797}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE