VideoPoet: A Large Language Model for Zero-Shot Video Generation

Dan KondratyukLijun YuXiuye GuJosé LezamaJonathan HuangGrant SchindlerRachel HornungVighnesh BirodkarJimmy YanMing-Chang Chiu

article2024ICML566 citationsBest Paper Award

Demonstrates that a unified decoder-only language model trained on discrete multimodal tokens can match or outperform diffusion approaches across diverse video generation and editing tasks without task-specific architectural changes.

Listen

Recent advances in generative artificial intelligence have enabled automated video creation, but the field relies almost entirely on diffusion models that require complex, separate modules or modifications to handle different tasks. In contrast, large language model architectures have shown immense versatility across language, speech, and robotics, yet their application to high-quality video generation has remained largely understudied. The article introduces VideoPoet, a unified decoder-only language model framework designed to evaluate and demonstrate that language models can perform diverse, high-fidelity video generation tasks within a single architecture.

To achieve this, the approach tokenizes multimodal inputs—including text embeddings, discrete visual tokens, and audio tokens—into a shared vocabulary of approximately 300,000 codes. The system leverages a two-stage pretraining and task-adaptation protocol, training an 8-billion-parameter model on roughly two trillion tokens across one billion image-text pairs and 270 million videos. This process incorporates alternating gradient descent across multiple tasks, such as text-to-video, image-to-video, video future prediction, inpainting, outpainting, and stylization. A custom spatial super-resolution transformer then upsamples base outputs to higher visual resolutions.

The article demonstrates several significant findings. First, the unified model achieves state-of-the-art results across standard zero-shot video generation benchmarks, attaining strong performance on datasets such as MSR-VTT and UCF-101. Second, side-by-side human evaluations indicate that the system is highly competitive with leading video diffusion models, earning clear preference for motion realism and interestingness. Third, scaling model capacity to 8 billion parameters substantially enhances temporal consistency, prompt fidelity, spatial reasoning, and counting. Finally, the framework demonstrates flexible task chaining, smoothly combining capabilities like animating a static image into a video and subsequent video stylization without generative degradation.

These findings prove that large language model architectures represent a viable, high-performance alternative to diffusion methods for video generation. Adopting language models allows organizations to leverage mature training infrastructure, hardware optimizations, and unified multitask scaling without maintaining disparate specialized models. The primary trade-off involves computational cost during scaling and runtime. Next steps include exploring additional acceleration techniques for inference, expanding model capabilities to direct text generation, and implementing governance strategies such as digital watermarking to mitigate ethical risks and deceptive misuse. Limitations include visual fidelity upper bounds set by discrete tokenizers, challenges rendering fine-grained details during rapid motion, and baseline aesthetic differences from excluding certain copyrighted datasets.

  • Paper: Phenaki: Variable Length Video Generation From Open Domain Textual Description, Ruben Villegas et al. (2023). Phenaki introduced autoregressive token-based video generation from text sequences using spatiotemporal visual tokenizers, directly establishing the foundation for VideoPoet's decoder-only visual token modeling.
  • Paper: VideoGPT: Video Generation using VQ-VAE and Transformers, Wilson Yan et al. (2021). VideoGPT established the core paradigm of compressing video into discrete VQ-VAE tokens and generating them sequentially using transformer language models.
  • Paper: Scaling Autoregressive Models for Content-Rich Text-to-Image Generation, Jiahui Yu et al. (2022). Parti demonstrated that scaling autoregressive sequence-to-sequence language models over discrete visual tokens achieves competitive generative quality compared to diffusion models, inspiring VideoPoet's LLM-driven video synthesis.
  • Paper: Zero-Shot Text-to-Image Generation, Aditya Ramesh et al. (2021). This work pioneered zero-shot visual generation using large autoregressive transformers over discrete image tokens, establishing the core framework extended to video in VideoPoet.
  • Paper: Video Diffusion Models, Jonathan Ho et al. (2022). Provides the foundational video generation benchmarks, conditioning setups, and spatiotemporal modeling concepts that VideoPoet positions its LLM-based architecture against.
  • Paper: Imagen Video: High Definition Video Generation with Diffusion Models, Jonathan Ho et al. (2022). Establishes standard text-to-video cascaded architectures and spatial-temporal super-resolution pipelines that VideoPoet adapts into its custom super-resolution transformer.
  • Paper: Make-A-Video: Text-to-Video Generation without Text-Video Data, Uriel Singer et al. (2023). Make-A-Video introduces critical zero-shot text-to-video evaluation methodologies and task adaptation strategies that VideoPoet directly benchmarks against.
  • Paper: ViViT: A Video Vision Transformer, Anurag Arnab et al. (2021). Introduces spatiotemporal tubelet tokenization and factorized transformer attention mechanisms essential for processing video as discrete sequence tokens in transformer architectures.
Cover for VideoPoet: A Large Language Model for Zero-Shot Video Generation

Abstract

We present VideoPoet, a model for synthesizing high-quality videos from a large variety of conditioning signals. VideoPoet employs a decoder-only transformer architecture that processes multimodal inputs – including images, videos, text, and audio. The training protocol follows that of Large Language Models (LLMs), consisting of two stages: pretraining and task-specific adaptation. During pretraining, VideoPoet incorporates a mixture of multimodal generative objectives within an autoregressive Transformer framework. The pretrained LLM serves as a foundation that is adapted to a range of video generation tasks. We present results demonstrating the model’s state-of-the-art capabilities in zero-shot video generation, specifically highlighting the generation of high-fidelity motions. Project page: https://sites.research.google/videopoet/.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Model Overview
  • 3.1. Tokenization
  • 3.2. Language Model Backbone
  • 3.3. Super-Resolution
  • 4. LLM Pretraining for Generation
  • 4.1. Task Prompt Design
  • 4.2. Training Strategy
  • 5. Experiments
  • 5.1. Experimental Setup
  • 5.2. Pretraining Task Analysis
  • 5.3. Comparison with the State-of-the-Art
  • 5.4. Runtime
  • 5.5. LLM's Diverse Capabilities in Video Generation
  • 5.6. Limitations
  • 6. Conclusion
  • Acknowledgements
  • Impact Statement
  • References
  • A. Appendix
  • A.1. Responsible AI and Fairness Analysis
  • A.2. Model Scale and Performance
  • A.2.1. QUALITATIVE COMPARISON OF 1B AND 8B MODELS
  • A.3. Additional Generated Examples
  • A.4. Video Stylization
  • A.5. Additional Implementation and Evaluation Details
  • A.5.1. ADDITIONAL IMPLEMENTATION DETAILS
  • A.5.2. SUPER-RESOLUTION IMPLEMENTATION DETAILS
  • A.5.3. ADDITIONAL EVALUATION DETAILS
  • A.5.4. ADDITIONAL HUMAN EVALUATION DETAILS
  • A.5.5. ZERO-SHOT TEXT-TO-VIDEO EVALUATION SETTINGS
  • A.5.6. SELF-SUPERVISED TASKS EVALUATION SETTINGS
  • A.5.7. STYLIZATION EVALUATION ON DAVIS

Knowls

  1. Knowl 1 — VideoPoet Architecture and Multimodal Sequence Layout

    model/method

    VideoPoet is a generative foundation model for video, audio, and image synthesis built on a decoder-only prefix Large Language Model (LLM) backbone. It models multiple modalities within a single shared discrete token space.

    The system utilizes three modality representations:

    1. Visual Modality: Video clips and images are quantized into discrete visual tokens using the MAGVIT-v2 tokenizer. A 17-frame, 2.125-second video at 128×128128 \times 128 resolution (sampled at 8 fps) produces a latent grid of shape (5,16,16)(5, 16, 16), which is flattened into 1,280 tokens. For 128×224128 \times 224 portrait resolution, the latent grid is (5,28,16)(5, 28, 16), totaling 2,240 tokens. Static images are encoded as 1-frame videos of shape (1,16,16)(1, 16, 16) (256 tokens).
    2. Audio Modality: Audio clips of 2.125 seconds are encoded via a pretrained SoundStream tokenizer with a 4-level Residual Vector Quantizer (RVQ), yielding 106 temporal latent frames. Each RVQ level uses a disjoint vocabulary of 1,024 codes (totaling 4,096 audio codes). Tokens are predicted sequentially from lower to higher RVQ levels across the clip.
    3. Text Conditioning: Pretrained continuous text embeddings are extracted using a frozen T5 XL encoder and mapped into the LLM embedding space via a learned linear projection layer.

    The total unified vocabulary consists of approximately 300,000 tokens: 256 special control codes (such as \<bos\>, \<task\>, \<res\>, \<bov_i\>, \<eov_i\>, \<bov_o\>, \<eov_o\>, and \<eos\>), 262,144 visual codes (2182^{18}), 4,096 audio codes, and a small English text token vocabulary.

    The input sequence format consists of two parts:

    • Prefix Input: Bidirectional self-attention over conditioning tokens (task tokens, text embeddings, input visual tokens, input audio tokens).
    • Output Generation: Causal autoregressive attention predicting target visual and/or audio tokens, with training loss applied exclusively to output token positions.
  2. Knowl 2 — Multimodal Pretraining Task Mixture and Prompt Design

    model/method

    VideoPoet is pretrained on a mixture of multimodal generative tasks within a single transformer, where each task is defined by a prefix input and an autoregressively predicted output:

    • Text-to-Video (T2V): Prefix contains text embeddings; output is video tokens.
    • Text-to-Image (T2I): Prefix contains text embeddings; output is a single-frame video token sequence without \<eos\> and \<eov_o\> tokens, enabling smooth autoregressive continuation into video.
    • Image-to-Video (I2V): Prefix contains the initial frame tokens (shape (1,16,16)(1, 16, 16)); output is the remaining video tokens.
    • Video Future Prediction (FP): Prefix contains variable-length initial video tokens; output is future video tokens.
    • Video Inpainting/Outpainting: Masked video tokens encoded using the COMMIT framework serve as input prefix; output is the reconstructed unmasked video.
    • Video Stylization: Prefix contains optical flow tokens, monocular depth tokens, and text prompt embeddings; output is stylized video tokens.
    • Audio-to-Video and Video-to-Audio: Conditioned on one modality's discrete tokens to predict the other.
    • Audio-Video Continuation (AVCont): Prefix contains the initial visual frame and its corresponding audio; output is the remaining synchronized visual and audio tokens.
    • Unconditioned Video Generation: No input prefix; output is visual video tokens.

    Task identification is mediated by a special \<task\> token indicating the unique output format. Tasks sharing identical output types (e.g., T2V, I2V, and unconditioned generation) share the same \<task\> token, allowing the LLM to dynamically adapt to varying prefix signals.

  3. Knowl 3 — Custom Non-Autoregressive Video Super-Resolution Transformer

    model/method

    To scale generated video outputs from base resolution (128×224128 \times 224 or 128×128128 \times 128) up to high resolution (896×512896 \times 512) without prohibitive autoregressive sequence lengths, VideoPoet employs a two-stage spatial super-resolution (SR) cascade in discrete token space:

    • Stage 1 (1B parameters): Upsamples 17×224×12817 \times 224 \times 128 videos to 17×448×25617 \times 448 \times 256 pixels (target token sequence shape (5,56,32)(5, 56, 32)).
    • Stage 2 (500M parameters): Upsamples 17×448×25617 \times 448 \times 256 to 17×896×51217 \times 896 \times 512 pixels (target token sequence shape (5,112,64)(5, 112, 64)).

    Both stages use a custom non-autoregressive transformer operating on MAGVIT-v2 tokens with the following architectural designs:

    1. Multi-Axis Windowed Attention: To avoid quadratic memory on long token sequences, transformer blocks consist of three sub-layers performing local windowed self-attention along three orthogonal axes: spatial vertical, spatial horizontal, and temporal. In Stage 1, window shapes are (1,56,4)(1, 56, 4), (1,8,32)(1, 8, 32), and (5,8,8)(5, 8, 8). In Stage 2, window shapes are (1,112,2)(1, 112, 2), (1,4,64)(1, 4, 64), and (5,8,8)(5, 8, 8). Cross-attention layers attend to low-resolution (LR) token sequences using isomorphic windows at half spatial resolution.
    2. Token Factorization: To manage the 262,144-class visual codebook, token prediction is factorized into k=2k = 2 classification heads of 512 classes each (512×512=262,144512 \times 512 = 262,144).
    3. Training and Inference: Models are trained using masked token modeling on 64M high-quality video pairs. Ground-truth LR tokens are noise-augmented via random discrete resampling, and conditioning signals (LR tokens and T5 XL text embeddings) are dropped independently 10% of the time. Sampling uses 24 non-autoregressive refinement steps with classifier-free guidance scales of 4.0 (text) / 1.0 (LR) for Stage 1, and 8.0 (text) / 2.0 (LR) for Stage 2.
  4. Knowl 4 — Alternating Gradient Descent and Two-Stage Curriculum Pretraining

    model/method

    VideoPoet optimizes training efficiency and multimodal dynamics via two dedicated training strategies:

    1. Alternating Gradient Descent (AGD) for Variable Sequence Lengths: Tasks with different sequence lengths (e.g., single-frame image generation vs. multi-frame long video prediction) are grouped into batches by length. The optimizer alternates gradient updates across these groups, achieving a near 0% token padding ratio without sequence packing artifacts.
    2. Two-Stage Pretraining Curriculum: Uniform temporal sampling of image and video data leads to poor motion modeling. Pretraining is structured in two distinct phases:
      • Phase 1 (First 25% iterations): 90% image data and 10% video data to accelerate visual concept understanding and spatial fidelity.
      • Phase 2 (Remaining 75% iterations): 90% video data and 10% image data to prioritize motion dynamics and temporal coherence.
    3. Task-Specific Alignment Finetuning: Pretrained models are fine-tuned on a filtered subset of millions of high-quality video clips. This fine-tuning stage mitigates decoding collapse (repetitive token prediction cycles) and allows the application of higher classifier-free guidance scales during sampling.
  5. Knowl 5 — Zero-Shot Text-to-Video Benchmark Performance

    data/table

    VideoPoet was evaluated in a zero-shot setting on the MSR-VTT and UCF-101 benchmarks. Videos were evaluated at 16 frames resized to 256×256256 \times 256. CLIP similarity (CLIPSIM) used CLIP ViT-B/16 (MSR-VTT), Fréchet Video Distance (FVD) used an I3D model trained on Kinetics-400 (evaluated across 2,048 samples with 20 repeats on MSR-VTT, and 10,000 samples on UCF-101), and Inception Score (IS) used a C3D model on UCF-101.

    Model MSR-VTT UCF-101
    CLIPSIM ↑\uparrow FVD ↓\downarrow FVD ↓\downarrow IS ↑\uparrow
    CogVideo (EN) (2022) 0.2631 1294 702 25.27
    MagicVideo (2022) - 998 655 -
    Video LDM (2023b) 0.2929 - 551 33.45
    ModelScopeT2V (2023a) 0.2930 550 - -
    InternVid (2023d) 0.2951 - 617 21.04
    VideoFactory (2023c) 0.3005 - 410 -
    Make-A-Video (2022) 0.3049 - 367 33.00
    Show-1 (2023a) 0.3072 538 394 35.42
    VideoPoet (Pretrain, 8B) 0.3049 213 355 38.44
    VideoPoet (Task adapt, 8B) 0.3123 - - -

    VideoPoet (8B) achieves state-of-the-art zero-shot FVD (213 on MSR-VTT, 355 on UCF-101) and Inception Score (38.44 on UCF-101), outperforming both diffusion-based and prior transformer-based video generation models.

  6. Knowl 6 — Ablation of Pretraining Task Mixtures across Video Benchmark Tasks

    data/table

    Ablation study evaluating the effect of pretraining task mixtures on downstream zero-shot performance using a 300M parameter model (trained for 300k steps with learning rate 10−310^{-3} and batch size 1,024) compared against the scaled 8B model. Downstream evaluation benchmarks: Text-to-Video on MSR-VTT (CLIPSIM ↑\uparrow) and UCF-101 (FVD ↓\downarrow), Frame Prediction on Kinetics-600 (K600 FVD ↓\downarrow), and Central Inpainting/Outpainting on Something-Something V2 (SSv2 FVD ↓\downarrow).

    Method Pretraining Tasks Zero-shot Evaluation Benchmark
    T2I T2V Uncond FP Painting AVCont T2V FP Inpaint Outpaint
    MSR-VTT UCF101 K600 SSv2 SSv2
    CLIPSIM ↑\uparrow FVD ↓\downarrow FVD ↓\downarrow FVD ↓\downarrow FVD ↓\downarrow
    T2V ✓ 0.244 822 759 2,333 2,310
    T2V+I ✓ ✓ 0.247 1,025 794 2,118 1,916
    SSL ✓ ✓ ✓ ✓ 0.226 1,742 700 1,093 1,500
    NO T2I ✓ ✓ ✓ ✓ ✓ 0.235 1,008 755 95 389
    ALL ✓ ✓ ✓ ✓ ✓ ✓ 0.240 1,085 729 127 636
    ALL (8B) ✓ ✓ ✓ ✓ ✓ ✓ 0.305 355 687 4.7 13.76

    Key takeaways:

    • Pure self-supervised pretraining (SSL) fails on text-guided tasks (MSR-VTT CLIPSIM drops to 0.226, UCF-101 FVD degrades to 1,742), proving paired text data is essential.
    • The complete multitask mixture (ALL) provides the best overall generalist performance across diverse task modalities.
    • Scaling the model to 8B parameters with larger data compute yields dramatic improvements across all tasks, notably reducing SSv2 inpainting FVD from 127 to 4.7 and outpainting FVD from 636 to 13.76.
  7. Knowl 7 — Zero-Shot Video Stylization via Structural Conditioning

    model/method

    VideoPoet performs video stylization without requiring diffusion adapter modules (such as ControlNet cross-attention) or latent blending. Instead, stylization is formulated as a multimodal sequence-to-sequence translation task:

    1. Structural Feature Extraction: Dense optical flow is estimated using RAFT, and monocular depth maps are estimated using MiDaS. Flow and depth channels are normalized and concatenated along the channel dimension to match the 3-channel layout of RGB video.
    2. Tokenizer Compatibility: The concatenated optical flow and depth representation is tokenized directly using the frozen MAGVIT-v2 tokenizer without retraining or architectural modification.
    3. Inference Guidance: The structure tokens and new text style prompt are supplied as the prefix input to VideoPoet. The text prompt controls appearance/content (e.g., 'oil painting', 'cyberpunk', 'cartoon'), while optical flow and depth tokens enforce geometric structure and motion consistency.

    On the DAVIS 2016 benchmark across 20 videos and 40 text style prompts, VideoPoet achieved a frame-text CLIPSIM of 0.3417, outperforming Control-A-Video (CLIPSIM 0.3246). In side-by-side human evaluation, raters preferred VideoPoet over Control-A-Video by 77.5% on video quality and 70.0% on text fidelity.

  8. Knowl 8 — Human Side-by-Side Evaluation of VideoPoet against Video Diffusion Models

    empirical result

    A side-by-side human evaluation was conducted comparing VideoPoet (task-adapted 8B model with prompt expansion and fixed negative prompts) against seven state-of-the-art video models: Show-1, VideoCrafter, Phenaki, Pika, Runway Gen-2, WALT, and Lumiere. Evaluation was performed on a fixed prompt bank of ~250 prompts selected preferentially for explicit motion descriptions across 5 human judgment dimensions:

    1. Text Fidelity: VideoPoet was preferred over Phenaki (71% vs. 29%), Show-1 (61% vs. 39%), VideoCrafter (62% vs. 38%), Runway Gen-2 (72% vs. 28%), Pika (76% vs. 24%), and WALT (55% vs. 45%). VideoPoet was comparable to Lumiere (48% vs. 52%).
    2. Video Quality: VideoPoet won against Phenaki (76%), Show-1 (68%), VideoCrafter (60%), Runway Gen-2 (56%), Pika (74%), and WALT (61%). Lumiere was the only model preferred over VideoPoet (59% vs. 41%).
    3. Motion Interestingness: VideoPoet outperformed all baselines: Show-1 (72%), VideoCrafter (64%), Runway Gen-2 (82%), Pika (72%), WALT (66%), and Lumiere (65%), while remaining comparable to Phenaki (48% vs. 52%).
    4. Motion Realism: VideoPoet won against Phenaki (76%), Show-1 (58%), VideoCrafter (58%), Runway Gen-2 (58%), Pika (84%), and WALT (57%). Lumiere was preferred over VideoPoet (61% vs. 39%).
    5. Temporal Consistency: VideoPoet won against Phenaki (66%) and Pika (56%), while being less preferred against Show-1 (40%), VideoCrafter (40%), Runway Gen-2 (40%), WALT (36%), and Lumiere (37%).

    Overall, VideoPoet demonstrates competitive performance against top diffusion models, exhibiting particular strengths in motion magnitude, realism, and interestingness.

  9. Knowl 9 — Autoregressive Temporal Extension and Zero-Shot Task Chaining

    model/method

    Because VideoPoet uses an autoregressive decoder-only language model over causal visual tokens, it enables flexible video manipulation through temporal continuation and task composition:

    1. Long Video Generation: Videos longer than 10 seconds are synthesized iteratively by conditioning each subsequent generation on the final 1 second of visual tokens from the previous clip, maintaining temporal continuity without noticeable distortion.
    2. Image Animation (Image-to-Video): A still image is tokenized into a single frame (1,16,16)(1, 16, 16) without padding. The LLM predicts subsequent frames autoregressively, driven by text prompts.
    3. Zero-Shot Task Chaining: Multiple independent generative capabilities are chained sequentially in teacher-forced token space without custom adapter pipelines. For example:
      • An input still image is animated to video via Image-to-Video.
      • The resulting video is spatially outpainted to expand its aspect ratio.
      • The outpainted video is transformed with video-to-video stylization using text prompts, optical flow, and depth tokens. Because each generation stage produces clean discrete tokens that remain within the training distribution, artifacts do not compound across chained tasks.
  10. Knowl 10 — Inference Latency and Generation Runtime

    empirical result

    On a hardware cluster of TPUv5p (4 chips), VideoPoet's generation pipeline achieves an amortized runtime of approximately 5 seconds per 1 second of output video. For a batch size of 4 videos generating 17 frames at 8 fps (2.125 seconds total):

    • Base LLM Autoregressive Token Generation: 34 seconds.
    • Detokenizer (MAGVIT-v2 tokens to pixels): 1.3 seconds.
    • Super-Resolution (Cascade to 896×512896 \times 512): 6.8 seconds.
  11. Knowl 11 — Limitations of Discrete Token-Based Video Modeling

    limitation

    VideoPoet exhibits three primary limitations inherent to its token-based LLM formulation and training data:

    1. Compression Artifacts: The discrete quantization and spatiotemporal compression of the visual tokenizer (MAGVIT-v2) establish an intrinsic upper bound on fine visual fidelity during pixel reconstruction.
    2. Static Per-Frame Aesthetics: On still frames and static scenes, per-frame aesthetics lag behind state-of-the-art image diffusion models due to excluding curated copyrighted datasets (such as LAION) during pretraining.
    3. Fine-Grained Dynamic Details: Generating small objects and intricate textures undergoing rapid or large-magnitude motions remains challenging within discrete visual token representation.

Coverage note — None was omitted; all key contributions including architecture, tokenization, pretraining task mixture, super-resolution transformer, training curriculum, empirical benchmark evaluations, human studies, and capabilities have been covered.

References

  1. 1.Agostinelli, A., Denk, T. I., Borsos, Z., Engel, J., Verzetti, M., Caillon, A., Huang, Q., Jansen, A., Roberts, A., Tagliasacchi, M., et al. Musiclm: Generating music from text. arXiv preprint arXiv:2301.11325, 2023.
  2. 2.Akbari, H., Kondratyuk, D., Cui, Y., Hornung, R., Wang, H., and Adam, H. Alternating gradient descent and mixture-of-experts for integrated multimodal perception. arXiv preprint arXiv:2305.06324, 2023.
  3. 3.Anil, R., Dai, A. M., Firat, O., Johnson, M., Lepikhin, D., Passos, A., Shakeri, S., Taropa, E., Bailey, P., Chen, Z., et al. Palm 2 technical report. arXiv preprint arXiv:2305.10403, 2023.
  4. 4.Bar-Tal, O., Chefer, H., Tov, O., Herrmann, C., Paiss, R., Zada, S., Ephrat, A., Hur, J., Li, Y., Michaeli, T., et al. Lumiere: A space-time diffusion model for video generation. arXiv preprint arXiv:2401.12945, 2024.
  5. 5.Blattmann, A., Dockhorn, T., Kulal, S., Mendelevitch, D., Kilian, M., Lorenz, D., Levi, Y., English, Z., Voleti, V., Letts, A., et al. Stable video diffusion: Scaling latent video diffusion models to large datasets. arXiv preprint arXiv:2311.15127, 2023a.
  6. 6.Blattmann, A., Rombach, R., Ling, H., Dockhorn, T., Kim, S. W., Fidler, S., and Kreis, K. Align your latents: High-resolution video synthesis with latent diffusion models. In CVPR, pp. 22563–22575, 2023b.
  7. 7.Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258, 2021.
  8. 8.Brooks, T., Holynski, A., and Efros, A. A. Instructpix2pix: Learning to follow image editing instructions. In CVPR, pp. 18392–18402, 2023.
  9. 9.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. NeurIPS, 33: 1877–1901, 2020.
  10. 10.Carreira, J., Noland, E., Banki-Horvath, A., Hillier, C., and Zisserman, A. A short note about kinetics-600. arXiv preprint arXiv:1808.01340, 2018.
  11. 11.Ceylan, D., Huang, C.-H. P., and Mitra, N. J. Pix2video: Video editing using image diffusion. In CVPR, pp. 23206–23217, 2023.
  12. 12.Chai, W., Guo, X., Wang, G., and Lu, Y. Stablevideo: Text-driven consistency-aware diffusion video editing. In CVPR, pp. 23040–23050, 2023.
  13. 13.Chang, H., Zhang, H., Jiang, L., Liu, C., and Freeman, W. T. Maskgit: Masked generative image transformer. In CVPR, pp. 11315–11325, 2022.
  14. 14.Chang, H., Zhang, H., Barber, J., Maschinot, A., Lezama, J., Jiang, L., Yang, M.-H., Murphy, K., Freeman, W. T., Rubinstein, M., et al. Muse: Text-to-image generation via masked generative transformers. arXiv preprint arXiv:2301.00704, 2023.
  15. 15.Chen, H., Xia, M., He, Y., Zhang, Y., Cun, X., Yang, S., Xing, J., Liu, Y., Chen, Q., Wang, X., et al. Videocrafter1: Open diffusion models for high-quality video generation. arXiv preprint arXiv:2310.19512, 2023a.
  16. 16.Chen, W., Wu, J., Xie, P., Wu, H., Li, J., Xia, X., Xiao, X., and Lin, L. Control-a-video: Controllable text-to-video generation with diffusion models. arXiv preprint arXiv:2305.13840, 2023b.
  17. 17.Chiu, M.-C., Chen, P.-Y., and Ma, X. Better may not be fairer: A study on subgroup discrepancy in image classification. In ICCV, pp. 4956–4966, 2023.
  18. 18.Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al. PaLM: Scaling language modeling with pathways. arXiv:2204.02311, 2022.
  19. 19.Ding, M., Yang, Z., Hong, W., Zheng, W., Zhou, C., Yin, D., Lin, J., Zou, X., Shao, Z., Yang, H., et al. Cogview: Mastering text-to-image generation via transformers. NeurIPS, pp. 19822–19835, 2021.
  20. 20.Driess, D., Xia, F., Sajjadi, M. S., Lynch, C., Chowdhery, A., Ichter, B., Wahid, A., Tompson, J., Vuong, Q., Yu, T., et al. Palm-e: An embodied multimodal language model. arXiv preprint arXiv:2303.03378, 2023.
  21. 21.Du, N., Huang, Y., Dai, A. M., Tong, S., Lepikhin, D., Xu, Y., Krikun, M., Zhou, Y., Yu, A. W., Firat, O., et al. GLaMs: Efficient scaling of language models with mixture-of-experts. In ICML, 2022.
  22. 22.Esser, P., Rombach, R., and Ommer, B. Taming transformers for high-resolution image synthesis. In CVPR, pp. 12868–12878, 2020.
  23. 23.Esser, P., Chiu, J., Atighehchian, P., Granskog, J., and Germanidis, A. Structure and content-guided video synthesis with diffusion models. In CVPR, pp. 7346–7356, 2023.
  24. 24.Feng, R., Weng, W., Wang, Y., Yuan, Y., Bao, J., Luo, C., Chen, Z., and Guo, B. Ccedit: Creative and controllable video editing via diffusion models. arXiv preprint arXiv:2309.16496, 2023.
  25. 25.Ge, S., Nah, S., Liu, G., Poon, T., Tao, A., Catanzaro, B., Jacobs, D., Huang, J.-B., Liu, M.-Y., and Balaji, Y. Preserve your own correlation: A noise prior for video diffusion models. In CVPR, pp. 22930–22941, 2023.
  26. 26.Geyer, M., Bar-Tal, O., Bagon, S., and Dekel, T. Tokenflow: Consistent diffusion features for consistent video editing. arXiv preprint arXiv:2307.10373, 2023.
  27. 27.Goyal, R., Ebrahimi Kahou, S., Michalski, V., Materzynska, J., Westphal, S., Kim, H., Haenel, V., Fruend, I., Yianilos, P., Mueller-Freitag, M., et al. The “something something” video database for learning and evaluating visual common sense. In ICCV, 2017.
  28. 28.Guo, Y., Yang, C., Rao, A., Wang, Y., Qiao, Y., Lin, D., and Dai, B. Animatediff: Animate your personalized text-to-image diffusion models without specific tuning. arXiv preprint arXiv:2307.04725, 2023.
  29. 29.Gupta, A., Tian, S., Zhang, Y., Wu, J., Martın-Martın, R., and Fei-Fei, L. Maskvit: Masked visual pre-training for video prediction. arXiv preprint arXiv:2206.11894, 2022.
  30. 30.Gupta, A., Yu, L., Sohn, K., Gu, X., Hahn, M., Fei-Fei, L., Essa, I., Jiang, L., and Lezama, J. Photorealistic video generation with diffusion models. arXiv preprint arXiv:2312.06662, 2023.
  31. 31.He, Y., Yang, T., Zhang, Y., Shan, Y., and Chen, Q. Latent video diffusion models for high-fidelity long video generation. arXiv preprint arXiv:2211.13221, 2(3):4, 2023.
  32. 32.Hershey, S., Chaudhuri, S., Ellis, D. P. W., Gemmeke, J. F., Jansen, A., Moore, C., Plakal, M., Platt, D., Saurous, R. A., Seybold, B., Slaney, M., Weiss, R., and Wilson, K. Cnn architectures for large-scale audio classification. In ICASSP, 2017. URL https://arxiv.org/abs/1609.09430.
  33. 33.Ho, J. and Salimans, T. Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598, 2022.
  34. 34.Ho, J., Chan, W., Saharia, C., Whang, J., Gao, R., Gritsenko, A., Kingma, D. P., Poole, B., Norouzi, M., Fleet, D. J., et al. Imagen video: High definition video generation with diffusion models. arXiv preprint arXiv:2210.02303, 2022a.
  35. 35.Ho, J., Salimans, T., Gritsenko, A., Chan, W., Norouzi, M., and Fleet, D. J. Video diffusion models. arXiv:2204.03458, 2022b.
  36. 36.Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., et al. Training compute-optimal large language models. arXiv preprint arXiv:2203.15556, 2022.
  37. 37.Hong, W., Ding, M., Zheng, W., Liu, X., and Tang, J. Cogvideo: Large-scale pretraining for text-to-video generation via transformers. arXiv preprint arXiv:2205.15868, 2022.
  38. 38.Hu, A., Russell, L., Yeo, H., Murez, Z., Fedoseev, G., Kendall, A., Shotton, J., and Corrado, G. Gaia-1: A generative world model for autonomous driving. arXiv preprint arXiv:2309.17080, 2023.
  39. 39.Li, R., Allal, L. B., Zi, Y., Muennighoff, N., Kocetkov, D., Mou, C., Marone, M., Akiki, C., Li, J., Chim, J., et al. StarCoder: may the source be with you! arXiv:2305.06161, 2023.
  40. 40.Liew, J. H., Yan, H., Zhang, J., Xu, Z., and Feng, J. Magicedit: High-fidelity and temporally coherent video editing. arXiv preprint arXiv:2308.14749, 2023.
  41. 41.Meng, C., He, Y., Song, Y., Song, J., Wu, J., Zhu, J.-Y., and Ermon, S. Sdedit: Guided image synthesis and editing with stochastic differential equations. arXiv preprint arXiv:2108.01073, 2021.
  42. 42.Nash, C., Carreira, J., Walker, J., Barr, I., Jaegle, A., Malinowski, M., and Battaglia, P. Transframer: Arbitrary frame prediction with generative models. arXiv preprint arXiv:2203.09494, 2022.
  43. 43.OpenAI. GPT-4 technical report. arXiv:2303.08774, 2023.
  44. 44.Perazzi, F., Pont-Tuset, J., McWilliams, B., Van Gool, L., Gross, M., and Sorkine-Hornung, A. A benchmark dataset and evaluation methodology for video object segmentation. In CVPR, 2016.
  45. 45.Pika. Pika 1.0, 2023. URL https://pika.art/launch.
  46. 46.Podell, D., English, Z., Lacey, K., Blattmann, A., Dockhorn, T., Muller, J., Penna, J., and Rombach, R. Sdxl: Improving latent diffusion models for high-resolution image synthesis. arXiv preprint arXiv:2307.01952, 2023.
  47. 47.Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(1):5485–5551, 2020.
  48. 48.Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I. Zero-shot text-to-image generation. arXiv preprint arXiv:2102.12092, 2021.
  49. 49.Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 1(2):3, 2022.
  50. 50.Ranftl, R., Lasinger, K., Hafner, D., Schindler, K., and Koltun, V. Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer. IEEE TPAMI, 44(3):1623–1637, 2020.
  51. 51.Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. High-resolution image synthesis with latent diffusion models. In CVPR, pp. 10684–10695, 2022.
  52. 52.Rubenstein, P. K., Asawaroengchai, C., Nguyen, D. D., Bapna, A., Borsos, Z., Quitry, F. d. C., Chen, P., Badawy, D. E., Han, W., Kharitonov, E., et al. Audiopalm: A large language model that can speak and listen. arXiv preprint arXiv:2306.12925, 2023.
  53. 53.Runway. Gen2, 2023. URL https://runwayml.com/.
  54. 54.Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E. L., Ghasemipour, K., Gontijo Lopes, R., Karagol Ayan, B., Salimans, T., et al. Photorealistic text-to-image diffusion models with deep language understanding. NeurIPS, 35:36479–36494, 2022.
  55. 55.Saito, M., Saito, S., Koyama, M., and Kobayashi, S. Train sparsely, generate densely: Memory-efficient unsupervised training of high-resolution temporal gan. IJCV, 128(10):2586–2606, 2020.
  56. 56.Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., Wortsman, M., et al. Laion-5b: An open large-scale dataset for training next generation image-text models. Advances in Neural Information Processing Systems, 35:25278–25294, 2022.
  57. 57.Schumann, C., Ricco, S., Prabhu, U., Ferrari, V., and Pantofaru, C. A step toward more inclusive people annotations for fairness. In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society, pp. 916–925, 2021.
  58. 58.Schumann, C., Olanubi, G. O., Wright, A., Monk, E., Heldreth, C., and Ricco, S. Consensus and subjectivity of skin tone annotation for ML fairness. In Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2023. URL https://openreview.net/forum?id=L9I9FhHfS3.
  59. 59.Singer, U., Polyak, A., Hayes, T., Yin, X., An, J., Zhang, S., Hu, Q., Yang, H., Ashual, O., Gafni, O., et al. Make-a-video: Text-to-video generation without text-video data. arXiv preprint arXiv:2209.14792, 2022.
  60. 60.Soomro, K., Zamir, A. R., and Shah, M. Ucf101: A dataset of 101 human actions classes from videos in the wild. arXiv preprint arXiv:1212.0402, 2012.
  61. 61.Sun, D., Herrmann, C., Reda, F., Rubinstein, M., Fleet, D. J., and Freeman, W. T. Disentangling architecture and training for optical flow. In ECCV, 2022.
  62. 62.Tang, Z., Yang, Z., Zhu, C., Zeng, M., and Bansal, M. Any-to-any generation via composable diffusion. arXiv preprint arXiv:2305.11846, 2023.
  63. 63.Tu, Z., Talebi, H., Zhang, H., Yang, F., Milanfar, P., Bovik, A., and Li, Y. Maxvit: Multi-axis vision transformer. In ECCV, pp. 459–479, 2022.
  64. 64.Unterthiner, T., Van Steenkiste, S., Kurach, K., Marinier, R., Michalski, M., and Gelly, S. Towards accurate generative models of video: A new metric & challenges. arXiv preprint arXiv:1812.01717, 2018.
  65. 65.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is all you need. NeurIPS, 30, 2017.
  66. 66.Villegas, R., Babaeizadeh, M., Kindermans, P.-J., Moraldo, H., Zhang, H., Saffar, M. T., Castro, S., Kunze, J., and Erhan, D. Phenaki: Variable length video generation from open domain textual description. arXiv preprint arXiv:2210.02399, 2022.
  67. 67.Voleti, V., Jolicoeur-Martineau, A., and Pal, C. Mcvd-masked conditional video diffusion for prediction, generation, and interpolation. NeurIPS, 35:23371–23385, 2022.
  68. 68.Wang, J., Yuan, H., Chen, D., Zhang, Y., Wang, X., and Zhang, S. Modelscope text-to-video technical report. arXiv preprint arXiv:2308.06571, 2023a.
  69. 69.Wang, W., Xie, K., Liu, Z., Chen, H., Cao, Y., Wang, X., and Shen, C. Zero-shot video editing using off-the-shelf image diffusion models. arXiv preprint arXiv:2303.17599, 2023b.
  70. 70.Wang, W., Yang, H., Tuo, Z., He, H., Zhu, J., Fu, J., and Liu, J. Videofactory: Swap attention in spatiotemporal diffusions for text-to-video generation. arXiv preprint arXiv:2305.10874, 2023c.
  71. 71.Wang, Y., He, Y., Li, Y., Li, K., Yu, J., Ma, X., Chen, X., Wang, Y., Luo, P., Liu, Z., et al. Internvid: A large-scale video-text dataset for multimodal understanding and generation. arXiv preprint arXiv:2307.06942, 2023d.
  72. 72.Wu, C., Huang, L., Zhang, Q., Li, B., Ji, L., Yang, F., Sapiro, G., and Duan, N. Godiva: Generating open-domain videos from natural descriptions. arXiv preprint arXiv:2104.14806, 2021.
  73. 73.Xu, J., Mei, T., Yao, T., and Rui, Y. Msr-vtt: A large video description dataset for bridging video and language. In CVPR, pp. 5288–5296, 2016.
  74. 74.Yan, W., Zhang, Y., Abbeel, P., and Srinivas, A. Videogpt: Video generation using vq-vae and transformers. arXiv preprint arXiv:2104.10157, 2021.
  75. 75.Yu, J., Xu, Y., Koh, J. Y., Luong, T., Baid, G., Wang, Z., Vasudevan, V., Ku, A., Yang, Y., Ayan, B. K., et al. Scaling autoregressive models for content-rich text-to-image generation. arXiv preprint arXiv:2206.10789, 2022.
  76. 76.Yu, L., Cheng, Y., Sohn, K., Lezama, J., Zhang, H., Chang, H., Hauptmann, A. G., Yang, M.-H., Hao, Y., Essa, I., et al. MAGVIT: Masked generative video transformer. In CVPR, 2023a.
  77. 77.Yu, L., Cheng, Y., Wang, Z., Kumar, V., Macherey, W., Huang, Y., Ross, D. A., Essa, I., Bisk, Y., Yang, M.-H., et al. SPAE: Semantic pyramid autoencoder for multimodal generation with frozen llms. In NeurIPS, 2023b.
  78. 78.Yu, L., Lezama, J., Gundavarapu, N. B., Versari, L., Sohn, K., Minnen, D., Cheng, Y., Gupta, A., Gu, X., Hauptmann, A. G., et al. Language model beats diffusion–tokenizer is key to visual generation. In ICLR, 2024.
  79. 79.Yu, S., Sohn, K., Kim, S., and Shin, J. Video probabilistic diffusion models in projected latent space. In CVPR, pp. 18456–18466, 2023c.
  80. 80.Zeghidour, N., Luebs, A., Omran, A., Skoglund, J., and Tagliasacchi, M. Soundstream: An end-to-end neural audio codec. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 30:495–507, 2021.
  81. 81.Zeng, Y., Wei, G., Zheng, J., Zou, J., Wei, Y., Zhang, Y., and Li, H. Make pixels dance: High-dynamic video generation. arXiv preprint arXiv:2311.10982, 2023.
  82. 82.Zhang, D. J., Wu, J. Z., Liu, J.-W., Zhao, R., Ran, L., Gu, Y., Gao, D., and Shou, M. Z. Show-1: Marrying pixel and latent diffusion models for text-to-video generation. arXiv preprint arXiv:2309.15818, 2023a.
  83. 83.Zhang, L., Rao, A., and Agrawala, M. Adding conditional control to text-to-image diffusion models. In CVPR, pp. 3836–3847, 2023b.
  84. 84.Zhang, Y., Jiang, L., Turk, G., and Yang, D. Auditing gender presentation differences in text-to-image models. arXiv preprint arXiv:2302.03675, 2023c.
  85. 85.Zhou, C., Liu, P., Xu, P., Iyer, S., Sun, J., Mao, Y., Ma, X., Efrat, A., Yu, P., Yu, L., et al. Lima: Less is more for alignment. arXiv preprint arXiv:2305.11206, 2023.
  86. 86.Zhou, D., Wang, W., Yan, H., Lv, W., Zhu, Y., and Feng, J. Magicvideo: Efficient video generation with latent diffusion models. arXiv preprint arXiv:2211.11018, 2022.
  87. 87.Zitkovich, B., Yu, T., Xu, S., Xu, P., Xiao, T., Xia, F., Wu, J., Wohlhart, P., Welker, S., Wahid, A., et al. RT-2: Vision-language-action models transfer web knowledge to robotic control. In CoRL, 2023.

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/