Towards Accurate Generative Models of Video: A New Metric & Challenges
Thomas UnterthinerSjoerd van SteenkisteKarol KurachRaphael MarinierMarcin MichalskiSylvain Gelly
Introduces Fréchet Video Distance (FVD), a metric validated against human judgment to evaluate visual quality, temporal coherence, and sample diversity in video generation, alongside the StarCraft 2 Videos benchmark.
Deep generative video models offer substantial potential for forecasting, simulation, and visual reasoning, but development has been constrained by inadequate evaluation standards and a lack of scalable benchmarks. Existing evaluation methods, such as Structural Similarity (SSIM) and Peak Signal to Noise Ratio (PSNR), evaluate videos frame by frame. Consequently, they fail to measure temporal continuity or visual diversity across entire video distributions. Furthermore, these traditional metrics require direct frame-by-frame alignment with a ground-truth reference, making them unusable for evaluating open-ended, unconditional video generation.
To address these limitations, the article introduces Fréchet Video Distance (FVD), a metric designed to evaluate the visual quality, temporal realism, and sample diversity of generated video sequences, alongside StarCraft 2 Videos (SCV), a standardized suite of four synthetic benchmark datasets designed to assess complex reasoning and temporal memory. The evaluation framework utilizes an Inflated 3D Convolutional Network (I3D) trained on human action recognition to compute feature distances across video distributions. The study evaluates these tools through comprehensive noise sensitivity experiments, an extensive human evaluation study comparing thousands of model variations, and baseline benchmarking of state-of-the-art video architectures.
Key findings demonstrate that FVD substantially outperforms conventional metrics in reflecting human judgment. In controlled tests where models exhibited identical scores on legacy metrics, FVD aligned with human preferences in 74.9% to 81.0% of comparisons, whereas SSIM and PSNR failed to differentiate model quality. Human agreement with FVD rankings escalates rapidly once model scores diverge by more than 50 points, establishing a clear threshold for meaningful visual improvements. Furthermore, baseline evaluations across the StarCraft 2 suite revealed that current leading video generation architectures fail to handle complex multi-agent interactions or maintain consistency over extended time horizons, frequently generating blurry artifacts or failing to remember tracked entities.
These results show that relying on legacy frame-level metrics introduces significant risk of misjudging model performance, misallocating development resources, and failing to detect critical motion distortions. By providing a dependable distribution-level assessment that operates with or without ground-truth sequences, FVD enables objective, reproducible comparisons across different generative architectures. Concurrently, the StarCraft 2 benchmark provides controlled environments with customizable complexity, allowing developers to isolate and diagnose specific failure modes such as relational reasoning and long-term retention.
Organizations developing or deploying video generation systems should transition from frame-by-frame metrics to FVD as the primary evaluation standard while maintaining consistent sample sizes to ensure valid comparisons. When benchmarking model capabilities, teams should adopt the scalable StarCraft 2 environment to stress-test complex entity tracking and temporal dynamics before deploying models to complex real-world tasks. Although FVD requires a fixed sample size to avoid estimation bias and synthetic benchmarks cannot fully replace real-world video complexity, the experimental results provide high confidence that FVD and the StarCraft 2 suite establish a reliable foundation for monitoring and advancing generative video technology.
- Paper: Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset, João Carreira et al. (2017). Introduces the Inflated 3D ConvNet (I3D) architecture and Kinetics dataset, providing the foundational spatiotemporal feature representation upon which Fréchet Video Distance (FVD) is computed.
- Paper: MoCoGAN: Decomposing Motion and Content for Video Generation, Sergey Tulyakov et al. (2018). Presents early deep generative modeling architectures that decompose video into motion and content, highlighting the baseline methodologies and evaluation shortcomings addressed by FVD.
- Paper: Generating Videos with Scene Dynamics, Carl Vondrick et al. (2016). Establishes fundamental generative adversarial frameworks for synthesizing video dynamics, contextualizing the initial attempts at video generation that required better qualitative metrics.
- Paper: Video Pixel Networks, Nal Kalchbrenner et al. (2017). Develops deep autoregressive models for raw video prediction, serving as a primary generative baseline and illustrating the need for sample-quality metrics beyond pixel-level loss.
- Paper: Deep multi-scale video prediction beyond mean square error, Michael Mathieu et al. (2015). Addresses the limitations of standard error metrics in predictive video synthesis using adversarial formulations, motivating the need for holistic distribution-level metrics like FVD.
- Paper: Learning Spatiotemporal Features with 3D Convolutional Networks, Du Tran et al. (2015). Provides early 3D convolutional architectures for capturing spatiotemporal video features, establishing the technical background for deep video representations used in generative evaluation.
- Paper: Adversarial Video Generation on Complex Datasets, Aidan Clark et al. (2019). Applies and builds directly on the evaluation paradigm of Fréchet Video Distance to scale generative adversarial networks on complex real-world video datasets.
- Paper: VideoGPT: Video Generation using VQ-VAE and Transformers, Wilson Yan et al. (2021). Utilizes FVD as a primary evaluation benchmark to measure visual fidelity and temporal coherence in transformer-based video synthesis.
- Paper: Video Diffusion Models, Jonathan Ho et al. (2022). Extends diffusion architectures to video generation, relying heavily on FVD as a standard metric to quantify improvements over GAN-based video models.
- Paper: VBench: Comprehensive Benchmark Suite for Video Generative Models, Ziqi Huang et al. (2023). Expands on metric-driven evaluation frameworks like FVD by proposing a multi-dimensional benchmark suite to systematically dissect video generative quality.
- Paper: Align Your Latents: High-Resolution Video Synthesis with Latent Diffusion Models, Andreas Blattmann et al. (2023). Employs FVD to benchmark high-resolution latent video diffusion models and demonstrate superior motion consistency.
- Paper: Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets, Andreas Blattmann et al. (2023). Demonstrates large-scale dataset curation and latent diffusion modeling, reporting FVD scores to evaluate zero-shot text-to-video performance.
- Paper: Make-A-Video: Text-to-Video Generation without Text-Video Data, Uriel Singer et al. (2023). Applies space-time factorized diffusion layers for text-to-video generation and uses FVD to demonstrate state-of-the-art visual and temporal realism.
- Paper: Photorealistic Video Generation with Diffusion Models, Agrim Gupta et al. (2024). Leverages transformer-based diffusion architectures for photorealistic video generation, evaluating temporal quality and realism using FVD.
- Paper: Lumiere: A Space-Time Diffusion Model for Video Generation, Omer Bar-Tal et al. (2024). Develops a unified space-time U-Net architecture for coherent video synthesis and benchmarks sample quality against prior models using FVD.
- Paper: Phenaki: Variable Length Video Generation From Open Domain Textual Description, Ruben Villegas et al. (2023). Generates variable-length open-domain videos using causal transformers, employing FVD to assess temporal coherence and benchmark performance against prior baselines.
