Towards Accurate Generative Models of Video: A New Metric & Challenges

Thomas UnterthinerSjoerd van SteenkisteKarol KurachRaphael MarinierMarcin MichalskiSylvain Gelly

article2018arXiv1,414 citations

Introduces Fréchet Video Distance (FVD), a metric validated against human judgment to evaluate visual quality, temporal coherence, and sample diversity in video generation, alongside the StarCraft 2 Videos benchmark.

Listen

Deep generative video models offer substantial potential for forecasting, simulation, and visual reasoning, but development has been constrained by inadequate evaluation standards and a lack of scalable benchmarks. Existing evaluation methods, such as Structural Similarity (SSIM) and Peak Signal to Noise Ratio (PSNR), evaluate videos frame by frame. Consequently, they fail to measure temporal continuity or visual diversity across entire video distributions. Furthermore, these traditional metrics require direct frame-by-frame alignment with a ground-truth reference, making them unusable for evaluating open-ended, unconditional video generation.

To address these limitations, the article introduces Fréchet Video Distance (FVD), a metric designed to evaluate the visual quality, temporal realism, and sample diversity of generated video sequences, alongside StarCraft 2 Videos (SCV), a standardized suite of four synthetic benchmark datasets designed to assess complex reasoning and temporal memory. The evaluation framework utilizes an Inflated 3D Convolutional Network (I3D) trained on human action recognition to compute feature distances across video distributions. The study evaluates these tools through comprehensive noise sensitivity experiments, an extensive human evaluation study comparing thousands of model variations, and baseline benchmarking of state-of-the-art video architectures.

Key findings demonstrate that FVD substantially outperforms conventional metrics in reflecting human judgment. In controlled tests where models exhibited identical scores on legacy metrics, FVD aligned with human preferences in 74.9% to 81.0% of comparisons, whereas SSIM and PSNR failed to differentiate model quality. Human agreement with FVD rankings escalates rapidly once model scores diverge by more than 50 points, establishing a clear threshold for meaningful visual improvements. Furthermore, baseline evaluations across the StarCraft 2 suite revealed that current leading video generation architectures fail to handle complex multi-agent interactions or maintain consistency over extended time horizons, frequently generating blurry artifacts or failing to remember tracked entities.

These results show that relying on legacy frame-level metrics introduces significant risk of misjudging model performance, misallocating development resources, and failing to detect critical motion distortions. By providing a dependable distribution-level assessment that operates with or without ground-truth sequences, FVD enables objective, reproducible comparisons across different generative architectures. Concurrently, the StarCraft 2 benchmark provides controlled environments with customizable complexity, allowing developers to isolate and diagnose specific failure modes such as relational reasoning and long-term retention.

Organizations developing or deploying video generation systems should transition from frame-by-frame metrics to FVD as the primary evaluation standard while maintaining consistent sample sizes to ensure valid comparisons. When benchmarking model capabilities, teams should adopt the scalable StarCraft 2 environment to stress-test complex entity tracking and temporal dynamics before deploying models to complex real-world tasks. Although FVD requires a fixed sample size to avoid estimation bias and synthetic benchmarks cannot fully replace real-world video complexity, the experimental results provide high confidence that FVD and the StarCraft 2 suite establish a reliable foundation for monitoring and advancing generative video technology.

arXiv: 1812.01717
  • Paper: Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset, João Carreira et al. (2017). Introduces the Inflated 3D ConvNet (I3D) architecture and Kinetics dataset, providing the foundational spatiotemporal feature representation upon which Fréchet Video Distance (FVD) is computed.
  • Paper: MoCoGAN: Decomposing Motion and Content for Video Generation, Sergey Tulyakov et al. (2018). Presents early deep generative modeling architectures that decompose video into motion and content, highlighting the baseline methodologies and evaluation shortcomings addressed by FVD.
  • Paper: Generating Videos with Scene Dynamics, Carl Vondrick et al. (2016). Establishes fundamental generative adversarial frameworks for synthesizing video dynamics, contextualizing the initial attempts at video generation that required better qualitative metrics.
  • Paper: Video Pixel Networks, Nal Kalchbrenner et al. (2017). Develops deep autoregressive models for raw video prediction, serving as a primary generative baseline and illustrating the need for sample-quality metrics beyond pixel-level loss.
  • Paper: Deep multi-scale video prediction beyond mean square error, Michael Mathieu et al. (2015). Addresses the limitations of standard error metrics in predictive video synthesis using adversarial formulations, motivating the need for holistic distribution-level metrics like FVD.
  • Paper: Learning Spatiotemporal Features with 3D Convolutional Networks, Du Tran et al. (2015). Provides early 3D convolutional architectures for capturing spatiotemporal video features, establishing the technical background for deep video representations used in generative evaluation.
Cover for Towards Accurate Generative Models of Video: A New Metric & Challenges

Abstract

Recent advances in deep generative models have lead to remarkable progress in synthesizing high quality images. Following their successful application in image processing and representation learning, an important next step is to consider videos. Learning generative models of video is a much harder task, requiring a model to capture the temporal dynamics of a scene, in addition to the visual presentation of objects. While recent attempts at formulating generative models of video have had some success, current progress is hampered by (1) the lack of qualitative metrics that consider visual quality, temporal coherence, and diversity of samples, and (2) the wide gap between purely synthetic video data sets and challenging real-world data sets in terms of complexity. To this extent we propose Fréchet Video Distance (FVD), a new metric for generative models of video, and StarCraft 2 Videos (SCV), a benchmark of game play from custom starcraft 2 scenarios that challenge the current capabilities of generative models of video. We contribute a large-scale human study, which confirms that FVD correlates well with qualitative human judgment of generated videos, and provide initial benchmark results on SCV.

Table of Contents

  • 1 Introduction
  • 2 Fréchet Video Distance
  • 3 Starcraft 2 Videos
  • 4 Experiments
  • 4.1 Noise Study
  • 4.2 Effect of Sample Size on FVD
  • 4.3 Human Evaluation
  • 4.3.1 Resolution of FVD
  • 4.4 Baseline Results on SC2 Benchmark Data Sets
  • 4.5 Correlation of FVD with SSIM and PSNR
  • 5 Conclusion
  • References
  • A Noise Study
  • B SCV Data Generation
  • B.1 Scenario Parameters
  • C Benchmark Hyperparameters
  • D Examples of Video Models on SCV

Knowls

  1. Knowl 1 — Fréchet Video Distance

    definition

    The Fréchet Video Distance (FVD) is an evaluation metric for video generative models that assesses visual quality, temporal coherence, and sample diversity by computing the 2-Wasserstein (Fréchet) distance between feature representations of real and generated video distributions modeled as multivariate Gaussians:

    FVD(PR,PG)=∥μR−μG∥22+Tr(ΣR+ΣG−2(ΣRΣG)1/2)\text{FVD}(P_R, P_G) = \|\mu_R - \mu_G\|_2^2 + \text{Tr}\left(\Sigma_R + \Sigma_G - 2(\Sigma_R \Sigma_G)^{1/2}\right)

    where μR,ΣR\mu_R, \Sigma_R and μG,ΣG\mu_G, \Sigma_G denote the empirical mean vectors and covariance matrices of video feature representations sampled from the real data distribution PRP_R and generated model distribution PGP_G, respectively.

    The feature representation is extracted from the logits layer (pre-classification layer) of an Inflated 3D ConvNet (I3D) pre-trained on RGB video sequences of the Kinetics-400 action recognition dataset. Unlike frame-by-frame metrics such as PSNR or SSIM, FVD operates over entire video distributions, capturing spatiotemporal dynamics without requiring point-to-point paired ground-truth videos, making it applicable to unconditional video synthesis.

  2. Knowl 2 — Kernel Video Distance

    model/method

    Kernel Video Distance (KVD) is a non-parametric alternative to Fréchet Video Distance that computes the Maximum Mean Discrepancy (MMD) between empirical video feature distributions without assuming Gaussianity.

    Given samples x1,…,xm∈Rdx_1, \dots, x_m \in \mathbb{R}^d from real videos PRP_R and y1,…,yn∈Rdy_1, \dots, y_n \in \mathbb{R}^d from generated videos PGP_G embedded via a pre-trained I3D network, the unbiased estimator of the squared MMD distance is:

    MMD2(X,Y)=1m(m−1)∑i=1m∑j≠imk(xi,xj)−2mn∑i=1m∑i=1nk(xi,yj)+1n(n−1)∑i=1n∑j≠ink(yi,yj)\text{MMD}^2(X, Y) = \frac{1}{m(m - 1)} \sum_{i=1}^m \sum_{j \neq i}^m k(x_i, x_j) - \frac{2}{mn} \sum_{i=1}^m \sum_{i=1}^n k(x_i, y_j) + \frac{1}{n(n - 1)} \sum_{i=1}^n \sum_{j \neq i}^n k(y_i, y_j)

    where k(a,b)=(aTb+1)3k(a, b) = (a^T b + 1)^3 is a polynomial kernel evaluating the similarity between the feature representations.

  3. Knowl 3 — Human Evaluation and Correlation with Video Quality Metrics

    data/table

    A large-scale human evaluation study evaluated how well automatic metrics (FVD, KVD, per-frame FID averaged over time, SSIM, and PSNR) align with human perceptual judgments of video quality. Two experimental regimes were evaluated across video prediction models (CDNA, SV2P, SVP-FP, SAVP) trained on the BAIR robot pushing dataset:

    1. Equal Metric (eq.): Pairs of models that have identical scores according to one metric are evaluated by human raters to test if other metrics or humans can distinguish quality differences.
    2. Spread Metric (spr.): Models selected across deciles (10th to 90th percentile) of a metric are compared to test whether ranking according to that metric agrees with human choices.
    Metric eq. FVD eq. SSIM eq. PSNR eq. KVD spr. FVD spr. SSIM spr. PSNR spr. KVD
    FVD N/A 74.9% 81.0% 63.0% 71.9% 58.4% 63.5% 63.1%
    SSIM 51.5% N/A 44.6% 43.6% 61.8% 51.2% 45.9% 50.2%
    PSNR 56.3% 21.4% N/A 48.8% 54.1% 37.0% 44.8% 54.1%
    KVD 40.6% 70.4% 73.8% N/A 69.4% 56.8% 63.8% 59.1%
    Avg. FID 35.5% 71.2% 52.0% 43.5% 62.4% 62.7% 57.6% 51.2%
    Among raters 79.3% 77.8% 84.4% 74.3% 83.3% 69.9% 72.5% 74.1%

    FVD demonstrates the highest agreement with human judgments across both spread and equal conditions. When models achieve identical SSIM or PSNR scores, FVD successfully distinguishes between their subjective visual quality with 74.9% and 81.0% human agreement, respectively, while neither SSIM nor PSNR can reliably separate models when FVD is held equal.

  4. Knowl 4 — Perceptual Resolution Threshold of Fréchet Video Distance

    empirical result

    Human perceptual distinction between video prediction models depends systematically on the numerical difference in their Fréchet Video Distance (FVD) scores:

    1. When the difference in FVD between two models is less than 50 points (ΔFVD<50\Delta \text{FVD} < 50), human rater agreement with the lower-FVD model is near chance level (50%50\% to 60%60\%).
    2. When the difference in FVD exceeds 50 points (ΔFVD≥50\Delta \text{FVD} \ge 50), human agreement rises rapidly, surpassing 80%80\% at ΔFVD≈200\Delta \text{FVD} \approx 200 and approaching 95%95\% at ΔFVD≥300\Delta \text{FVD} \ge 300.

    Consequently, an FVD margin of at least 50 points is required to reliably indicate human-perceptible quality differences between generative video models.

  5. Knowl 5 — StarCraft 2 Videos Benchmark Suite

    definition

    StarCraft 2 Videos (SCV) is a synthetic video generation benchmark created using the StarCraft II Learning Environment (SC2LE). Each dataset consists of 10,000 training, 2,000 validation, and 2,000 test video replays available at 64×6464 \times 64 and 128×128128 \times 128 spatial resolutions, organized into four distinct scenarios:

    1. Move Unit to Border (MUtB): A single unit randomly selected from 6 types (Marine, Siege Tank, Drone, Zergling, Colossus, Archon) and 4 colors spawns in the center and moves toward a border destination; videos span 11–27 frames (recorded every 6th game frame). Tests basic unit movement animations and trajectory modeling.
    2. Collect Mineral Shards (CMS): Two Marine units greedily collect 20 randomly distributed mineral shards; videos span 99 frames (recorded every 4th game frame). Tests multi-agent path planning, state updates (disappearing shards upon arrival), and sequential interactions.
    3. Brawl: Two opposing 9-unit armies (Terran vs. Zerg) engage in combat with varied unit types, attack animations, projectiles, health pools, and positions; videos span up to 99 frames (recorded every 4th game frame). Tests multi-entity dynamics, non-rigid motion, and fine-grained visual interactions.
    4. Road Trip with Medivac (RTwM): A flying Medivac unit collects 1 to 4 units of varying types/colors, visits two intermediate beacons, and unloads the identical units at a final beacon; videos span 32 frames (recorded every 8th game frame). Tests long-term memory, object permanence, and temporal consistency.
  6. Knowl 6 — Baseline Performance of Generative Video Models on BAIR, KTH, and SCV

    data/table

    Four conditional video prediction architectures—Convolutional Dynamic Neural Advection (CDNA), Stochastic Variational Video Prediction (SV2P), Stochastic Video Generation with Fixed Prior (SVP-FP), and Stochastic Adversarial Video Prediction (SAVP)—were evaluated across BAIR, KTH, and all four SCV benchmark tasks using FVD (lower is better):

    Model BAIR KTH SCV-MUtB SCV-CMS SCV-Brawl SCV-RTwM
    64×6464\times 64 128×128128\times 128 64×6464\times 64 128×128128\times 128 64×6464\times 64 128×128128\times 128 64×6464\times 64 128×128128\times 128
    CDNA 296.5 150.8 486.1 51.4 440.8 515.3 877.1 1016.6 1089.3 1295.4
    SV2P 262.5 136.8 423.9 710.5 430.4 316.0 859.9 995.9 1068.7 1026.1
    SVP-FP 315.5 208.4 276.7 121.3 379.8 442.1 714.5 1240.7 1022.9 2031.4
    SAVP 116.4 78.0 479.7 204.4 188.8 192.5 192.9 150.3 698.6 1055.4

    All models were conditioned on 2 context frames to predict 14 frames (BAIR, MUtB, CMS, Brawl), 10 context frames to predict 10 frames (KTH), or 2 context frames to predict up to 32 frames (RTwM). FVD was computed using 256 samples on BAIR and 1024 samples on KTH and SCV.

    SAVP achieved the best overall FVD scores across most benchmarks due to adversarial loss terms. While models achieved reasonable results on MUtB, all models failed to model multi-unit interactions and unit identity in Brawl and completely failed to solve the long-term temporal dependencies in RTwM.

  7. Knowl 7 — Sensitivity of Video Feature Embeddings to Static and Temporal Perturbations

    empirical result

    Evaluation of FVD across synthetic noise corruptions on BAIR, HMDB51, and Kinetics-400 demonstrates distinct sensitivity profiles:

    1. Static Perturbations: Black rectangle occlusions, Gaussian blur, Gaussian noise, and salt & pepper noise are detected monotonically by both 2D frame-level Inception embeddings and 3D spatio-temporal (I3D) embeddings.
    2. Temporal Perturbations: Frame-level distortions including local adjacent frame swapping, global sequence frame swapping, interleaving frames from multiple videos, and abrupt switching between videos are poorly captured by 2D Inception frame averages, but are detected with strong monotonic rank correlation by 3D I3D embeddings.

    Among I3D configurations, features extracted from the pre-classification logits layer of an I3D model pre-trained on Kinetics-400 RGB frames yield the highest rank correlation across all noise types.

  8. Knowl 8 — Sample-Size Bias in Empirical Fréchet Video Distance Estimation

    limitation

    Estimating the mean vectors and covariance matrices from a finite sample of NN video sequences introduces an empirical estimation bias in Fréchet Video Distance. Even when evaluating two non-overlapping subsets drawn from the exact same underlying distribution PRP_R, the empirical FVD is strictly greater than zero and scales inversely with sample size (e.g., FVD ≈500\approx 500 at N=32N=32, dropping to ≈20\approx 20 at N=4096N=4096 on BAIR). Consequently, quantitative FVD comparisons across different models are only valid when evaluated using an identical number of sample sequences.

  9. Knowl 9 — Correlation Structure Between FVD, SSIM, and PSNR

    empirical result

    Across an evaluation of more than 20,000 video prediction model checkpoints, the empirical correlation between metrics exhibits the following structure:

    1. SSIM vs. PSNR: High positive correlation (Pearson's r=0.730r = 0.730, Kendall's τ=0.648\tau = 0.648), reflecting their shared formulation based on pairwise frame-by-frame pixel differences against ground-truth frames.
    2. SSIM vs. FVD: Moderate, statistically significant negative correlation (Pearson's r=−0.640r = -0.640, Kendall's τ=−0.189\tau = -0.189), showing that SSIM captures some overlapping visual degradations but diverges on temporal coherence.
    3. PSNR vs. FVD: Weak correlation (Pearson's r=−0.278r = -0.278, Kendall's τ=−0.007\tau = -0.007), indicating that PSNR is ineffective at measuring distribution-level video realism and temporal consistency.

Coverage note — None was omitted; all primary contributions—including the mathematical definition and kernel variants of FVD, the SCV benchmark suite and scenarios, human evaluation studies, noise sensitivity benchmarks, resolution thresholds, and state-of-the-art model baseline evaluations—are fully covered.

References

  1. 1.M. Babaeizadeh, C. Finn, D. Erhan, R. H. Campbell, and S. Levine. Stochastic variational video prediction. International Conference on Learning Representations (ICLR), 2017.
  2. 2.M. Bińkowski, D. J. Sutherland, M. Arbel, and A. Gretton. Demystifying MMD GANs. International Conference on Learning Representations (ICLR), 2018.
  3. 3.A. Brock, J. Donahue, and K. Simonyan. Large scale gan training for high fidelity natural image synthesis. arXiv, 2018.
  4. 4.W. Byeon, Q. Wang, R. K. Srivastava, P. Koumoutsakos, P. Vlachas, Z. Wan, T. Sapsis, F. Raue, S. Palacio, T. Breuel, et al. Contextvp: Fully context-aware video prediction. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, pages 1122–1126, 2018.
  5. 5.J. Carreira and A. Zisserman. Quo vadis, action recognition? a new model and the kinetics dataset. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017.
  6. 6.J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei. Imagenet: A large-scale hierarchical image database. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2009.
  7. 7.E. Denton and R. Fergus. Stochastic video generation with a learned prior. International Conference on Machine Learning (ICML), 2018.
  8. 8.D. Dowson and B. Landau. The frchet distance between multivariate normal distributions. Journal of Multivariate Analysis, 12(3):450 – 455, 1982.
  9. 9.F. Ebert, C. Finn, A. X. Lee, and S. Levine. Self-supervised visual planning with temporal skip connections. In Conference on Robot Learning, pages 344–356, 2017.
  10. 10.C. Finn, I. Goodfellow, and S. Levine. Unsupervised learning for physical interaction through video prediction. Advances in Neural Information Processing Systems (NIPS), 2016.
  11. 11.I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio. Generative adversarial nets. Advances in neural information processing systems (NIPS), 2014.
  12. 12.A. Gretton, K. M. Borgwardt, M. J. Rasch, B. Schölkopf, and A. Smola. A kernel two-sample test. Journal of Machine Learning Research, 13(Mar):723–773, 2012.
  13. 13.E. Haller and M. Leordeanu. Unsupervised object segmentation in video by efficient selection of highly probable positive features. IEEE International Conference on Computer Vision (ICCV), 2017.
  14. 14.M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. Advances in Neural Information Processing Systems (NIPS), 2017.
  15. 15.S. Hochreiter and J. Schmidhuber. Long short-term memory. Neural computation, 9(8):1735–1780, 1997.
  16. 16.Q. Huynh-Thu and M. Ghanbari. The accuracy of psnr in predicting video quality for different video scenes and frame rates. Telecommunication Systems, 2012.
  17. 17.P. Isola, J.-Y. Zhu, T. Zhou, and A. A. Efros. Image-to-image translation with conditional adversarial networks. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017.
  18. 18.H. Jiang, D. Sun, V. Jampani, M.-H. Yang, E. Learned-Miller, and J. Kautz. Super slomo: High quality estimation of multiple intermediate frames for video interpolation. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018.
  19. 19.N. Kalchbrenner, A. van den Oord, K. Simonyan, I. Danihelka, O. Vinyals, A. Graves, and K. Kavukcuoglu. Video pixel networks. International Conference on Machine Learning (ICML), 2017.
  20. 20.T. Karras, T. Aila, S. Laine, and J. Lehtinen. Progressive growing of GANs for improved quality, stability, and variation. International Conference on Learning Representations (ICLR), 2018.
  21. 21.W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijayanarasimhan, F. Viola, T. Green, T. Back, P. Natsev, M. Suleyman, and A. Zisserman. The kinetics human action video dataset. arXiv, 2017.
  22. 22.H. Kuehne, H. Jhuang, R. Stiefelhagen, and T. Serre. Hmdb51: A large video database for human motion recognition. In High Performance Computing in Science and Engineering 12, pages 571–582. Springer, 2013.
  23. 23.B. M. Lake, T. D. Ullman, J. B. Tenenbaum, and S. J. Gershman. Building machines that learn and think like people. Behavioral and Brain Sciences, 40, 2017.
  24. 24.A. X. Lee, R. Zhang, F. Ebert, P. Abbeel, C. Finn, and S. Levine. Stochastic adversarial video prediction. arXiv, 2018.
  25. 25.A. Lerer, S. Gross, and R. Fergus. Learning physical intuition of block towers by example. In Proceedings of the 33rd International Conference on International Conference on Machine Learning-Volume 48, pages 430–438. JMLR. org, 2016.
  26. 26.B. Lotter, G. Kreiman, and D. Cox. Deep Predictive Coding Networks for Video Prediction and Unsupervised Learning. International Conference on Learning Representations (ICLR), 2017.
  27. 27.P. Luc, C. Couprie, S. Chintala, and J. Verbeek. Semantic segmentation using adversarial networks. arXiv, 2016.
  28. 28.M. Lucic, K. Kurach, M. Michalski, S. Gelly, and O. Bousquet. Are gans created equal? a large-scale study. Advances in Neural Information Processing Systems (NIPS), 2018.
  29. 29.M. Mathieu, C. Couprie, and Y. LeCun. Deep multi-scale video prediction beyond mean square error. International Conference on Learning Representations (ICLR), 2016.
  30. 30.N. Ponomarenko, L. Jin, O. Ieremeiev, V. Lukin, K. Egiazarian, J. Astola, B. Vozel, K. Chehdi, M. Carli, F. Battisti, et al. Image database tid2013: Peculiarities, results and perspectives. Signal Processing: Image Communication, 30:57–77, 2015.
  31. 31.K. Preuer, P. Renz, T. Unterthiner, S. Hochreiter, and G. Klambauer. Frchet chemnet distance: A metric for generative models for molecules in drug discovery. Journal of Chemical Information and Modeling, 58(9):1736–1741, 2018.
  32. 32.I. Radosavovic, P. Dollr, R. Girshick, G. Gkioxari, and K. He. Data distillation: Towards omni-supervised learning. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018.
  33. 33.M. Ranzato, A. Szlam, J. Bruna, M. Mathieu, R. Collobert, and S. Chopra. Video (language) modeling: a baseline for generative models of natural videos. arXiv, 2014.
  34. 34.M. Saito, E. Matsumoto, and S. Saito. Temporal generative adversarial nets with singular value clipping. International Conference on Computer Vision (ICCV), 2017.
  35. 35.M. S. Sajjadi, B. Schölkopf, and M. Hirsch. Enhancenet: Single image super-resolution through automated texture synthesis. International Conference on Computer Vision (ICCV), 2017.
  36. 36.C. Schuldt, I. Laptev, and B. Caputo. Recognizing human actions: a local svm approach. In Proceedings of the 17th International Conference on Pattern Recognition (ICPR), volume 3, pages 32–36. IEEE, 2004.
  37. 37.K. Soomro, A. R. Zamir, and M. Shah. Ucf101: A dataset of 101 human actions classes from videos in the wild. arXiv, 2012.
  38. 38.N. Srivastava, E. Mansimov, and R. Salakhudinov. Unsupervised learning of video representations using lstms. International Conference on Machine Learning (ICML), 2015.
  39. 39.C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna. Rethinking the inception architecture for computer vision. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016.
  40. 40.L. Theis, A. van den Oord, and M. Bethge. A note on the evaluation of generative models. International Conference on Learning Representations (ICLR), 2016.
  41. 41.S. Tulyakov, M.-Y. Liu, X. Yang, and J. Kautz. Mocogan: Decomposing motion and content for video generation. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018.
  42. 42.T. Unterthiner, B. Nessler, C. Seward, G. Klambauer, M. Heusel, H. Ramsauer, and S. Hochreiter. Coulomb GANs: Provably Optimal Nash Equilibria via Potential Fields. International Conference on Learning Representations (ICLR), 2018.
  43. 43.A. Vaswani, S. Bengio, E. Brevdo, F. Chollet, A. N. Gomez, S. Gouws, L. Jones, L. Kaiser, N. Kalchbrenner, N. Parmar, R. Sepassi, N. Shazeer, and J. Uszkoreit. Tensor2tensor for neural machine translation. arXiv, 2018.
  44. 44.R. Villegas, D. Erhan, H. Lee, et al. Hierarchical long-term video prediction without supervision. In International Conference on Machine Learning, pages 6033–6041, 2018.
  45. 45.O. Vinyals, T. Ewalds, S. Bartunov, P. Georgiev, A. S. Vezhnevets, M. Yeo, A. Makhzani, H. Kuttler, J. Agapiou, J. Schrittwieser, et al. StarCraft II: A new challenge for reinforcement learning. arXiv, 2017.
  46. 46.C. Vondrick, H. Pirsiavash, and A. Torralba. Generating videos with scene dynamics. Proceedings of the 30th International Conference on Neural Information Processing Systems (NIPS), 2016.
  47. 47.T.-C. Wang, M.-Y. Liu, J.-Y. Zhu, N. Yakovenko, A. Tao, J. Kautz, and B. Catanzaro. Video-to-video synthesis. In Advances in Neural Information Processing Systems, pages 1152–1164, 2018.
  48. 48.Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli. Image quality assessment: from error visibility to structural similarity. IEEE Transactions on Image Processing, 2004.
  49. 49.R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang. The unreasonable effectiveness of deep features as a perceptual metric. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018.

Citation

MLA
Unterthiner, T., et al. “Towards Accurate Generative Models of Video: A New Metric & Challenges”. arXiv, 2018, http://arxiv.org/abs/1812.01717v2.
APA
Unterthiner, T., Steenkiste, S. van ., Kurach, K., Marinier, R., Michalski, M., & Gelly, S. (2018). Towards Accurate Generative Models of Video: A New Metric & Challenges. arXiv. http://arxiv.org/abs/1812.01717v2
Chicago
Unterthiner, T., S. van . Steenkiste, K. Kurach, R. Marinier, M. Michalski, and S. Gelly. 2018. “Towards Accurate Generative Models of Video: A New Metric & Challenges”. arXiv. http://arxiv.org/abs/1812.01717v2.
Harvard
Unterthiner, T. et al. (2018) “Towards Accurate Generative Models of Video: A New Metric & Challenges”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1812.01717v2.
Vancouver
1. Unterthiner T, Steenkiste S van, Kurach K, Marinier R, Michalski M, Gelly S (2018) Towards Accurate Generative Models of Video: A New Metric & Challenges. arXiv

BibTeX

@article{unterthiner2018towards,
  title = {Towards Accurate Generative Models of Video: A New Metric & Challenges},
  author = {Unterthiner, Thomas and Steenkiste, Sjoerd van and Kurach, Karol and Marinier, Raphael and Michalski, Marcin and Gelly, Sylvain},
  year = {2018},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1812.01717v2},
  eprint = {1812.01717}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/