YouTube-8M: A Large-Scale Video Classification Benchmark
Sami Abu-El-HaijaNisarg KothariJoonseok LeePaul NatsevGeorge TodericiBalakrishnan VaradarajanSudheendra Vijayanarasimhan
Introduces YouTube-8M, a massive multi-label benchmark spanning eight million videos and 4,800 visual entities, providing precomputed deep frame features and standard baselines to lower the barrier for large-scale video classification research.
While major advances in computer vision have been powered by massive image datasets, progress in automated video understanding has lagged due to the absence of similarly large, diverse benchmarks. Existing video resources are largely restricted to human actions or sports and contain fewer than 500 categories. Furthermore, storing and processing raw video at scale presents extreme computational and infrastructural hurdles for most research teams.
To solve this, the article introduces YouTube-8M, a public benchmark for large-scale multi-label video classification, and evaluates several baseline machine learning models on it. The dataset comprises more than 8 million videos—totaling over 500,000 hours of content—annotated with 4,800 visually recognizable entity classes organized across 24 top-level domains. To make the data computationally accessible, the authors extracted pre-computed frame-level features from 1.9 billion video frames using a deep image classification network, compressed the features eightfold via dimensionality reduction and quantization, and benchmarked various frame- and video-level classification architectures.
The findings demonstrate the power and practicality of this resource. First, the pre-computed features substantially lower technical barriers, enabling baseline models to converge in under a day on a single standard machine. Second, simple task-independent video-level representations (combining mean, standard deviation, and top ordinal statistics) paired with Mixture-of-Experts models achieved 30.0% mean Average Precision and 63.3% top-1 accuracy, matching or exceeding more complex sequence-based deep networks such as Long Short-Term Memory models. Third, representations pre-trained on YouTube-8M demonstrated strong transferability to other datasets, boosting performance on Sports-1M and establishing a new state of the art on ActivityNet by increasing mean Average Precision from 53.8% to 77.6%.
These results show that when underlying static frame features are sufficiently rich, high-performing video classifiers do not necessarily require computationally expensive temporal modeling over raw pixels. In practice, this dramatically reduces model training time, compute costs, and storage footprints while enabling high-accuracy video categorization across diverse real-world domains.
Organizations and researchers working on video analysis should adopt the pre-extracted features and open-source codebase to train classifiers efficiently without massive hardware investments. To build on this work, future development should focus on explicitly modeling noisy and missing annotations, incorporating compact audio and motion features, and exploring multiple-instance learning to better localize themes across video timelines.
The primary limitation of the benchmark is its reliance on machine-generated topic labels, which human evaluation showed to have high precision (78.8%) but low recall (14.5%), meaning many valid labels are absent. Additionally, because the provided features cap video analysis at the first six minutes and exclude explicit motion channels, performance on highly motion-dependent tasks may require cautious interpretation.
- Paper: UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild, Khurram Soomro et al. (2012). UCF101 established the foundational benchmark and methodology for unconstrained video action classification from YouTube footage, contextualizing the scale and multi-label shift in YouTube-8M.
- Paper: ActivityNet: A large-scale video benchmark for human activity understanding, Fabian Caba Heilbron et al. (2015). ActivityNet provides the core precedent for large-scale untrimmed web video categorization, directly motivating the development of massive-scale video benchmarks like YouTube-8M.
- Paper: Beyond short snippets: Deep networks for video classification, Joe Yue-Hei Ng et al. (2015). This work demonstrates how deep frame-level CNN feature pooling and LSTM modeling across entire video sequences achieve high-accuracy video classification, inspiring the baseline pipeline used in YouTube-8M.
- Paper: ImageNet Large Scale Visual Recognition Challenge, Olga Russakovsky et al. (2014). The ImageNet challenge defined large-scale visual recognition benchmarking and pre-trained CNN backbones, which YouTube-8M explicitly references and utilizes for frame-level feature extraction.
- Paper: Two-Stream Convolutional Networks for Action Recognition in Videos, Karen Simonyan et al. (2014). This seminal paper introduces deep convolutional representations for action recognition in video, establishing the deep learning baseline paradigms extended by YouTube-8M.
- Paper: Learning Spatiotemporal Features with 3D Convolutional Networks, Du Tran et al. (2015). Du Tran et al. demonstrate how spatiotemporal 3D convolutional representations can be learned directly from large-scale supervised video data like Sports-1M.
- Paper: Unsupervised Learning of Video Representations using LSTMs, Nitish Srivastava et al. (2015). This paper establishes sequence-to-sequence LSTM architectures for extracting representation vectors from YouTube video streams.
- Paper: Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset, João Carreira et al. (2017). Carreira and Zisserman introduce the Kinetics dataset and Inflated 3D ConvNet (I3D) architecture, extending large-scale video pretraining and benchmarking following YouTube-8M.
- Paper: The Kinetics Human Action Video Dataset, Will Kay et al. (2017). The Kinetics benchmark builds upon YouTube-8M's philosophy of large-scale YouTube data curation to provide trimmed video action classification at scale.
- Paper: Can Spatiotemporal 3D CNNs Retrace the History of 2D CNNs and ImageNet?, Kensho Hara et al. (2017). This empirical study investigates whether large-scale video datasets enable training deep 3D CNNs from scratch without overfitting, answering questions raised by the scale of YouTube-8M.
- Paper: HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips, Antoine Miech et al. (2019). HowTo100M scales YouTube-based video representation learning from millions of visual entity labels to joint text-video embeddings learned across hundreds of millions of narrated clips.
- Paper: VideoBERT: A Joint Model for Video and Language Representation Learning, Chen Sun et al. (2019). VideoBERT extends large-scale YouTube video learning to multimodal transformer pre-training across visual tokens and speech transcriptions.
- Paper: Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval, Max Bain et al. (2021). Frozen in Time presents an end-to-end visual transformer architecture for video-text retrieval trained on millions of web video-text pairs.
- Paper: Revisiting Unreasonable Effectiveness of Data in Deep Learning Era, Chen Sun et al. (2017). This work explores how scaling weakly labeled internet datasets to hundreds of millions of examples systematically improves deep visual representations.
- Paper: Exploring the Limits of Weakly Supervised Pretraining, Dhruv Mahajan et al. (2018). Mahajan et al. push the boundaries of weakly supervised large-scale pretraining using billions of social media samples with noisy multi-label metadata.
