YouTube-8M: A Large-Scale Video Classification Benchmark

Sami Abu-El-HaijaNisarg KothariJoonseok LeePaul NatsevGeorge TodericiBalakrishnan VaradarajanSudheendra Vijayanarasimhan

article2016arXiv1,420 citations

Introduces YouTube-8M, a massive multi-label benchmark spanning eight million videos and 4,800 visual entities, providing precomputed deep frame features and standard baselines to lower the barrier for large-scale video classification research.

Listen

While major advances in computer vision have been powered by massive image datasets, progress in automated video understanding has lagged due to the absence of similarly large, diverse benchmarks. Existing video resources are largely restricted to human actions or sports and contain fewer than 500 categories. Furthermore, storing and processing raw video at scale presents extreme computational and infrastructural hurdles for most research teams.

To solve this, the article introduces YouTube-8M, a public benchmark for large-scale multi-label video classification, and evaluates several baseline machine learning models on it. The dataset comprises more than 8 million videos—totaling over 500,000 hours of content—annotated with 4,800 visually recognizable entity classes organized across 24 top-level domains. To make the data computationally accessible, the authors extracted pre-computed frame-level features from 1.9 billion video frames using a deep image classification network, compressed the features eightfold via dimensionality reduction and quantization, and benchmarked various frame- and video-level classification architectures.

The findings demonstrate the power and practicality of this resource. First, the pre-computed features substantially lower technical barriers, enabling baseline models to converge in under a day on a single standard machine. Second, simple task-independent video-level representations (combining mean, standard deviation, and top ordinal statistics) paired with Mixture-of-Experts models achieved 30.0% mean Average Precision and 63.3% top-1 accuracy, matching or exceeding more complex sequence-based deep networks such as Long Short-Term Memory models. Third, representations pre-trained on YouTube-8M demonstrated strong transferability to other datasets, boosting performance on Sports-1M and establishing a new state of the art on ActivityNet by increasing mean Average Precision from 53.8% to 77.6%.

These results show that when underlying static frame features are sufficiently rich, high-performing video classifiers do not necessarily require computationally expensive temporal modeling over raw pixels. In practice, this dramatically reduces model training time, compute costs, and storage footprints while enabling high-accuracy video categorization across diverse real-world domains.

Organizations and researchers working on video analysis should adopt the pre-extracted features and open-source codebase to train classifiers efficiently without massive hardware investments. To build on this work, future development should focus on explicitly modeling noisy and missing annotations, incorporating compact audio and motion features, and exploring multiple-instance learning to better localize themes across video timelines.

The primary limitation of the benchmark is its reliance on machine-generated topic labels, which human evaluation showed to have high precision (78.8%) but low recall (14.5%), meaning many valid labels are absent. Additionally, because the provided features cap video analysis at the first six minutes and exclude explicit motion channels, performance on highly motion-dependent tasks may require cautious interpretation.

arXiv: 1609.08675
Cover for YouTube-8M: A Large-Scale Video Classification Benchmark

Abstract

Many recent advancements in Computer Vision are attributed to large datasets. Open-source software packages for Machine Learning and inexpensive commodity hardware have reduced the barrier of entry for exploring novel approaches at scale. It is possible to train models over millions of examples within a few days. Although large-scale datasets exist for image understanding, such as ImageNet, there are no comparable size video classification datasets.

In this paper, we introduce YouTube-8M, the largest multi-label video classification dataset, composed of ~8 million videos (500K hours of video), annotated with a vocabulary of 4800 visual entities. To get the videos and their labels, we used a YouTube video annotation system, which labels videos with their main topics. While the labels are machine-generated, they have high-precision and are derived from a variety of human-based signals including metadata and query click signals. We filtered the video labels (Knowledge Graph entities) using both automated and manual curation strategies, including asking human raters if the labels are visually recognizable. Then, we decoded each video at one-frame-per-second, and used a Deep CNN pre-trained on ImageNet to extract the hidden representation immediately prior to the classification layer. Finally, we compressed the frame features and make both the features and video-level labels available for download.

We trained various (modest) classification models on the dataset, evaluated them using popular evaluation metrics, and report them as baselines. Despite the size of the dataset, some of our models train to convergence in less than a day on a single machine using TensorFlow. We plan to release code for training a TensorFlow model and for computing metrics.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 YouTube-8M Dataset
  • 3.1 Vocabulary Construction
  • 3.2 Collecting Videos
  • 3.3 Features
  • 3.4 Dataset Statistics
  • 3.5 Human Rated Test Set
  • 4 Baseline Approaches
  • 4.1 Models from Frame Features
  • 4.1.1 Frame-Level Models and Average Pooling
  • 4.1.2 Deep Bag of Frame (DBoF) Pooling
  • 4.1.3 Long Short-Term Memory (LSTM)
  • 4.2 Video level representations
  • 4.2.1 First, second order and ordinal statistics
  • 4.2.2 Feature normalization
  • 4.3 Models from Video Features
  • 4.3.1 Logistic Regression
  • 4.3.2 Hinge Loss
  • 4.3.3 Mixture of Experts (MoE)
  • 5 Experiments
  • 5.1 Evaluation Metrics
  • 5.2 Results on YouTube-8M
  • 5.2.1 Human Rated Test Set
  • 5.3 Results on Sports-1M
  • 5.4 Results on ActivityNet
  • 6 Conclusions
  • References

Knowls

  1. Knowl 1 — YouTube-8M Dataset Specification and Curation Pipeline

    definition

    YouTube-8M is a large-scale multi-label video dataset consisting of 8,264,650 videos totaling over 500,000 hours of video (mean duration of 229.6 seconds). The dataset vocabulary contains 4,800 visually recognizable Knowledge Graph entities spanning 24 top-level categories, with an average of 1.8 entity labels per video.

    The dataset creation pipeline consists of the following steps:

    1. Initial Whitelist/Blacklist Filtering: Starting from millions of Knowledge Graph topics, entities were filtered using a manually curated whitelist of 25 visual types (e.g., sport, tourist attraction, food, animal) and a blacklist of non-visual types (e.g., music compositions, software, album), yielding ∼50,000\sim 50,000 candidate entities.
    2. Visualness and Recognizability Curation: Three human raters scored each entity on a 1-to-5 scale for visual recognizability by laypersons. Entities with an average score ≤2.5\le 2.5 were retained (∼10,000\sim 10,000 entities).
    3. Video Sampling: Candidate videos having at least 1,000 views and a duration between 120 and 500 seconds were retrieved and labeled using the YouTube video annotation system.
    4. Frequency Thresholding: Entities with fewer than 200 associated videos and videos without any remaining valid entity labels were removed, resulting in the final 4,800 classes and 8,264,650 videos.
    5. Dataset Splitting: Videos were randomly split into Train (70%, 5,786,881 videos), Validate (20%, 1,652,167 videos), and Test (10%, 825,602 videos).
  2. Knowl 2 — YouTube-8M Frame Feature Extraction and Compression Pipeline

    model/method

    To enable practical computation over 500,000 hours of video, standard frame-level features were pre-extracted and compressed across 1.9 billion video frames as follows:

    1. Frame Decoding: Each video is decoded at 1 frame per second for up to the first 360 seconds (6 minutes).
    2. Deep Feature Extraction: Each decoded frame is passed through an Inception network pre-trained on ImageNet. The 2048-dimensional ReLU activations from the layer immediately prior to classification (pool_3/_reshape) are extracted.
    3. Dimensionality Reduction: Principal Component Analysis (PCA) with whitening reduces feature dimensionality from 2048 to 1024. The PCA projection matrix and mean vector are computed over all frames in the Training split.
    4. Non-Uniform Quantization: Each 32-bit floating point coordinate is quantized into an 8-bit unsigned integer (256 distinct values) using optimally computed non-uniform quantization bin boundaries.

    This two-step compression achieves an 8-fold data size reduction. Reconstructing and training models on uncompressed 32-bit 2048-dimensional features changes evaluation metrics by less than 1% compared to training on the compressed features.

  3. Knowl 3 — Video-Level Statistical and Ordinal Feature Aggregation

    model/method

    Let a video vv of length FvF_v frames be represented by a sequence of 1024-dimensional frame features x1v,x2v,…,xFvvx^v_1, x^v_2, \dots, x^v_{F_v}, where xjv∈R1024x^v_j \in \mathbb{R}^{1024}. A fixed-length unsupervised video descriptor ϕ(x1:Fvv)\phi(x^v_{1:F_v}) is constructed by computing statistics across frames reconstructed in the unquantized ReLU activation space:

    ϕ(x1:Fvv)=[μ(x1:Fvv)σ(x1:Fvv)TopK(x1:Fvv)]\phi(x^v_{1:F_v}) = \begin{bmatrix} \mu(x^v_{1:F_v}) \\ \sigma(x^v_{1:F_v}) \\ \text{Top}_K(x^v_{1:F_v}) \end{bmatrix}

    where:

    • μ(x1:Fvv)=1Fv∑j=1Fvxjv∈R1024\mu(x^v_{1:F_v}) = \frac{1}{F_v} \sum_{j=1}^{F_v} x^v_j \in \mathbb{R}^{1024} is the mean feature vector.
    • σ(x1:Fvv)∈R1024\sigma(x^v_{1:F_v}) \in \mathbb{R}^{1024} is the per-dimension standard deviation across all frames.
    • TopK(x1:Fvv)∈R1024⋅K\text{Top}_K(x^v_{1:F_v}) \in \mathbb{R}^{1024 \cdot K} is formed by concatenating the top KK ordinal values (highest values) along each dimension over the video. With K=5K=5, this yields a 5120-dimensional ordinal vector and a 7168-dimensional total aggregated descriptor.

    Prior to classification, global normalization is applied to ϕ(x1:Fvv)\phi(x^v_{1:F_v}): mean subtraction, PCA decorrelation and whitening, and L2L_2 normalization.

  4. Knowl 4 — Deep Bag of Frames (DBoF) Video Pooling Model

    model/method

    The Deep Bag of Frames (DBoF) architecture maps kk randomly sampled frame-level features x1v,…,xkv∈RNx^v_1, \dots, x^v_k \in \mathbb{R}^{N} (N=1024N = 1024) of a video vv into a single video-level prediction:

    1. Up-Projection: Each frame feature xjvx^v_j is mapped through a shared fully connected layer of MM units (M=8192M = 8192) with ReLU activations to produce an MM-dimensional sparse code.
    2. Batch Normalization & Pooling: A batch normalization layer is applied to the sparse codes, followed by element-wise max pooling across the kk frames to create a single MM-dimensional video representation.
    3. Classification: The pooled representation is passed through a fully connected hidden layer of 1024 units with ReLU activations, followed by a final logistic or softmax classification layer.

    The parameters are trained end-to-end using Stochastic Gradient Descent with AdaGrad, a learning rate of 0.1, and an L2L_2 weight decay penalty of 0.0005.

  5. Knowl 5 — Mixture of Experts (MoE) Video Classifier

    model/method

    The Mixture of Experts (MoE) classifier models the probability of entity ee given an aggregated video feature x∈RDx \in \mathbb{R}^D as a linear combination of ∣He∣|\mathcal{H}_e| expert logistic models gated by a softmax routing distribution:

    p(e∣x)=∑h∈Hep(h∣x)σ(uhTx)p(e \mid x) = \sum_{h \in \mathcal{H}_e} p(h \mid x) \sigma(u_h^T x)

    where σ(z)=1/(1+exp⁡(−z))\sigma(z) = 1 / (1 + \exp(-z)), uh∈RDu_h \in \mathbb{R}^D are the logistic weights of expert hh, and the gating probability p(h∣x)p(h \mid x) over ∣He∣+1|\mathcal{H}_e| + 1 states is:

    p(h∣x)=exp⁡(whTx)1+∑h′∈Heexp⁡(wh′Tx)p(h \mid x) = \frac{\exp(w_h^T x)}{1 + \sum_{h' \in \mathcal{H}_e} \exp(w_{h'}^T x)}

    The (∣He∣+1)(|\mathcal{H}_e|+1)-th state is a fixed non-existence state that produces output 0. Given ground truth g∈{0,1}g \in \{0, 1\}, predicted probability py∣x=p(y=1∣x)p_{y \mid x} = p(y=1 \mid x), expert probability py∣h,x=σ(uhTx)p_{y \mid h, x} = \sigma(u_h^T x), and gating probability ph∣x=p(h∣x)p_{h \mid x} = p(h \mid x), the log-loss is L(py∣x,g)=−glog⁡py∣x−(1−g)log⁡(1−py∣x)\mathcal{L}(p_{y \mid x}, g) = -g \log p_{y \mid x} - (1-g) \log(1 - p_{y \mid x}). The analytic gradients are:

    ∂L(py∣x,g)∂wh=xph∣x(py∣h,x−py∣x)(py∣x−g)py∣x(1−py∣x)\frac{\partial \mathcal{L}(p_{y \mid x}, g)}{\partial w_h} = x \frac{p_{h \mid x} (p_{y \mid h, x} - p_{y \mid x}) (p_{y \mid x} - g)}{p_{y \mid x} (1 - p_{y \mid x})}

    ∂L(py∣x,g)∂uh=xph∣xpy∣h,x(1−py∣h,x)(py∣x−g)py∣x(1−py∣x)\frac{\partial \mathcal{L}(p_{y \mid x}, g)}{\partial u_h} = x \frac{p_{h \mid x} p_{y \mid h, x} (1 - p_{y \mid h, x}) (p_{y \mid x} - g)}{p_{y \mid x} (1 - p_{y \mid x})}

    For 4,800 one-vs-all entity classifiers, independent 2-expert MoE models (∣He∣=2|\mathcal{H}_e|=2) are trained using AdaGrad with a learning rate of 1.0 and batch size 32.

  6. Knowl 6 — Long Short-Term Memory (LSTM) Video Modeling

    model/method

    Sequence modeling directly on sequential 1024-dimensional frame features uses a stacked Recurrent Neural Network with the following configuration:

    • Network Structure: 2 stacked LSTM layers, each with 1024 hidden units.
    • Temporal Unrolling: The network is unrolled for 60 iterations during training, establishing a temporal gradient horizon of 60 seconds.
    • Per-Frame Loss Weighting: At each unrolled step t∈{1,…,N}t \in \{1, \dots, N\} (where N=60N=60), the per-frame loss weight is set to wt=t/Nw_t = t / N, increasing linearly from 1/N1/N for the first step to 1.01.0 at the final step.
    • Inference: The concatenated hidden states of both LSTM layers at the final video frame are used as the video descriptor for the final classification layer.
  7. Knowl 7 — Video Multi-Label Classification Evaluation Metrics

    definition

    Multi-label video classification models are evaluated using three primary metrics across a video set VV where each video vv has ground-truth binary label set GvG_v and predicted entity rankings rankv,e\text{rank}_{v,e}:

    1. Mean Average Precision (mAP): For each entity, prediction scores are quantized into buckets of 10−410^{-4}. At a score threshold τ\tau, precision P(τ)P(\tau) and recall R(τ)R(\tau) across videos t∈Tt \in T with binary ground truth gtg_t and predicted score yty_t are:

    P(τ)=∑t∈TI(yt≥τ)gt∑t∈TI(yt≥τ),R(τ)=∑t∈TI(yt≥τ)gt∑t∈TgtP(\tau) = \frac{\sum_{t \in T} \mathbb{I}(y_t \ge \tau) g_t}{\sum_{t \in T} \mathbb{I}(y_t \ge \tau)}, \quad R(\tau) = \frac{\sum_{t \in T} \mathbb{I}(y_t \ge \tau) g_t}{\sum_{t \in T} g_t}

    Average precision is approximated over discrete thresholds τj=j/10000\tau_j = j / 10000 (j=1,…,10000j=1, \dots, 10000) by:

    AP=∑j=110000P(τj)[R(τj)−R(τj+1)]\text{AP} = \sum_{j=1}^{10000} P(\tau_j) [R(\tau_j) - R(\tau_{j+1})]

    The metric mAP is the unweighted arithmetic mean of AP across all classes.

    1. Hit@kk: The proportion of test videos for which at least one ground-truth label appears in the top kk model predictions:

    Hit@k=1∣V∣∑v∈V⋁e∈GvI(rankv,e≤k)\text{Hit@}k = \frac{1}{|V|} \sum_{v \in V} \bigvee_{e \in G_v} \mathbb{I}(\text{rank}_{v,e} \le k)

    1. Precision at Equal Recall Rate (PERR): The precision measured when the number of retrieved top predictions for video vv equals the total number of ground-truth labels ∣Gv∣|G_v|:

    PERR=1∣{v∈V:∣Gv∣>0}∣∑v∈V:∣Gv∣>01∣Gv∣∑e∈GvI(rankv,e≤∣Gv∣)\text{PERR} = \frac{1}{|\{v \in V : |G_v| > 0\}|} \sum_{v \in V : |G_v| > 0} \frac{1}{|G_v|} \sum_{e \in G_v} \mathbb{I}(\text{rank}_{v,e} \le |G_v|)

  8. Knowl 8 — Benchmark Baseline Results on YouTube-8M

    data/table

    The performance of baseline video classification models on the YouTube-8M validation set (1,652,167 videos) is summarized below:

    Input Features Modeling Approach mAP (%) Hit@1 (%) PERR (%)
    Frame-level, {x1:Fvv}\{x^v_{1:F_v}\} Logistic + Average Pooling 11.0 50.8 42.2
    Frame-level, {x1:Fvv}\{x^v_{1:F_v}\} Deep Bag of Frames (DBoF) 26.9 62.7 55.1
    Frame-level, {x1:Fvv}\{x^v_{1:F_v}\} LSTM (2-layer, 1024-dim) 26.6 64.5 57.3
    Video-level, μ\mu Online Hinge Loss (SVM) 17.0 56.3 47.9
    Video-level, μ\mu Logistic Regression 28.1 60.5 53.0
    Video-level, μ\mu Mixture-of-2-Experts 29.6 62.3 54.9
    Video-level, [μ;σ;Top5][\mu; \sigma; \text{Top}_5] Mixture-of-2-Experts 30.0 63.3 55.8

    On the exhaustively human-rated test set (>8,000 videos), the top three methods achieve:

    • Mixture-of-2-Experts ([μ;σ;Top5][\mu; \sigma; \text{Top}_5]): Hit@1 = 70.1%, PERR = 29.1%, Hit@5 = 84.8%.
    • LSTM: Hit@1 = 69.1%, PERR = 30.5%, Hit@5 = 84.7%.
    • DBoF: Hit@1 = 68.6%, PERR = 29.0%, Hit@5 = 83.5%.

    These results demonstrate that independent binary Mixture of Experts trained on summary video-level statistics perform competitively with or better than complex sequence (LSTM) and bag-of-frames (DBoF) deep networks on mAP, while naive frame-level prediction averaging fails significantly.

  9. Knowl 9 — Precision and Recall Characteristics of Machine Ground Truth Labels

    empirical result

    Exhaustive evaluation by 3 human raters per video on a uniform sample of over 8,000 test partition videos determined the quality of the automated YouTube topic annotation system ground-truth labels:

    • Precision: 78.8% relative to human consensus (comparable to typical inter-rater human agreement of ∼80%\sim 80\% on video annotation tasks).
    • Recall: 14.5% relative to exhaustive human labeling.

    The low recall arises because the annotation system identifies only the primary central topics of a video rather than exhaustively labeling all visible objects and actions. Consequently, the benchmark constitutes an evaluation test bed for algorithms handling high-precision but incomplete (missing) training labels.

  10. Knowl 10 — Transfer Learning Performance on ActivityNet and Sports-1M

    data/table

    Feature representations and models pre-trained on YouTube-8M transfer effectively to other video understanding benchmarks.

    On ActivityNet (untrimmed video classification):

    Approach mAP (%) Hit@1 (%) Hit@5 (%)
    Heilbron et al. (2015) 43.0 - -
    Ma et al. (2015) 53.8 - -
    Mixture-of-2-Experts (μ\mu, ActivityNet PCA) 69.1 68.7 85.4
    + Pretrained PCA on YT-8M 74.1 72.5 89.3
    Mixture-of-2-Experts ([μ;σ;Top5][\mu; \sigma; \text{Top}_5], ActivityNet PCA) 74.2 72.3 89.6
    + Pretrained PCA on YT-8M 77.6 74.9 91.6
    LSTM (trained from scratch on ActivityNet) 57.9 63.4 81.0
    LSTM + Pretrained on YT-8M (fixed weights) 75.6 74.2 92.4

    Transferring the YouTube-8M pre-trained PCA and representation improves the state-of-the-art mAP on ActivityNet from 53.8% to 77.6% without using motion or optical flow features.

    On Sports-1M (1.2 million videos, 487 sports classes):

    • Logistic Regression (μ\mu): mAP = 58.0%, Hit@1 = 60.1%, Hit@5 = 79.6%.
    • Mixture-of-2-Experts ([μ;σ;Top5][\mu; \sigma; \text{Top}_5]): mAP = 61.3%, Hit@1 = 63.2%, Hit@5 = 82.6%.
    • LSTM (trained from scratch on Sports-1M): mAP = 66.7%, Hit@1 = 64.9%, Hit@5 = 85.6%.
    • LSTM + Pretrained on YT-8M and fine-tuned: mAP = 67.6%, Hit@1 = 65.7%, Hit@5 = 86.2%.

Coverage note — None was omitted; all key contributed definitions, architectures (DBoF, LSTM, MoE, feature statistics), metrics, dataset statistics, and empirical evaluation/transfer results are fully covered.

References

  1. 1.Freebase: A community-curated database of well-known people, places, and things. https://www.freebase.com.
  2. 2.Google I/O 2013 - semantic video annotations in the Youtube Topics API: Theory and applications. https://www.youtube.com/watch?v=wf_77z1H-vQ.
  3. 3.Knowledge Graph Search API. https://developers.google.com/knowledge-graph/.
  4. 4.Tensorflow: Image recognition. https://www.tensorflow.org/tutorials/image_recognition.
  5. 5.M. Blank, L. Gorelick, E. Shechtman, M. Irani, and R. Basri. Actions as space-time shapes. In Proceedings of the International Conference on Computer Vision (ICCV), 2005.
  6. 6.J. Deng, W. Dong, R. Socher, L. jia Li, K. Li, and L. Fei-fei. Imagenet: A large-scale hierarchical image database. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2009.
  7. 7.M. Everingham, L. V. Gool, C. K. I. Williams, J. Winn, and A. Zisserman. The pascal visual object classes (voc) challenge, 2009.
  8. 8.L. Fei-fei, R. Fergus, and P. Perona. One-shot learning of object categories. IEEE Transactions on Pattern Analysis and Machine Intelligence, 28, 2006.
  9. 9.R. Girshick. Fast R-CNN. In Proceedings of the International Conference on Computer Vision (ICCV), 2015.
  10. 10.G. Griffin, A. Holub, and P. Perona. Caltech-256 object category dataset. Technical Report 7694, California Institute of Technology, 2007.
  11. 11.K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. CoRR, abs/1512.03385, 2015.
  12. 12.F. C. Heilbron, V. Escorcia, B. Ghanem, and J. C. Niebles. Activitynet: A large-scale video benchmark for human activity understanding. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 961–970, 2015.
  13. 13.S. Hochreiter and J. Schmidhuber. Long short-term memory. Neural Computing, 9(8), Nov. 1997.
  14. 14.S. Ioffe and C. Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In Proceedings of the International Conference on Machine Learning (ICML), pages 448–456, 2015.
  15. 15.H. Jegou, F. Perronnin, M. Douze, J. Sanchez, P. Perez, and C. Schmid. Aggregating local image descriptors into compact codes. IEEE Trans. Pattern Anal. Mach. Intell., 34(9), Sept. 2012.
  16. 16.Y. Jiang, J. Liu, A. Roshan Zamir, G. Toderici, I. Laptev, M. Shah, and R. Sukthankar. THUMOS challenge: Action recognition with a large number of classes. http://crcv.ucf.edu/THUMOS14, 2014.
  17. 17.Y.-G. Jiang, Z. Wu, J. Wang, X. Xue, and S.-F. Chang. Exploiting feature and class relationships in video categorization with regularized deep neural networks. arXiv preprint arXiv:1502.07209, 2015.
  18. 18.M. I. Jordan. Hierarchical mixtures of experts and the em algorithm. Neural Computation, 6, 1994.
  19. 19.A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei. Large-scale video classification with convolutional neural networks. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 1725–1732, Columbus, Ohio, USA, 2014.
  20. 20.A. Krizhevsky, I. Sutskever, and G. E. Hinton. ImageNet classification with deep convolutional neural networks. In Advances in Neural Information Processing Systems (NIPS), pages 1097–1105, 2012.
  21. 21.H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre. Hmdb: a large video database for human motion recognition. In Proceedings of the International Conference on Computer Vision (ICCV), 2011.
  22. 22.I. Laptev and T. Lindeberg. Space-time interest points. In Proceedings of the International Conference on Computer Vision (ICCV), 2003.
  23. 23.I. Laptev, M. Marszalek, C. Schmid, and B. Rozenfeld. Learning realistic human actions from movies. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2008.
  24. 24.S. Ma, S. A. Bargal, J. Zhang, L. Sigal, and S. Sclaroff. Do less and achieve more: Training cnns for action recognition utilizing action images from the web. CoRR, abs/1512.07155, 2015.
  25. 25.V. Mnih and G. Hinton. Learning to label aerial images from noisy data. In Proceedings of the 29th Annual International Conference on Machine Learning (ICML), June 2012.
  26. 26.J. Y.-H. Ng, M. J. Hausknecht, S. Vijayanarasimhan, O. Vinyals, R. Monga, and G. Toderici. Beyond short snippets: Deep networks for video classification. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 4694–4702, 2015.
  27. 27.F. Perronnin and C. Dance. Fisher kernels on visual vocabularies for image categorization. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2007.
  28. 28.A. Quattoni and A. Torralba. Recognizing indoor scenes. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2009.
  29. 29.S. Reed, H. Lee, D. Anguelov, C. Szegedy, D. Erhan, and A. Rabinovich. Training deep neural networks on noisy labels with bootstrapping. ArXiv e-prints, Dec. 2014.
  30. 30.O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei. ImageNet Large Scale Visual Recognition Challenge. International Journal of Computer Vision (IJCV), 115(3):211–252, 2015.
  31. 31.P. Sermanet, D. Eigen, X. Zhang, M. Mathieu, R. Fergus, and Y. LeCun. Overfeat: Integrated recognition, localization and detection using convolutional networks. In International Conference on Learning Representations (ICLR).
  32. 32.J. Shotton, J. Winn, C. Rother, and A. Criminisi. Textonboost: Joint appearance, shape and context modeling for multi-class object. In Proceedings of the European Conference on Computer Vision (ECCV), 2006.
  33. 33.K. Soomro, A. R. Zamir, and M. Shah. UCF101: A dataset of 101 human actions classes from videos in the wild. In CRCV-TR-12-01, 2012.
  34. 34.B. Thomee, D. A. Shamma, G. Friedland, B. Elizalde, K. Ni, D. Poland, D. Borth, and L. Li. The new data and new challenges in multimedia research. CoRR, abs/1503.01817, 2015.
  35. 35.D. Tran, L. D. Bourdev, R. Fergus, L. Torresani, and M. Paluri. C3D: generic features for video analysis. CoRR, abs/1412.0767, 2014.
  36. 36.H. Wang, M. M. Ullah, A. Kläser, I. Laptev, and C. Schmid. Evaluation of local spatio-temporal features for action recognition. In Proc. BMVC, 2009.
  37. 37.S. Wiesler, A. Richard, R. Schlüter, and H. Ney. Mean-normalized stochastic gradient for large-scale deep learning. In IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2014, Florence, Italy, May 4-9, 2014, pages 180–184. IEEE, 2014.
  38. 38.J. Xiao, K. A. Ehinger, J. Hays, A. Torralba, A. Oliva, and J. Xiao. Sun database: Exploring a large collection of scene categories, 2013.
  39. 39.Z. Xu, Y. Yang, and A. G. Hauptmann. A discriminative cnn video representation for event detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2015.
  40. 40.H.-F. Yu, P. Jain, P. Kar, and I. Dhillon. Large-scale multi-label learning with missing labels. In Proceedings of The 31st International Conference on Machine Learning (ICML), pages 593–601, 2014.
  41. 41.M. D. Zeiler and R. Fergus. Visualizing and understanding convolutional networks. CoRR, abs/1311.2901, 2013.

Citation

MLA
Abu-El-Haija, S., et al. “YouTube-8M: A Large-Scale Video Classification Benchmark”. arXiv, 2016, http://arxiv.org/abs/1609.08675v1.
APA
Abu-El-Haija, S., Kothari, N., Lee, J., Natsev, P., Toderici, G., Varadarajan, B., & Vijayanarasimhan, S. (2016). YouTube-8M: A Large-Scale Video Classification Benchmark. arXiv. http://arxiv.org/abs/1609.08675v1
Chicago
Abu-El-Haija, S., N. Kothari, J. Lee, et al. 2016. “YouTube-8M: A Large-Scale Video Classification Benchmark”. arXiv. http://arxiv.org/abs/1609.08675v1.
Harvard
Abu-El-Haija, S. et al. (2016) “YouTube-8M: A Large-Scale Video Classification Benchmark”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1609.08675v1.
Vancouver
1. Abu-El-Haija S, Kothari N, Lee J, Natsev P, Toderici G, Varadarajan B, Vijayanarasimhan S (2016) YouTube-8M: A Large-Scale Video Classification Benchmark. arXiv

BibTeX

@article{abuelhaija2016youtube,
  title = {YouTube-8M: A Large-Scale Video Classification Benchmark},
  author = {Abu-El-Haija, Sami and Kothari, Nisarg and Lee, Joonseok and Natsev, Paul and Toderici, George and Varadarajan, Balakrishnan and Vijayanarasimhan, Sudheendra},
  year = {2016},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1609.08675v1},
  eprint = {1609.08675}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission