ActivityNet: A large-scale video benchmark for human activity understanding

Fabian Caba HeilbronVictor EscorciaBernard GhanemJuan Carlos Niebles

article2015CVPR3,040 citations

Introduces ActivityNet, a large-scale video benchmark spanning 203 rich daily-living activity categories across 849 hours of untrimmed video to evaluate trimmed recognition, untrimmed classification, and temporal activity detection.

Listen

Current computer vision systems for recognizing human actions remain limited because existing benchmarks emphasize only simple movements in short, manually trimmed clips drawn mostly from sports or a narrow set of categories. This gap matters because people spend far more time on everyday tasks such as household work and personal care than on the activities that dominate current datasets.

The document introduces ActivityNet to address that shortfall. Its goal is to supply a large, diverse collection of untrimmed videos that span a wide range of complex daily activities and to demonstrate how the new resource can support three practical tasks: classifying entire videos, classifying trimmed activity segments, and detecting activity instances inside long videos.

The authors built the benchmark through crowdsourced collection and annotation of online videos, resulting in 203 activity classes, roughly 137 videos per class, and 849 total hours of video. They also compared ActivityNet against seven prior datasets by mapping every class in those collections onto a common activity hierarchy and by plotting scale in terms of both number of classes and samples per class.

ActivityNet covers activity types that are almost absent from earlier benchmarks, such as household activities, personal care, and work-related tasks, while existing collections remain heavily skewed toward sports and simple actions. In overall size it ranks second among published resources yet leads in breadth of activity coverage. The taxonomy organizes classes into three levels, allowing algorithms to be evaluated at varying degrees of granularity.

These characteristics mean researchers can now train and test models on realistic, untrimmed footage that better matches the distribution of daily human behavior. Improved performance on such data should translate into more reliable systems for applications such as video search, surveillance, and assistive technologies.

The authors are releasing the full set of annotations and an evaluation toolkit so that the community can adopt the benchmark immediately. Further expansion of the dataset and development of baseline algorithms for the three evaluation scenarios are the logical next steps.

The summary rests on an extended abstract; the complete paper may contain additional experimental results or updated statistics not shown here.

Cover for ActivityNet: A large-scale video benchmark for human activity understanding

Abstract

In spite of many dataset efforts for human action recognition, current computer vision algorithms are still severely limited in terms of the variability and complexity of the actions that they can recognize. This is in part due to the simplicity of current benchmarks, which mostly focus on simple actions and movements occurring on manually trimmed videos. In this paper we introduce ActivityNet, a new large-scale video benchmark for human activity understanding. Our benchmark aims at covering a wide range of complex human activities that are of interest to people in their daily living. In its current version, ActivityNet provides samples from 203 activity classes with an average of 137 untrimmed videos per class and 1.41 activity instances per video, for a total of 849 video hours. We illustrate three scenarios in which ActivityNet can be used to compare algorithms for human activity understanding: untrimmed video classification, trimmed activity classification and activity detection.

Knowls

  1. Knowl 1 — ActivityNet Dataset Scale and Properties

    definition

    ActivityNet is a large-scale video benchmark designed for human activity understanding across a broad spectrum of complex, daily living activities. In its release configuration, the dataset contains:

    • Activity categories: 203 distinct activity classes organized hierarchically.
    • Video volume and duration: A total of 849 hours of untrimmed video footage.
    • Dataset scale: An average of 137 untrimmed videos per activity class.
    • Annotation density: An average of 1.41 temporal activity instances annotated per video recording.

    Videos in ActivityNet are untrimmed recordings sourced online, where each activity instance is annotated with its precise start and end temporal boundaries.

  2. Knowl 2 — Hierarchical Semantic Taxonomy for Human Activities

    model/method

    ActivityNet organizes human activities into a multi-tier semantic tree taxonomy spanning high-level activity domains down to specific leaf activities. The taxonomy comprises eight top-level parent categories:

    1. Personal care
    2. Eating & drinking
    3. Household activities
    4. Caring & helping
    5. Work-related (including education and working activities)
    6. Sports & exercises
    7. Socializing & leisure
    8. Simple actions

    Each leaf activity node connects to the root via intermediate semantic tiers. For example:

    Household activities⟶Housework⟶Interior cleaning⟶Cleaning windows\text{Household activities} \longrightarrow \text{Housework} \longrightarrow \text{Interior cleaning} \longrightarrow \text{Cleaning windows}

    Personal care⟶Grooming⟶Grooming oneself⟶Brushing teeth\text{Personal care} \longrightarrow \text{Grooming} \longrightarrow \text{Grooming oneself} \longrightarrow \text{Brushing teeth}

    This structure balances everyday domestic activities (which constitute the majority of human daily routines) with athletic, social, and occupational actions.

  3. Knowl 3 — ActivityNet Evaluation Scenarios for Video Understanding

    experimental setup

    ActivityNet defines three distinct benchmarking scenarios to evaluate and compare computer vision algorithms for video activity understanding:

    1. Untrimmed Video Classification: Given a full, unsegmented video containing one or more activity instances along with background or non-action content, the goal is to predict the set of activity categories present in the video without requiring temporal boundary localization.
    2. Trimmed Activity Classification: Given a pre-trimmed video segment containing only the temporal duration of a specific activity execution, the goal is to predict the correct activity category label.
    3. Activity Detection (Temporal Localization): Given an untrimmed video, the goal is to simultaneously classify the activities present and detect their temporal extent by predicting start time and end time intervals (tstart,tend)(t_{\text{start}}, t_{\text{end}}) for each detected activity instance.
  4. Knowl 4 — Comparative Diversity and Scale of Video Action Benchmarks

    empirical result

    When existing human action benchmarks (including UCF101, HMDB51, Hollywood, THUMOS'14, Sports-1M, Olympic, and MPII Human Pose/Activities) are mapped into ActivityNet's 8 top-level taxonomy categories, existing datasets show heavy category imbalance and limited diversity:

    • Category composition: Prior datasets are predominantly skewed toward Sports & exercises (e.g., UCF101, Olympic, Sports-1M), Socializing & leisure / Simple actions (e.g., Hollywood, HMDB51), or specialized fine-grained domains (e.g., cooking in MPII), largely omitting Household activities, Personal care, and Work-related categories.
    • Scale vs. Diversity tradeoff: While datasets like Sports-1M offer high volume (over 10310^3 classes and 10310^3 samples per class) strictly in the sports domain, and standard action recognition benchmarks typically offer around 100 classes with 100 samples per class, ActivityNet ranks as the second largest overall dataset by instance count while providing the highest diversity across varied daily human activity domains.

Coverage note — None was omitted; all contributed aspects from this 1-page extended abstract (dataset statistics, taxonomy design, benchmark tasks, and comparative dataset analysis) are fully covered.

References

  1. 1.Andrej Karpathy, George Toderici, Sanketh Shetty, Thomas Leung, Rahul Sukthankar, and Li Fei-Fei. Large-scale video classification with convolutional neural networks. In CVPR, 2014.
  2. 2.Amir Roshan Zamir Khurram Soomro and Mubarak Shah. A dataset of 101 human action classes from videos in the wild. Technical report, University of Central Florida, 2012.
  3. 3.H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre. HMDB: A large video database for human motion recognition. In ICCV, 2011. doi: 10.1109/ICCV.2011.6126543.
  4. 4.Ivan Laptev, Marcin Marszałek, Cordelia Schmid, and Benjamin Rozenfeld. Learning realistic human actions from movies. In CVPR, 2008.
  5. 5.Juan Carlos Niebles, Chih-Wei Chen, and Li Fei-Fei. Modeling temporal structure of decomposable motion segments for activity classification. In ECCV, 2010.
  6. 6.Marcus Rohrbach, Sikandar Amin, Mykhaylo Andriluka, and Bernt Schiele. A database for fine grained activity detection of cooking activities. In CVPR, 2012.
  7. 7.Thumos14. Thumos challenge 2014. http://crcv.ucf.edu/THUMOS14, 2013.
  8. 8.U.S. Department of Labor. American time use survey. http://www.bls.gov/tus/, 2013.

Citation

MLA
Heilbron, F. C., et al. “ActivityNet: A Large-scale Video Benchmark for Human Activity Understanding”. 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2015, pp. 961–70, https://doi.org/10.1109/CVPR.2015.7298698.
APA
Heilbron, F. C., Escorcia, V., Ghanem, B., & Niebles, J. C. (2015). ActivityNet: A large-scale video benchmark for human activity understanding. 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 961–970. https://doi.org/10.1109/CVPR.2015.7298698
Chicago
Heilbron, F. C., V. Escorcia, B. Ghanem, and J. C. Niebles. 2015. “ActivityNet: A Large-scale Video Benchmark for Human Activity Understanding”. 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 961–70. https://doi.org/10.1109/CVPR.2015.7298698.
Harvard
Heilbron, F.C. et al. (2015) “ActivityNet: A large-scale video benchmark for human activity understanding”, 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 961–970. Available at: https://doi.org/10.1109/CVPR.2015.7298698.
Vancouver
1. Heilbron FC, Escorcia V, Ghanem B, Niebles JC (2015) ActivityNet: A large-scale video benchmark for human activity understanding. In: 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 961–970

BibTeX

@inproceedings{Heilbron_2015, title={ActivityNet: A large-scale video benchmark for human activity understanding}, url={http://dx.doi.org/10.1109/CVPR.2015.7298698}, DOI={10.1109/cvpr.2015.7298698}, booktitle={2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Heilbron, Fabian Caba and Escorcia, Victor and Ghanem, Bernard and Niebles, Juan Carlos}, year={2015}, month=June, pages={961–970} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE