ActivityNet: A large-scale video benchmark for human activity understanding
Fabian Caba HeilbronVictor EscorciaBernard GhanemJuan Carlos Niebles
Introduces ActivityNet, a large-scale video benchmark spanning 203 rich daily-living activity categories across 849 hours of untrimmed video to evaluate trimmed recognition, untrimmed classification, and temporal activity detection.
Current computer vision systems for recognizing human actions remain limited because existing benchmarks emphasize only simple movements in short, manually trimmed clips drawn mostly from sports or a narrow set of categories. This gap matters because people spend far more time on everyday tasks such as household work and personal care than on the activities that dominate current datasets.
The document introduces ActivityNet to address that shortfall. Its goal is to supply a large, diverse collection of untrimmed videos that span a wide range of complex daily activities and to demonstrate how the new resource can support three practical tasks: classifying entire videos, classifying trimmed activity segments, and detecting activity instances inside long videos.
The authors built the benchmark through crowdsourced collection and annotation of online videos, resulting in 203 activity classes, roughly 137 videos per class, and 849 total hours of video. They also compared ActivityNet against seven prior datasets by mapping every class in those collections onto a common activity hierarchy and by plotting scale in terms of both number of classes and samples per class.
ActivityNet covers activity types that are almost absent from earlier benchmarks, such as household activities, personal care, and work-related tasks, while existing collections remain heavily skewed toward sports and simple actions. In overall size it ranks second among published resources yet leads in breadth of activity coverage. The taxonomy organizes classes into three levels, allowing algorithms to be evaluated at varying degrees of granularity.
These characteristics mean researchers can now train and test models on realistic, untrimmed footage that better matches the distribution of daily human behavior. Improved performance on such data should translate into more reliable systems for applications such as video search, surveillance, and assistive technologies.
The authors are releasing the full set of annotations and an evaluation toolkit so that the community can adopt the benchmark immediately. Further expansion of the dataset and development of baseline algorithms for the three evaluation scenarios are the logical next steps.
The summary rests on an extended abstract; the complete paper may contain additional experimental results or updated statistics not shown here.
- Paper: UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild, Khurram Soomro et al. (2012). UCF101 exemplifies the short, action-focused benchmark tradition that ActivityNet explicitly expands beyond in category breadth and video realism.
- Paper: Action recognition by dense trajectories, Heng Wang et al. (2011). Dense trajectories provide a key pre-deep-learning action-recognition approach that helps frame the methods ActivityNet evaluates on its broader benchmark.
- Paper: Two-Stream Convolutional Networks for Action Recognition in Videos, Karen Simonyan et al. (2014). The two-stream model establishes a major appearance-and-motion approach that helps orient readers to action-recognition methods evaluated on ActivityNet.
- Paper: Dense-Captioning Events in Videos, Ranjay Krishna et al. (2017). ActivityNet Captions builds on ActivityNet’s untrimmed-video foundation to localize events and describe them with temporally grounded language.
- Paper: Proposal-Based Multiple Instance Learning for Weakly-Supervised Temporal Action Localization, Huan Ren et al. (2023). This work applies ActivityNet’s long-video benchmark to weakly supervised temporal localization, extending its detection setting with proposal-based learning.
