Scaling Egocentric Vision: The EPIC-KITCHENS Dataset
Dima DamenHazel DoughtyGiovanni Maria FarinellaSanja FidlerAntonino FurnariEvangelos KazakosDavide MoltisantiJonathan MunroToby PerrettWill Price
Introduces EPIC-KITCHENS, a large-scale first-person video benchmark of 55 hours of unscripted daily activities with participant-narrated action and object annotations to advance egocentric action recognition, detection, and anticipation across novel environments.
First-person (egocentric) computer vision provides vital insights into human behavior, intentions, and interactions with physical objects, with strong applications in assistive devices, smart home automation, and robotics. However, progress in this domain has been heavily constrained by the lack of large-scale, unscripted video datasets recorded in diverse, real-world settings. Existing benchmarks are mostly small, recorded from third-person perspectives, or rely on artificial scripts and staged laboratory environments that fail to capture natural human multi-tasking and daily variations.
The article introduces and evaluates EPIC-KITCHENS, a large-scale first-person video benchmark designed to advance egocentric vision. The primary objective is to capture naturalistic human-object interactions in home environments and benchmark baseline artificial intelligence models on object detection, action recognition, and action anticipation.
The benchmark dataset was collected by 32 participants representing 10 nationalities across 4 cities in North America and Europe over three consecutive days. Using head-mounted cameras, participants recorded 55 hours of unscripted kitchen activities across 432 sequences, totaling 11.5 million frames. The annotation process used participant audio narrations recorded directly after the activities to capture true user intent, followed by crowdsourced refinement to establish 39,564 precise action segments and 454,255 bounding boxes across 125 verb classes and 331 noun classes. The researchers then evaluated baseline deep learning models across two test settings: environments seen during training (S1) and completely unseen environments (S2).
The baseline evaluations yielded several critical findings. First, state-of-the-art object detection achieves low accuracy on fine-grained and low-frequency objects; detection mean average precision at an intersection-over-union of 0.5 was 35.4% in seen environments and 33.1% in unseen environments. Second, recognizing combined actions (both verb and noun correctly) is very difficult, yielding only 20.5% top-1 accuracy in seen kitchens and dropping by nearly half to 10.9% in unseen kitchens. Third, action anticipation one second before execution presents a substantial hurdle, achieving only 4.6% top-1 accuracy in seen environments and 1.7% in unseen environments. Finally, while object detection models generalized relatively well across seen and unseen environments, action recognition models suffered severe performance degradation when applied to unseen environments.
These findings demonstrate that current computer vision algorithms are far from achieving reliable performance in naturalistic, unconstrained environments. Developing systems for smart wearables and assistive living requires AI models that can generalize across different domestic layouts, handle rare objects, and understand longer-term user context rather than relying on narrow, sequential assumptions.
Moving forward, machine learning researchers and practitioners should use the public benchmark and leaderboards to develop architectures capable of long-term temporal modeling, multi-scale action history tracking, and few-shot learning for rare objects. Real-time inference efficiency should also be prioritized to make these models viable for deployment on wearable devices.
The primary limitations include reliance on single-person activities in kitchen environments, incomplete participant narrations for secondary actions (such as closing doors or drawers), and long-tail class imbalances inherent to natural human behavior. Nonetheless, the high consistency and low error rates across quality checks confirm that the dataset provides a robust and credible benchmark for next-generation egocentric vision research.
- Paper: Hollywood in Homes: Crowdsourcing Data Collection for Activity Understanding, Gunnar A. Sigurdsson et al. (2016). It introduces the Charades benchmark for unscripted daily indoor activities, laying essential methodological groundwork for crowdsourced video collection of complex home tasks.
- Paper: ActivityNet: A large-scale video benchmark for human activity understanding, Fabian Caba Heilbron et al. (2015). It provides the foundational framework and evaluation paradigms for large-scale human activity understanding in untrimmed, daily-life videos.
- Paper: Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset, João Carreira et al. (2017). It establishes standard deep 3D spatio-temporal architectures and pre-training methodologies essential for baseline video action recognition.
- Paper: Dense-Captioning Events in Videos, Ranjay Krishna et al. (2017). It introduces the paradigm of temporally dense event detection and narration in long-form videos that directly informs narrative-based video annotations.
- Paper: UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild, Khurram Soomro et al. (2012). It serves as a foundational benchmark for action recognition that set initial standards for action categorization from video sequences.
- Paper: Ego4D: Around the World in 3,000 Hours of Egocentric Video, Kristen Grauman et al. (2021). It scales first-person visual perception from kitchen environments to massive, multi-institutional, worldwide daily-life scenarios across thousands of hours.
- Paper: Video Swin Transformer, Ze Liu et al. (2021). It applies advanced vision transformer architectures to solve fine-grained temporal object state change and interaction tasks pioneered by egocentric benchmarks.
- Paper: HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips, Antoine Miech et al. (2019). It extends joint video-language representation learning by leveraging massive collections of instructional and narrated videos without manual labels.
- Paper: VideoBERT: A Joint Model for Video and Language Representation Learning, Chen Sun et al. (2019). It develops self-supervised multi-modal representation models specifically applied to instructional cooking videos and action-object associations.
