EPIC-KITCHENS VISOR Benchmark: VIdeo Segmentations and Object Relations
Ahmad DarkhalilDandan ShanBin ZhuJian MaAmlan KarRichard E. L. HigginsSanja FidlerDavid FouheyDima Damen
Presents a large-scale egocentric video benchmark featuring dense pixel-level masks and hand-object contact relations across complex cooking sequences, establishing evaluation challenges for long-term video object segmentation and interaction reasoning.
Existing computer vision models largely rely on static images or short video clips, struggling to capture how objects physically transform over extended interactions. In real-world environments like kitchens, objects undergo dynamic state changes—such as whole vegetables being peeled, sliced, and cooked into composite meals—while being actively manipulated by human hands. Past video datasets provided either short-term object segmentations without detailed action labels or broad action labels limited to coarse bounding boxes. The article addresses this gap by introducing a benchmark suite and dataset designed to capture precise pixel-level object transformations, hand-object contacts, and long-term entity tracking over time.
The main objective of the article is to establish a comprehensive pixel-level annotation pipeline and benchmark suite for first-person (egocentric) video. Specifically, the article evaluates object tracking, hand-object contact segmentation, and long-term source retrieval using untrimmed videos from the EPIC-KITCHENS collection.
To accomplish this, the authors created an interactive annotation pipeline that combined artificial intelligence segmentation tools with manual human verification over a 22-month period. Working across 36 hours of video from 179 recordings, the team generated 271,600 manual masks across 257 object classes, which were further expanded to 9.9 million dense masks through automated bidirectional interpolation. The dataset was structured into sequences averaging 12 seconds across consecutive actions, which is 2.5 to 4 times longer than previous benchmark datasets. The authors established three distinct evaluation challenges: tracking object masks across consecutive actions, segmenting hands alongside the specific objects they contact, and tracing a query object back across minutes of video to identify its source container.
The experimental findings demonstrate significant performance variations across tasks. Fine-tuning models directly on this dataset improved video object segmentation accuracy by roughly 13% over general pre-trained models, achieving a tracking score of 78.0 on the test set. However, performance dropped by about 7 to 9 points when evaluating in previously unseen kitchen environments. Hand detection and segmentation proved highly accurate, reaching over 90% accuracy, but segmenting manipulated objects and predicting contact states remained challenging, scoring between 24% and 34%. For long-term reasoning, baseline models struggled to locate when an object emerged from its source container across average time gaps of 5.4 minutes, but providing oracle temporal boundaries boosted source identification accuracy from 34% to over 94%.
These results indicate that while current computer vision architectures reliably detect hands and track rigid objects, they face severe limitations when tracking severe physical transformations, occlusions, and long-term scene interactions. Developing robust egocentric vision systems is critical for applications in robotics, assistive augmented reality, and workplace monitoring, where systems must maintain spatial awareness of tools and changing materials. The substantial drop in performance within unseen environments highlights the ongoing operational risk of deploying vision systems that fail to generalize across diverse physical settings.
Moving forward, researchers and developers should focus on multi-frame reasoning architectures that can infer contact states and locate temporal transitions rather than relying on single-frame detection. Teams should leverage the publicly available dataset and code repositories to evaluate how foundation models handle complex physical transformations over extended horizons.
The findings are constrained by several dataset limitations, including severe motion blur in first-person cameras, long-tailed distributions where common containers like refrigerators dominate source reasoning, and annotator ambiguities regarding complex boundaries like sauces and mixtures. Furthermore, benchmark evaluations for long-term reasoning relied on oracle-assisted baselines rather than fully automated end-to-end artificial intelligence systems, meaning practical deployments will face greater real-world performance degradation until temporal reasoning capabilities improve.
- Paper: Scaling Egocentric Vision: The EPIC-KITCHENS Dataset, Dima Damen et al. (2018). It introduces the core EPIC-KITCHENS egocentric video benchmark upon which VISOR directly builds its dense pixel-level segmentations and interaction annotations.
- Paper: The 2017 DAVIS Challenge on Video Object Segmentation, Jordi Pont-Tuset et al. (2017). It establishes the foundational multi-target semi-supervised video object segmentation benchmark and evaluation methodology adapted by modern video segmentation suites.
- Paper: A Benchmark Dataset and Evaluation Methodology for Video Object Segmentation, Federico Perazzi et al. (2016). It defines the standardized metrics and dense pixel-level annotation protocols for video object segmentation that inform subsequent temporal benchmark design.
- Paper: Ego4D: Around the World in 3,000 Hours of Egocentric Video, Kristen Grauman et al. (2021). It establishes large-scale egocentric video challenges focused on hands and interacting objects, providing essential context for first-person physical reasoning benchmarks.
- Paper: LVIS: A Dataset for Large Vocabulary Instance Segmentation, Agrim Gupta et al. (2019). It formalizes large-vocabulary instance segmentation and crowdsourced mask annotation methodologies essential for handling diverse everyday object categories.
- Paper: Fast Online Object Tracking and Segmentation: A Unifying Approach, Qiang Wang et al. (2018). It demonstrates unified bounding-box tracking and semi-supervised mask generation mechanisms relevant to AI-assisted video annotation pipelines.
- Paper: Mask R-CNN, Kaiming He et al. (2017). It provides the foundational deep learning framework for instance-level object masking and spatial alignment upon which video segmentation models depend.
- Paper: SAM 2: Segment Anything in Images and Videos, Nikhila Ravi et al. (2025). It extends promptable segmentation to video masklet tracking across complex motion, occlusion, and object deformation in dynamic footage.
- Paper: HOT3D: Hand and Object Tracking in 3D from Egocentric Multi-View Videos, Prithviraj Banerjee et al. (2025). It advances egocentric hand-object interaction understanding by capturing synchronized 3D tracking and multi-view poses during physical object manipulation.
- Paper: PACO: Parts and Attributes of Common Objects, Vignesh Ramanathan et al. (2023). It extends fine-grained object understanding by jointly benchmarking part-level segmentation and detailed attributes in egocentric and common object settings.
- Paper: HourVideo: 1-Hour Video-Language Understanding, Keshigeyan Chandrasegaran et al. (2024). It scales egocentric video understanding to hour-long temporal horizons, testing long-range reasoning over complex daily human activities.
- Paper: Egocentric Audio-Visual Object Localization, Chao Huang et al. (2023). It applies egocentric object localization to multimodal audio-visual signals while explicitly compensating for wearable camera egomotion.
- Paper: NeuralDome: A Neural Modeling Pipeline on Multi-View Human-Object Interactions, Juze Zhang et al. (2023). It continues human-object interaction research into neural 3D modeling and multi-view spatial reconstruction to overcome severe interaction occlusions.
