EPIC-KITCHENS VISOR Benchmark: VIdeo Segmentations and Object Relations

Ahmad DarkhalilDandan ShanBin ZhuJian MaAmlan KarRichard E. L. HigginsSanja FidlerDavid FouheyDima Damen

article2022NeurIPS174 citations

Presents a large-scale egocentric video benchmark featuring dense pixel-level masks and hand-object contact relations across complex cooking sequences, establishing evaluation challenges for long-term video object segmentation and interaction reasoning.

Listen

Existing computer vision models largely rely on static images or short video clips, struggling to capture how objects physically transform over extended interactions. In real-world environments like kitchens, objects undergo dynamic state changes—such as whole vegetables being peeled, sliced, and cooked into composite meals—while being actively manipulated by human hands. Past video datasets provided either short-term object segmentations without detailed action labels or broad action labels limited to coarse bounding boxes. The article addresses this gap by introducing a benchmark suite and dataset designed to capture precise pixel-level object transformations, hand-object contacts, and long-term entity tracking over time.

The main objective of the article is to establish a comprehensive pixel-level annotation pipeline and benchmark suite for first-person (egocentric) video. Specifically, the article evaluates object tracking, hand-object contact segmentation, and long-term source retrieval using untrimmed videos from the EPIC-KITCHENS collection.

To accomplish this, the authors created an interactive annotation pipeline that combined artificial intelligence segmentation tools with manual human verification over a 22-month period. Working across 36 hours of video from 179 recordings, the team generated 271,600 manual masks across 257 object classes, which were further expanded to 9.9 million dense masks through automated bidirectional interpolation. The dataset was structured into sequences averaging 12 seconds across consecutive actions, which is 2.5 to 4 times longer than previous benchmark datasets. The authors established three distinct evaluation challenges: tracking object masks across consecutive actions, segmenting hands alongside the specific objects they contact, and tracing a query object back across minutes of video to identify its source container.

The experimental findings demonstrate significant performance variations across tasks. Fine-tuning models directly on this dataset improved video object segmentation accuracy by roughly 13% over general pre-trained models, achieving a tracking score of 78.0 on the test set. However, performance dropped by about 7 to 9 points when evaluating in previously unseen kitchen environments. Hand detection and segmentation proved highly accurate, reaching over 90% accuracy, but segmenting manipulated objects and predicting contact states remained challenging, scoring between 24% and 34%. For long-term reasoning, baseline models struggled to locate when an object emerged from its source container across average time gaps of 5.4 minutes, but providing oracle temporal boundaries boosted source identification accuracy from 34% to over 94%.

These results indicate that while current computer vision architectures reliably detect hands and track rigid objects, they face severe limitations when tracking severe physical transformations, occlusions, and long-term scene interactions. Developing robust egocentric vision systems is critical for applications in robotics, assistive augmented reality, and workplace monitoring, where systems must maintain spatial awareness of tools and changing materials. The substantial drop in performance within unseen environments highlights the ongoing operational risk of deploying vision systems that fail to generalize across diverse physical settings.

Moving forward, researchers and developers should focus on multi-frame reasoning architectures that can infer contact states and locate temporal transitions rather than relying on single-frame detection. Teams should leverage the publicly available dataset and code repositories to evaluate how foundation models handle complex physical transformations over extended horizons.

The findings are constrained by several dataset limitations, including severe motion blur in first-person cameras, long-tailed distributions where common containers like refrigerators dominate source reasoning, and annotator ambiguities regarding complex boundaries like sauces and mixtures. Furthermore, benchmark evaluations for long-term reasoning relied on oracle-assisted baselines rather than fully automated end-to-end artificial intelligence systems, meaning practical deployments will face greater real-world performance degradation until temporal reasoning capabilities improve.

arXiv: 2209.13064
Cover for EPIC-KITCHENS VISOR Benchmark: VIdeo Segmentations and Object Relations

Abstract

We introduce VISOR, a new dataset of pixel annotations and a benchmark suite for segmenting hands and active objects in egocentric video. VISOR annotates videos from EPIC-KITCHENS, which comes with a new set of challenges not encountered in current video segmentation datasets. Specifically, we need to ensure both short- and long-term consistency of pixel-level annotations as objects undergo transformative interactions, e.g. an onion is peeled, diced and cooked - where we aim to obtain accurate pixel-level annotations of the peel, onion pieces, chopping board, knife, pan, as well as the acting hands. VISOR introduces an annotation pipeline, AI-powered in parts, for scalability and quality. In total, we publicly release 272K manual semantic masks of 257 object classes, 9.9M interpolated dense masks, 67K hand-object relations, covering 36 hours of 179 untrimmed videos. Along with the annotations, we introduce three challenges in video object segmentation, interaction understanding and long-term reasoning.

For data, code and leaderboards: http://epic-kitchens.github.io/VISOR

Table of Contents

  • 1 Introduction
  • 2 Related Efforts
  • 3 VISOR annotation pipeline
  • 3.1 Entities, frames and sub-sequences
  • 3.2 Tooling, annotation rules and annotator training
  • 3.3 VISOR Object Relations and Entities
  • 3.4 Dense annotations
  • 4 Dataset Statistics and Analysis
  • 5 Challenges and Baselines
  • 5.1 Semi-Supervised Video Object Segmentation (VOS)
  • 5.2 Hand-Object Segmentation (HOS) Relations
  • 5.3 Where Did This Come From (WDTCF)? A Taster Challenge
  • 6 Conclusion and Next Steps
  • References
  • A Appendix - Societal Impact and Resources Used
  • B Appendix - Entities, Frames, and Subsequences (Main §2.1)
  • B.1 Entity Preparation
  • B.2 Frame Extraction
  • B.3 Subsequence Examples
  • C Appendix - Annotation rules and annotator training (Main §2.2) / Annotator Training and Rules
  • C.1 Annotator Recruiting
  • C.2 Training Material
  • C.3 Rules
  • D Appendix – Tooling: The TORonto Annotation Suite (Main §2.2)
  • D.1 Segmentation Tools
  • D.2 Project Management
  • D.3 Annotator Agreement
  • D.4 Annotation Efficiency
  • E Appendix – Correction (Main §2.2) / Correction
  • E.1 Correction Interface for Collecting Comments
  • E.2 Correction on TORAS
  • E.3 Statistics of Collected Comments
  • F Appendix - VISOR Object Relations and Entities (Main §2.3)
  • F.1 Quality Control
  • F.2 Annotation Instructions for Hand-Object Segment Relations
  • F.3 Annotation Instructions for Exhaustive Annotation
  • G Appendix - Dense Annotations (Main §2.4)
  • H Appendix - VOS Benchmark Details (Main §4.1)
  • H.1 Data Preparation
  • H.2 Metrics and Evaluation
  • H.3 Baseline Training Details
  • H.4 Additional Results
  • I Appendix - HOS Benchmark Details (Main §4.2)
  • I.1 Data Preparation
  • I.2 Metrics and Evaluation
  • I.3 Baselines and Training Details
  • I.4 Additional Results
  • J Appendix - WDTCF Benchmark Details (Main §4.3)
  • J.1 Data Preparation and Annotation
  • J.2 Evaluation Details
  • J.3 Baseline Details
  • J.4 Additional Results
  • K Appendix - EPIC-KITCHENS VISOR - Datasheet for Dataset

Citation

MLA
Darkhalil, A., et al. “EPIC-KITCHENS VISOR Benchmark: VIdeo Segmentations and Object Relations”. arXiv, 2022, http://arxiv.org/abs/2209.13064v1.
APA
Darkhalil, A., Shan, D., Zhu, B., Ma, J., Kar, A., Higgins, R., Fidler, S., Fouhey, D., & Damen, D. (2022). EPIC-KITCHENS VISOR Benchmark: VIdeo Segmentations and Object Relations. arXiv. http://arxiv.org/abs/2209.13064v1
Chicago
Darkhalil, A., D. Shan, B. Zhu, et al. 2022. “EPIC-KITCHENS VISOR Benchmark: VIdeo Segmentations and Object Relations”. arXiv. http://arxiv.org/abs/2209.13064v1.
Harvard
Darkhalil, A. et al. (2022) “EPIC-KITCHENS VISOR Benchmark: VIdeo Segmentations and Object Relations”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2209.13064v1.
Vancouver
1. Darkhalil A, Shan D, Zhu B, Ma J, Kar A, Higgins R, Fidler S, Fouhey D, Damen D (2022) EPIC-KITCHENS VISOR Benchmark: VIdeo Segmentations and Object Relations. arXiv

BibTeX

@article{darkhalil2022epic,
  title = {EPIC-KITCHENS VISOR Benchmark: VIdeo Segmentations and Object Relations},
  author = {Darkhalil, Ahmad and Shan, Dandan and Zhu, Bin and Ma, Jian and Kar, Amlan and Higgins, Richard and Fidler, Sanja and Fouhey, David and Damen, Dima},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2209.13064v1},
  eprint = {2209.13064}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors