Scaling Egocentric Vision: The EPIC-KITCHENS Dataset

Dima DamenHazel DoughtyGiovanni Maria FarinellaSanja FidlerAntonino FurnariEvangelos KazakosDavide MoltisantiJonathan MunroToby PerrettWill Price

article2018arXiv1,408 citations

Introduces EPIC-KITCHENS, a large-scale first-person video benchmark of 55 hours of unscripted daily activities with participant-narrated action and object annotations to advance egocentric action recognition, detection, and anticipation across novel environments.

Listen

First-person (egocentric) computer vision provides vital insights into human behavior, intentions, and interactions with physical objects, with strong applications in assistive devices, smart home automation, and robotics. However, progress in this domain has been heavily constrained by the lack of large-scale, unscripted video datasets recorded in diverse, real-world settings. Existing benchmarks are mostly small, recorded from third-person perspectives, or rely on artificial scripts and staged laboratory environments that fail to capture natural human multi-tasking and daily variations.

The article introduces and evaluates EPIC-KITCHENS, a large-scale first-person video benchmark designed to advance egocentric vision. The primary objective is to capture naturalistic human-object interactions in home environments and benchmark baseline artificial intelligence models on object detection, action recognition, and action anticipation.

The benchmark dataset was collected by 32 participants representing 10 nationalities across 4 cities in North America and Europe over three consecutive days. Using head-mounted cameras, participants recorded 55 hours of unscripted kitchen activities across 432 sequences, totaling 11.5 million frames. The annotation process used participant audio narrations recorded directly after the activities to capture true user intent, followed by crowdsourced refinement to establish 39,564 precise action segments and 454,255 bounding boxes across 125 verb classes and 331 noun classes. The researchers then evaluated baseline deep learning models across two test settings: environments seen during training (S1) and completely unseen environments (S2).

The baseline evaluations yielded several critical findings. First, state-of-the-art object detection achieves low accuracy on fine-grained and low-frequency objects; detection mean average precision at an intersection-over-union of 0.5 was 35.4% in seen environments and 33.1% in unseen environments. Second, recognizing combined actions (both verb and noun correctly) is very difficult, yielding only 20.5% top-1 accuracy in seen kitchens and dropping by nearly half to 10.9% in unseen kitchens. Third, action anticipation one second before execution presents a substantial hurdle, achieving only 4.6% top-1 accuracy in seen environments and 1.7% in unseen environments. Finally, while object detection models generalized relatively well across seen and unseen environments, action recognition models suffered severe performance degradation when applied to unseen environments.

These findings demonstrate that current computer vision algorithms are far from achieving reliable performance in naturalistic, unconstrained environments. Developing systems for smart wearables and assistive living requires AI models that can generalize across different domestic layouts, handle rare objects, and understand longer-term user context rather than relying on narrow, sequential assumptions.

Moving forward, machine learning researchers and practitioners should use the public benchmark and leaderboards to develop architectures capable of long-term temporal modeling, multi-scale action history tracking, and few-shot learning for rare objects. Real-time inference efficiency should also be prioritized to make these models viable for deployment on wearable devices.

The primary limitations include reliance on single-person activities in kitchen environments, incomplete participant narrations for secondary actions (such as closing doors or drawers), and long-tail class imbalances inherent to natural human behavior. Nonetheless, the high consistency and low error rates across quality checks confirm that the dataset provides a robust and credible benchmark for next-generation egocentric vision research.

arXiv: 1804.02748
Cover for Scaling Egocentric Vision: The EPIC-KITCHENS Dataset

Abstract

First-person vision is gaining interest as it offers a unique viewpoint on people's interaction with objects, their attention, and even intention. However, progress in this challenging domain has been relatively slow due to the lack of sufficiently large datasets. In this paper, we introduce EPIC-KITCHENS, a large-scale egocentric video benchmark recorded by 32 participants in their native kitchen environments. Our videos depict nonscripted daily activities: we simply asked each participant to start recording every time they entered their kitchen. Recording took place in 4 cities (in North America and Europe) by participants belonging to 10 different nationalities, resulting in highly diverse cooking styles. Our dataset features 55 hours of video consisting of 11.5M frames, which we densely labeled for a total of 39.6K action segments and 454.3K object bounding boxes. Our annotation is unique in that we had the participants narrate their own videos (after recording), thus reflecting true intention, and we crowd-sourced ground-truths based on these. We describe our object, action and anticipation challenges, and evaluate several baselines over two test splits, seen and unseen kitchens. Dataset and Project page: this http URL

Table of Contents

  • 1 Introduction
  • 2 Related Datasets
  • 3 The EPIC-KITCHENS Dataset
  • 3.1 Data Collection
  • 3.2 Action Segment Annotations
  • 3.3 Active Object Bounding Box Annotations
  • 3.4 Verb and Noun Classes
  • 3.5 Annotation Quality Assurance
  • 4 Benchmarks and Baseline Results
  • 4.1 Object Detection Benchmark
  • 4.2 Action Recognition Benchmark
  • 4.3 Action Anticipation Benchmark
  • Discussion:
  • 5 Conclusion and Future Work
  • References

Knowls

  1. Knowl 1 — EPIC-KITCHENS Dataset Specification and Collection

    definition

    EPIC-KITCHENS is a large-scale egocentric (first-person) video dataset capturing non-scripted, natural daily activities recorded in native home kitchen environments by 32 participants across 4 cities (Bristol, UK; Toronto, Canada; Catania, Italy; Seattle, USA) representing 10 nationalities.

    Key dataset characteristics:

    • Volume and Duration: 55 hours of recording spanning 11.5 million frames across 432 continuous video sequences. On average, each participant recorded 13.6 sequences totaling 1.7 hours (with a maximum single recording of 4.6 hours).
    • Capture Hardware and Settings: Recorded with head-mounted GoPro cameras with viewpoint calibration ensuring outstretched hands appear roughly centrally. Videos were recorded with synchronized audio at Full HD resolution (1920×10801920 \times 1080), linear field of view, and 59.94 fps59.94\text{ fps} (with minor portions at 1280×7201280 \times 720 [1%] or 1920×14401920 \times 1440 [0.5%], and 30 fps30\text{ fps} [1%], 48 fps48\text{ fps} [1%], or 90 fps90\text{ fps} [0.2%]).
    • Activity Profile: Unscripted solo cooking, food preparation, cleaning, dishwashing, and routine kitchen interactions featuring natural multi-tasking and parallel-goal activities.
  2. Knowl 2 — Participant-Led Audio Narration Pipeline

    model/method

    To annotate long, unscripted egocentric video streams without interrupting natural activity during capture, EPIC-KITCHENS uses a participant-led post-recording narration workflow:

    1. Audio Commentary: After recording, each participant watches their own videos and narrates the actions performed in real time using a handheld audio recorder in their native/fluent language (17 in English, 7 in Italian, 6 in Spanish, 1 in Greek, and 1 in Chinese). Instructions direct them to use concise present-tense verb-object phrases (e.g., "cut kiwi", "pour water into kettle").
    2. Audio Chunking and Crowdsourced Transcription: Silence below a set decibel threshold is stripped to split audio into approximately 30-second speech chunks. Amazon Mechanical Turk (AMT) workers transcribe (and translate non-English narrations into English) with 3-worker redundancy and an edit distance agreement filter of 0 to at least one other transcription.
    3. Temporal Timestamp Alignment: Accurate start and end timestamps for each audio narration segment are extracted using automated closed-caption alignment tools.

    This pipeline yielded 39,596 action narrations (averaging one narration every 4.9s4.9\text{s} with a mean phrase length of 2.8 words).

  3. Knowl 3 — Action Segment Temporal Boundary Consensus Formulation

    equation

    To refine coarse narration timestamps into precise action segment boundaries Ai=[tsi,tei]A_i = [t_{s_i}, t_{e_i}] for each narrated phrase pip_i, Ka=4K_a = 4 crowdsourced annotators label start and end times on Amazon Mechanical Turk. Constraints require action length ≥0.5s\ge 0.5\text{s} and that an action cannot start before the preceding action's start time.

    Let Ai(j)A_i(j) denote the temporal boundary interval annotated by worker j∈{1,…,Ka}j \in \{1, \dots, K_a\} for action ii. The agreement score αi(j)\alpha_i(j) for annotator jj is calculated as the mean temporal Intersection over Union (IoU) across all other workers: αi(j)=1Ka∑k=1KaIoU(Ai(j),Ai(k))\alpha_i(j) = \frac{1}{K_a} \sum_{k=1}^{K_a} \text{IoU}\left(A_i(j), A_i(k)\right)

    The annotator with maximum overall agreement is selected as j^=arg⁡max⁡jαi(j)\hat{j} = \arg\max_j \alpha_i(j), and the annotator exhibiting maximum IoU with j^\hat{j} is identified as k^=arg⁡max⁡k≠j^IoU(Ai(j^),Ai(k))\hat{k} = \arg\max_{k \neq \hat{j}} \text{IoU}\left(A_i(\hat{j}), A_i(k)\right). The final ground-truth action segment AiA_i is formed by combining the top two agreeing annotations if their overlap is strong: Ai={Union(Ai(j^),Ai(k^)),if IoU(Ai(j^),Ai(k^))>0.5Ai(j^),otherwiseA_i = \begin{cases} \text{Union}\left(A_i(\hat{j}), A_i(\hat{k})\right), & \text{if } \text{IoU}\left(A_i(\hat{j}), A_i(\hat{k})\right) > 0.5 \\ A_i(\hat{j}), & \text{otherwise} \end{cases}

    This process produced 39,564 verified action segments with mean duration μ=3.7s\mu = 3.7\text{s} (σ=5.6s\sigma = 5.6\text{s}), covering 99.9%99.9\% of narrated phrases.

  4. Knowl 4 — Active Object Bounding Box Annotation Consensus Formulation

    equation

    EPIC-KITCHENS annotates bounding boxes exclusively for "active" objects—nouns OiO_i involved in the action phrase pip_i of segment Ai=[tsi,tei]A_i = [t_{s_i}, t_{e_i}]. Frames ff are sampled at 2 fps2\text{ fps} across the interval [tsi−2s,tei+2s][t_{s_i} - 2\text{s}, t_{e_i} + 2\text{s}] for tasks of up to 50 consecutive frames (25s25\text{s}) annotated by single workers.

    For each task, Ko=3K_o = 3 crowdsourced workers provide bounding boxes. Let BB(q,f,k)\text{BB}(q, f, k) represent the kk-th bounding box annotated by worker qq on frame ff. The agreement score β(q)\beta(q) for worker qq relative to other workers j≠qj \neq q is defined as: β(q)=∑f∑j≠qmax⁡k,lIoU(BB(j,f,k),BB(q,f,l))\beta(q) = \sum_{f} \sum_{j \neq q} \max_{k, l} \text{IoU}\left(\text{BB}(j, f, k), \text{BB}(q, f, l)\right)

    The bounding boxes provided by the worker maximizing β(q)\beta(q) are chosen as the final ground-truth annotations (resolving ties by selecting the tighter bounding boxes). Annotator quality is vetted by requiring IoU≥0.7\text{IoU} \ge 0.7 on golden reference frames. In total, 454,255 active object bounding boxes were collected (mean μ=1.64\mu = 1.64 boxes/frame, σ=0.92\sigma = 0.92).

  5. Knowl 5 — Verb and Noun Semantic Class Taxonomy

    definition

    To convert free-text narrations into discrete classes suitable for classification and detection, words extracted from the narrations are grouped into minimally overlapping semantic clusters:

    • Grammatical Extraction: Part-of-Speech (POS) tagging using SpaCy extracts the first verb and all subsequent nouns per phrase. When an active noun is replaced by a pronoun (e.g., "put it down"), the noun is propagated from the directly preceding narration phrase.
    • Taxonomy Clustering: Verbs are clustered manually, while nouns are clustered semi-automatically (handling compound nouns, merging synonymous objects like "cup" and "mug", and distinguishing distinct appliances like "washing machine" vs. "coffee machine").
    • Class Totals: The dataset defines ∣CV∣=125|C_V| = 125 verb classes and ∣CN∣=331|C_N| = 331 noun classes. Each action class is defined as a unique verb-noun pair (cv,cn)∈CV×CN(c_v, c_n) \in C_V \times C_N.
    • Super-Categories: Noun classes are categorized into 19 super-categories (9 food/drink categories and 10 non-food categories spanning kitchenware, appliances, cutlery, and furniture).
  6. Knowl 6 — Seen (S1) and Unseen (S2) Benchmark Split Protocol

    data/table

    To assess generalization to both known environments and completely novel environments, EPIC-KITCHENS partitions data into two benchmark testing splits, holding out 27%27\% of annotations for challenge leaderboards:

    • Seen Kitchens (S1): Sequences from 28 participants/kitchens are divided such that roughly 80%80\% of sequences per participant are assigned to Train/Val and 20%20\% to testing. Individual sequences are kept intact.
    • Unseen Kitchens (S2): Complete sequences from 4 participants/kitchens are held out exclusively for testing (7%7\% of total dataset frames).
    Split #Subjects #Sequences Duration (s) % Data Action Segments Bounding Boxes
    Train/Val 28 272 141,731 73% 28,561 326,388
    S1 Test (Seen) 28 106 39,084 20% 8,064 97,872
    S2 Test (Unseen) 4 54 13,231 7% 2,939 29,995
  7. Knowl 7 — Egocentric Object Detection Benchmark and Faster R-CNN Baselines

    data/table

    The object detection benchmark evaluates detection of active noun classes CNC_N in frames containing active object annotations (Icn∈CNI_{c_n \in C_N}). The baseline detector is Faster R-CNN with a ResNet-101 backbone pre-trained on MS-COCO, evaluated across many-shot classes (≥100\ge 100 training bounding boxes; 202 classes) and few-shot classes (10≤count<10010 \le \text{count} < 100; 88 classes).

    Mean Average Precision (mAP) under PASCAL VOC criteria across IoU thresholds ∈{0.05,0.50,0.75}\in \{0.05, 0.50, 0.75\}:

    Split IoU Threshold Few-shot mAP (%) Many-shot mAP (%) All mAP (%)
    S1 (Seen) IoU >0.05> 0.05 31.59 51.60 47.84
    IoU >0.50> 0.50 20.72 38.81 35.41
    IoU >0.75> 0.75 2.70 10.07 8.69
    S2 (Unseen) IoU >0.05> 0.05 23.19 49.30 46.64
    IoU >0.50> 0.50 16.95 34.95 33.11
    IoU >0.75> 0.75 2.46 8.68 8.05

    Object detection performance generalizes well from seen to unseen environments (35.41%35.41\% vs 33.11%33.11\% mAP @ IoU >0.50> 0.50), but few-shot classes exhibit significantly lower precision across all thresholds.

  8. Knowl 8 — Action Recognition Benchmark and TSN Baseline Performance

    data/table

    In the action recognition benchmark, given a pre-trimmed action segment Ai=[tsi,tei]A_i = [t_{s_i}, t_{e_i}], models predict the verb cv∈CVc_v \in C_V, noun cn∈CNc_n \in C_N, and full action pair (cv,cn)(c_v, c_n). Baselines evaluate Temporal Segment Networks (TSN) with BN-Inception (spatial RGB, TV-L1 optical flow, and equal-weight fusion) trained with joint independent loss heads for verbs and nouns, alongside Two-Stream CNN (2SCNN), random chance, and largest-class baselines.

    Evaluation results on Seen (S1) and Unseen (S2) test sets:

    Split Model Top-1 Accuracy (%) Top-5 Accuracy (%)
    Verb Noun Action Verb Noun Action
    S1 (Seen) Chance/Random 12.62 1.73 0.22 43.39 8.12 3.68
    Largest Class 22.41 4.50 1.59 70.20 18.89 14.90
    2SCNN (Fusion) 42.16 29.14 13.23 80.58 53.70 30.36
    TSN (RGB) 45.68 36.80 19.86 85.56 64.19 41.89
    TSN (Flow) 42.75 17.40 9.02 79.52 39.43 21.92
    TSN (Fusion) 48.23 36.71 20.54 84.09 62.32 39.79
    S2 (Unseen) Chance/Random 10.71 1.89 0.22 38.98 9.31 3.81
    Largest Class 22.26 4.80 0.10 63.76 19.44 17.17
    2SCNN (Fusion) 36.16 18.03 7.31 71.97 38.41 19.49
    TSN (RGB) 34.89 21.82 10.11 74.56 45.34 25.33
    TSN (Flow) 40.08 14.51 6.73 73.40 33.77 18.64
    TSN (Fusion) 39.40 22.70 10.89 74.29 45.72 25.26

    Joint action classification Top-1 accuracy drops from 20.54%20.54\% on seen kitchens (S1) to 10.89%10.89\% on unseen kitchens (S2) for TSN Fusion, showing that action recognition struggles to generalize across different environments.

  9. Knowl 9 — Action Anticipation Benchmark and Baseline Performance

    data/table

    The action anticipation task requires forecasting the upcoming action class Ca=(cv,cn)C_a = (c_v, c_n) before the action begins. Given an action starting at tsit_{s_i}, the model observes the preceding segment [tsi−(τa+τo),tsi−τa][t_{s_i} - (\tau_a + \tau_o), t_{s_i} - \tau_a], where τa=1.0s\tau_a = 1.0\text{s} is the anticipation horizon and τo=1.0s\tau_o = 1.0\text{s} is the observation duration.

    Performance of 2SCNN and TSN models trained to predict verbs and nouns jointly:

    Split Model Top-1 Accuracy (%) Top-5 Accuracy (%)
    Verb Noun Action Verb Noun Action
    S1 (Seen) 2SCNN (RGB) 29.76 15.15 4.32 76.03 38.56 15.21
    TSN (RGB) 31.81 16.22 6.00 76.56 42.15 18.21
    TSN (Flow) 29.64 10.30 2.93 73.70 30.09 10.92
    TSN (Fusion) 30.66 14.86 4.62 75.32 40.11 16.01
    S2 (Unseen) 2SCNN (RGB) 25.23 9.97 2.29 68.66 27.38 9.35
    TSN (RGB) 25.30 10.41 2.39 68.32 29.50 9.63
    TSN (Flow) 25.61 8.40 1.78 67.57 24.62 8.19
    TSN (Fusion) 25.37 9.76 1.74 68.25 27.24 9.05

    Action anticipation Top-1 Action accuracy peaks at 6.00%6.00\% on S1 and 2.39%2.39\% on S2 using TSN (RGB). In contrast to action recognition, optical flow and stream fusion fail to improve anticipation performance over RGB appearance alone.

  10. Knowl 10 — Annotation Quality Assurance Error Rates

    empirical result

    Manual verification across 300 randomly sampled instances per annotation category establishes the error rates of the crowdsourced pipeline:

    • Action Segment Boundaries (AiA_i): Checking whether temporal start and end timestamps completely bound the intended action without enclosing extraneous actions: error rate is 5.7%5.7\%.
    • Object Bounding Boxes (OiO_i): Checking whether bounding boxes tightly encompass target active objects with minimal background overlap and complete recall of active instances: error rate is 6.3%6.3\%.
    • Verb Classes (CVC_V): Checking the accuracy of verb clustering and assignment: error rate is 3.3%3.3\%.
    • Noun Classes (CNC_N): Checking the accuracy of noun clustering and assignment: error rate is 6.0%6.0\%.

Coverage note — None was omitted; all key contributions—dataset collection, multi-stage narration and consensus annotation formulations, class taxonomy, seen/unseen splits, baseline architectures, and benchmark results for object detection, action recognition, action anticipation, and annotation quality validation—are fully covered.

References

  1. 1.Abu-El-Haija, S., Kothari, N., Lee, J., Natsev, P., Toderici, G., Varadarajan, B., Vijayanarasimhan, S.: YouTube-8M: A Large-Scale Video Classification Benchmark. In: CoRR (2016)
  2. 2.Alletto, S., Serra, G., Calderara, S., Cucchiara, R.: Understanding social relationships in egocentric vision. In: Pattern Recognition (2015)
  3. 3.Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C.L., Parikh, D.: VQA: Visual Question Answering. In: ICCV (2015)
  4. 4.Banerjee, S., Pedersen, T.: An adapted lesk algorithm for word sense disambiguation using wordnet. In: CICLing (2002)
  5. 5.Carnegie Mellon University: CMU sphinx. https://cmusphinx.github.io/
  6. 6.Damen, D., Leelasawassuk, T., Haines, O., Calway, A., Mayol-Cuevas, W.: You-do, I-learn: Discovering task relevant objects and their modes of interaction from multi-user egocentric video. In: BMVC (2014)
  7. 7.Das, A., Kottur, S., Gupta, K., Singh, A., Yadav, D., Moura, J.M., Parikh, D., Batra, D.: Visual Dialog. In: CVPR (2017)
  8. 8.De La Torre, F., Hodgins, J., Bargteil, A., Martin, X., Macey, J., Collado, A., Beltran, P.: Guide to the Carnegie Mellon University Multimodal Activity (CMU-MMAC) database. In: Robotics Institute (2008)
  9. 9.Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., Fei-Fei, L.: Imagenet: A large-scale hierarchical image database. In: CVPR (2009)
  10. 10.Doughty, H., Damen, D., Mayol-Cuevas, W.: Who’s better? who’s best? pairwise deep ranking for skill determination. In: CVPR (2018)
  11. 11.Everingham, M., Van Gool, L., Williams, C.K.I., Winn, J., Zisserman, A.: The PASCAL Visual Object Classes (VOC) Challenge. In: IJCV (2010)
  12. 12.Fathi, A., Hodgins, J., Rehg, J.: Social interactions: A first-person perspective. In: CVPR (2012)
  13. 13.Fathi, A., Li, Y., Rehg, J.: Learning to recognize daily actions using gaze. In: ECCV (2012)
  14. 14.Fouhey, D.F., Kuo, W.c., Efros, A.A., Malik, J.: From lifestyle vlogs to everyday interactions. arXiv preprint arXiv:1712.02310 (2017)
  15. 15.Furnari, A., Battiato, S., Grauman, K., Farinella, G.M.: Next-active-object prediction from egocentric videos. In: JVCIR (2017)
  16. 16.Georgia Tech: Extended GTEA Gaze+. http://webshare.ipat.gatech.edu/coc-rim-wall-lab/web/yli440/egtea gp (2018)
  17. 17.Google: Google cloud speech api. https://cloud.google.com/speech
  18. 18.Goyal, R., Kahou, S.E., Michalski, V., Materzynska, J., Westphal, S., Kim, H., Haenel, V., Fr¨und, I., Yianilos, P., Mueller-Freitag, M., Hoppe, F., Thurau, C., Bax, I., Memisevic, R.: The ”something something” video database for learning and evaluating visual common sense. In: ICCV (2017)
  19. 19.He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: CVPR (2016)
  20. 20.Heidarivincheh, F., Mirmehdi, M., Damen, D.: Action completion: A temporal model for moment detection. In: BMVC (2018)
  21. 21.Huang, J., Rathod, V., Chow, D., Sun, C., Zhu, M., Fathi, A., Lu, Z.: Tensorflow Object Detection API. https://github.com/tensorflow/models/tree/master/research/object detection
  22. 22.Huang, J., Rathod, V., Sun, C., Zhu, M., Korattikara, A., Fathi, A., Fischer, I., Wojna, Z., Song, Y., Guadarrama, S., et al.: Speed/accuracy trade-offs for modern convolutional object detectors. In: CVPR (2017)
  23. 23.IBM: IBM watson speech to text. https://www.ibm.com/watson/services/speech-to-text
  24. 24.Ioffe, S., Szegedy, C.: Batch normalization: Accelerating deep network training by reducing internal covariate shift. In: ICML (2015)
  25. 25.Kalogeiton, V., Weinzaepfel, P., Ferrari, V., Schmid, C.: Joint learning of object and action detectors. In: ICCV (2017)
  26. 26.Karpathy, A., Fei-Fei, L.: Deep Visual-Semantic Alignments for Generating Image Descriptions. In: CVPR (2015)
  27. 27.Krizhevsky, A., Sutskever, I., Hinton, G.E.: Imagenet classification with deep convolutional neural networks. In: NIPS (2012)
  28. 28.Kuehne, H., Arslan, A., Serre, T.: The Language of Actions: Recovering the Syntax and Semantics of Goal-Directed Human Activities. In: CVPR (2014)
  29. 29.Lee, Y., Ghosh, J., Grauman, K.: Discovering important people and objects for egocentric video summarization. In: CVPR (2012)
  30. 30.Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll´ar, P., Zitnick, C.L.: Microsoft COCO: Common objects in context. In: ECCV (2014)
  31. 31.Mikolov, T., Chen, K., Corrado, G., Dean, J.: Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781 (2013)
  32. 32.Miller, G.: Wordnet: a lexical database for english. In: CACM (1995)
  33. 33.Moltisanti, D., Wray, M., Mayol-Cuevas, W., Damen, D.: Trespassing the boundaries: Labeling temporal bounds for object interactions in egocentric video. In: ICCV (2017)
  34. 34.Nair, A., Chen, D., Agrawal, P., Isola, P., Abbeel, P., Malik, J., Levine, S.: Combining self-supervised learning and imitation for vision-based rope manipulation. In: ICRA (2017)
  35. 35.Park, H.S., Hwang, J.J., Niu, Y., Shi, J.: Egocentric future localization. In: CVPR (2016)
  36. 36.Pirsiavash, H., Ramanan, D.: Detecting activities of daily living in first-person camera views. In: CVPR (2012)
  37. 37.Ren, S., He, K., Girshick, R., Sun, J.: Faster R-CNN: Towards real-time object detection with region proposal networks. In: NIPS (2015)
  38. 38.Rohrbach, A., Rohrbach, M., Tandon, N., Schiele, B.: A Dataset for Movie Description. In: CVPR (2015)
  39. 39.Rohrbach, M., Amin, S., Andriluka, M., Schiele, B.: A Database for Fine Grained Activity Detection of Cooking Activities. In: CVPR (2012)
  40. 40.Ryoo, M.S., Matthies, L.: First-person activity recognition: What are they doing to me? In: CVPR (2013)
  41. 41.Sigurdsson, G.A., Gupta, A., Schmid, C., Farhadi, A., Alahari, K.: Charades-ego: A large-scale dataset of paired third and first person videos. In: ArXiv (2018)
  42. 42.Sigurdsson, G.A., Varol, G., Wang, X., Farhadi, A., Laptev, I., Gupta, A.: Hollywood in homes: Crowdsourcing data collection for activity understanding. In: ECCV (2016)
  43. 43.Simonyan, K., Zisserman, A.: Two-stream convolutional networks for action recognition in videos. In: Advances in neural information processing systems. pp. 568–576 (2014)
  44. 44.Stein, S., McKenna, S.: Combining Embedded Accelerometers with Computer Vision for Recognizing Food Preparation Activities. In: UbiComp (2013)
  45. 45.Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., Rabinovich, A.: Going deeper with convolutions. In: CVPR (2015)
  46. 46.Tapaswi, M., Zhu, Y., Stiefelhagen, R., Torralba, A., Urtasun, R., Fidler, S.: MovieQA: Understanding stories in movies through question-answering. In: CVPR (2016)
  47. 47.Vondrick, C., Pirsiavash, H., Torralba, A.: Anticipating visual representations from unlabeled video. In: CVPR (2016)
  48. 48.Wang, L., Xiong, Y., Wang, Z., Qiao, Y., Lin, D., Tang, X., Val Gool, L.: Temporal segment networks: Towards good practices for deep action recognition. In: ECCV (2016)
  49. 49.Yamaguchi, K.: Bbox-annotator. https://github.com/kyamagu/bbox-annotator
  50. 50.Yeung, S., Russakovsky, O., Jin, N., Andriluka, M., Mori, G., Fei-Fei, L.: Every moment counts: Dense detailed labeling of actions in complex videos. IJCV (2018)
  51. 51.Yuanjun, X.: PyTorch Temporal Segment Network. https://github.com/yjxiong/tsn-pytorch (2017)
  52. 52.Zach, C., Pock, T., Bischof, H.: A duality based approach for realtime TV-L1 optical flow. In: Pattern Recognition (2007)
  53. 53.Zhang, T., McCarthy, Z., Jow, O., Lee, D., Goldberg, K., Abbeel, P.: Deep imitation learning for complex manipulation tasks from virtual reality teleoperation. In: ICRA (2018)
  54. 54.Zhao, H., Yan, Z., Wang, H., Torresani, L., Torralba, A.: SLAC: A Sparsely Labeled Dataset for Action Classification and Localization. arXiv preprint arXiv:1712.09374 (2017)
  55. 55.Zhou, B., Zhao, H., Puig, X., Fidler, S., Barriuso, A., Torralba, A.: Scene parsing through ade20k dataset. In: CVPR (2017)
  56. 56.Zhou, L., Xu, C., Corso, J.J.: Towards automatic learning of procedures from web instructional videos. arXiv preprint arXiv:1703.09788 (2017)

Citation

MLA
Damen, D., et al. “Scaling Egocentric Vision: The EPIC-KITCHENS Dataset”. arXiv, 2018, http://arxiv.org/abs/1804.02748v2.
APA
Damen, D., Doughty, H., Farinella, G. M., Fidler, S., Furnari, A., Kazakos, E., Moltisanti, D., Munro, J., Perrett, T., Price, W., & Wray, M. (2018). Scaling Egocentric Vision: The EPIC-KITCHENS Dataset. arXiv. http://arxiv.org/abs/1804.02748v2
Chicago
Damen, D., H. Doughty, G. M. Farinella, et al. 2018. “Scaling Egocentric Vision: The EPIC-KITCHENS Dataset”. arXiv. http://arxiv.org/abs/1804.02748v2.
Harvard
Damen, D. et al. (2018) “Scaling Egocentric Vision: The EPIC-KITCHENS Dataset”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1804.02748v2.
Vancouver
1. Damen D, Doughty H, Farinella GM, et al (2018) Scaling Egocentric Vision: The EPIC-KITCHENS Dataset. arXiv

BibTeX

@article{damen2018scaling,
  title = {Scaling Egocentric Vision: The EPIC-KITCHENS Dataset},
  author = {Damen, Dima and Doughty, Hazel and Farinella, Giovanni Maria and Fidler, Sanja and Furnari, Antonino and Kazakos, Evangelos and Moltisanti, Davide and Munro, Jonathan and Perrett, Toby and Price, Will and Wray, Michael},
  year = {2018},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1804.02748v2},
  eprint = {1804.02748}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors