Hollywood in Homes: Crowdsourcing Data Collection for Activity Understanding

Gunnar A. SigurdssonGül VarolXiaolong WangAli FarhadiIvan LaptevAbhinav Gupta

article2016ECCV1,439 citations

Introduces a crowdsourced video collection methodology that captures realistic home activities across hundreds of participants, establishing the Charades dataset to advance human action recognition and video description generation.

Listen

Computer vision models increasingly support assistive technologies, smart home systems, and robotics, yet their training has relied heavily on curated web media, movies, and sports footage. These sources do not reflect the mundane, multi-step activities of daily living, while existing indoor datasets collected in controlled laboratory settings suffer from narrow environments and limited scalability. To address this gap, the article introduces a distributed framework to crowdsource realistic video creation at scale and demonstrates its utility by benchmarking leading activity recognition and video captioning algorithms.

To construct the resulting Charades benchmark, the researchers deployed a three-stage crowdsourcing workflow using Amazon Mechanical Turk across three continents. Online contributors generated realistic household scripts guided by a vocabulary of 40 objects and 30 actions across 15 indoor scene types, recorded themselves acting out these scripts in their own homes, and subsequently verified and temporally annotated the recordings. Financial incentives, such as sign-up and retention bonuses, reduced the base cost to approximately $1 per video and drove worker retention up by 34% and individual output by 109%. The resulting collection spans 9,848 videos with an average length of 30 seconds, 27,847 natural-language descriptions, and 66,500 temporally localized intervals covering 157 distinct action classes.

The evaluation yielded several critical performance findings. First, state-of-the-art action recognition models struggled significantly in realistic indoor environments; the best individual baseline achieved only 17.2% mean average precision, which rose to 18.6% when combining all models, well below performance levels typically observed on standard web or sports benchmarks. Second, the vast majority of classification errors stemmed from subtle differences among actions involving the same object, such as holding versus taking an item, whereas actions without specific object interactions achieved a much higher accuracy of 38.9%. Third, automated description models generated grammatically fluent sentences but frequently failed to identify the correct core activities, indicating a tendency to default to broad linguistic priors rather than precise visual evidence.

These findings indicate that existing computer vision systems carry major blind spots when applied to unstructured, real-world human environments. Relying on current architectures for fine-grained assistive tasks introduces operational risks of misinterpreting user intent and complex human-object interactions. The article establishes that scaling realistic data collection is economically feasible via crowdsourcing, but closing the performance gap will require algorithms designed specifically to track subtle object state changes and context rather than just global motion.

Organizations developing computer vision and robotic systems should incorporate realistic household benchmarks like Charades into their evaluation pipelines to measure deployment readiness accurately. Future development should prioritize fine-grained object interaction modeling and multi-modal grounding before deploying autonomous systems in home settings. While the dataset provides high label precision (95.6%) and broad geographic reach, its reliance on scripted acting within a controlled vocabulary may still introduce subtle behavioral biases compared to entirely unscripted life, requiring continued testing as richer naturalistic data becomes available.

Cover for Hollywood in Homes: Crowdsourcing Data Collection for Activity Understanding

Abstract

Computer vision has a great potential to help our daily lives by searching for lost keys, watering flowers or reminding us to take a pill. To succeed with such tasks, computer vision methods need to be trained from real and diverse examples of our daily dynamic scenes. While most of such scenes are not particularly exciting, they typically do not appear on YouTube, in movies or TV broadcasts. So how do we collect sufficiently many diverse but boring samples representing our lives? We propose a novel Hollywood in Homes approach to collect such data. Instead of shooting videos in the lab, we ensure diversity by distributing and crowdsourcing the whole process of video creation from script writing to video recording and annotation. Following this procedure we collect a new dataset, Charades, with hundreds of people recording videos in their own homes, acting out casual everyday activities. The dataset is composed of 9,848 annotated videos with an average length of 30 seconds, showing activities of 267 people from three continents. Each video is annotated by multiple free-text descriptions, action labels, action intervals and classes of interacted objects. In total, Charades provides 27,847 video descriptions, 66,500 temporally localized intervals for 157 action classes and 41,104 labels for 46 object classes. Using this rich data, we evaluate and provide baseline results for several tasks including action recognition and automatic description generation. We believe that the realism, diversity, and casual nature of this dataset will present unique challenges and new opportunities for computer vision community.

Table of Contents

  • 1 Introduction
  • 2 Hollywood in Homes
  • 2.1 Generating Scripts
  • 2.2 Generating Videos
  • 2.3 Annotations
  • 3 Charades v1.0 Analysis
  • 4 Applications
  • 4.1 Action Classification
  • 4.2 Sentence Prediction
  • 5 Conclusions
  • 6 Acknowledgements
  • References

Knowls

  1. Knowl 1 — Hollywood in Homes Crowdsourcing Framework

    model/method

    The Hollywood in Homes framework is a distributed crowdsourcing methodology for creating diverse, realistic video datasets of everyday indoor activities using Amazon Mechanical Turk (AMT). The framework structures data collection into three sequential stages:

    1. Script Generation: Workers receive a designated room scene (selected from 15 indoor categories) alongside five randomly chosen seed objects and five seed actions derived from lexical analysis of movie scripts. Workers write a realistic, commonplace script (a short paragraph) depicting one or two people in their home using at least two of the provided objects and two actions.
    2. Video Direction and Acting: Independent workers read the generated scripts and record themselves in their own homes acting out the described activity in a video averaging 30 seconds.
    3. Verification and Annotation: Separate workers verify that the uploaded videos accurately depict the assigned script, followed by multi-stage annotation of interacted objects, action class labels, start and end timestamps for each action interval, and free-text video descriptions.
  2. Knowl 2 — Charades v1.0 Dataset

    definition

    Charades v1.0 is a large-scale video dataset of casual, everyday household activities collected in real domestic environments through crowdsourced scriptwriting, acting, and verification. The dataset specifications include:

    • Scale: 9,848 videos with an average duration of 30.1 seconds (totalling over 82 hours of footage).
    • Diversity: Recorded by 267 distinct people across three continents, with over 15% of videos featuring multiple people.
    • Scenes: 15 indoor home room categories (e.g., Living Room, Bedroom, Kitchen, Bathroom, Home Office, Dining Room, Entryway, Hallway, Stairs, Laundry Room, Walk-in Closet, Recreation Room, Pantry, Garage, Basement).
    • Vocabulary and Actions: Interacting with 46 object classes and 30 verbs, forming 157 fine-grained action classes based on (verb, proposition, noun) triplets.
    • Annotations: 66,500 temporally localized action intervals with an average duration of 12.8 seconds (averaging 6.8 relevant action instances per video), 41,104 object class labels, and 27,847 free-text video descriptions.
  3. Knowl 3 — Action Classification Baseline Performance on Charades

    data/table

    Action classification on Charades evaluates the multi-label classification accuracy of video models across all 157 action classes, measured in mean Average Precision (mAP, in %). Due to heavy visual diversity, clutter, and fine-grained object interactions, models achieve lower accuracy on Charades than on conventional sports or movie benchmarks:

    Method mAP (%)
    Random Baseline 5.9
    C3D Features + Linear SVM 10.9
    AlexNet (fc6 frame average) + Linear SVM 11.3
    Two-Stream Network (Class-Balanced) 11.9
    Two-Stream Network (VGG-16) 14.3
    Improved Dense Trajectories (IDT) + Fisher Vector + Linear SVM 17.2
    Combined (Late Fusion of all baselines) 18.6

    Handcrafted Improved Dense Trajectory (IDT) features outperform deep convolutional models on this benchmark (17.2% vs. 14.3% for two-stream networks and 10.9% for C3D), while late fusion across all baseline predictors achieves 18.6% mAP.

  4. Knowl 4 — Improved Dense Trajectories Parameter and Descriptor Ablation

    data/table

    The classification accuracy (mAP, in %) of Improved Dense Trajectories (IDT) on the 157 action classes of Charades is evaluated across combinations of local video descriptors—Histogram of Oriented Gradients (HOG), Histogram of Optical Flow (HOF), and Motion Boundary Histograms (MBH)—and across varying numbers of Gaussian Mixture Model (GMM) components (K∈{64,128,256}K \in \{64, 128, 256\}) used in Fisher vector encoding:

    GMM Vocabulary (KK) HOG HOF MBH HOG+MBH HOG+HOF+MBH
    K=64K=64 12.3 13.9 15.0 15.8 16.5
    K=128K=128 12.7 14.3 15.4 16.2 16.9
    K=256K=256 13.0 14.4 15.5 16.5 17.2

    Performance increases monotonically with vocabulary size (KK) and when combining shape, optical flow, and boundary descriptors, reaching the maximum mAP of 17.2% with K=256K=256 and all three descriptors.

  5. Knowl 5 — Fine-Grained Object Interaction Bottleneck in Action Recognition

    empirical result

    Analysis of the confusion matrix for the Combined baseline model on Charades reveals key error patterns across everyday actions:

    • Same-Object Actions: The majority of classification errors occur between distinct actions involving the exact same object (e.g., confusing putting on clothes with taking clothes from somewhere, or opening a box with closing a box).
    • Functionally Similar Objects: High confusion occurs across action classes involving objects with similar physical usage, such as substitutions among blanket, clothes, and towel, or between couch and bed.
    • Object-Free Actions: For the subset of action categories that do not involve explicit object interactions (e.g., standing up, sneezing), the model achieves 38.9% mAP, compared to the overall dataset average of 18.6% mAP.

    These results indicate that fine-grained human-object interaction recognition is the primary bottleneck in domestic activity understanding.

  6. Knowl 6 — Video Sentence Prediction Benchmark on Charades

    data/table

    Video description generation is evaluated on Charades under two distinct target settings: predicting the original human-written generation Script (1 reference sentence per video) and predicting the observed Description provided by post-hoc video viewers (an average of 2.4 reference sentences per video). Models are evaluated using CIDEr, BLEU1…4\text{BLEU}_{1\dots4}, ROUGEL\text{ROUGE}_L, and METEOR metrics:

    Script Target Description Target
    Metric RW Random NN S2VT Human RW Random NN S2VT Human
    CIDEr 0.03 0.08 0.11 0.17 0.51 0.04 0.05 0.07 0.14 0.53
    BLEU4\text{BLEU}_4 0.00 0.03 0.03 0.06 0.10 0.00 0.04 0.05 0.11 0.20
    BLEU3\text{BLEU}_3 0.01 0.07 0.07 0.12 0.16 0.02 0.09 0.10 0.18 0.29
    BLEU2\text{BLEU}_2 0.09 0.15 0.15 0.21 0.27 0.09 0.20 0.21 0.30 0.43
    BLEU1\text{BLEU}_1 0.37 0.29 0.29 0.36 0.43 0.38 0.40 0.40 0.49 0.62
    ROUGEL\text{ROUGE}_L 0.21 0.24 0.25 0.31 0.35 0.22 0.27 0.28 0.35 0.44
    METEOR 0.10 0.11 0.12 0.13 0.20 0.11 0.13 0.14 0.16 0.24

    The Sequence-to-Sequence Video-to-Text (S2VT) baseline outperforms Random Words (RW), Random Sentences (Random), and Nearest Neighbors (NN), achieving a CIDEr score of 0.17 on scripts and 0.14 on descriptions, but falls substantially short of human performance (0.51 and 0.53 CIDEr, respectively).

  7. Knowl 7 — Charades Dataset Train/Test Split Protocol

    experimental setup

    The Charades v1.0 dataset partition into training and test sets enforces four joint constraints:

    1. Worker Disjointness: An Amazon Mechanical Turk worker appearing in the training set cannot appear in the test set, preventing visual memorization of specific homes or actors.
    2. Distribution Matching: Category frequencies in the test set match the category frequencies in the training set.
    3. Minimum Class Representation: Every action class contains at least 25 training videos and at least 6 test videos.
    4. Dominance Prevention: No single worker contributes a disproportionate fraction of test set videos.

    Partitioning the pool of workers (80% training / 20% test) subject to these constraints produces:

    • Training Set: 7,985 videos comprising 49,809 annotated action intervals.
    • Test Set: 1,863 videos comprising 16,691 annotated action intervals.
  8. Knowl 8 — Crowdsourced Worker Incentive and Retention Mechanism

    model/method

    To overcome the reluctance of crowd workers to film videos in their own homes and keep costs manageable, the Hollywood in Homes collection uses a multi-tier financial incentive structure:

    • Base Pay: $1.00 per completed 30-second video.
    • Sign-up Bonus: A $5.00 bonus awarded on the worker's first video submission, which increased new worker recruitment by 211% at an overall cost increase of 17%.
    • Referral Bonus: A $5.00 bonus paid when a referred worker submitted at least 15 videos (claimed by 4% of the workforce).
    • Retention Bonus: A performance bonus awarded every 15th submitted video, increasing worker return rates by 34% and per-worker output by 109% (representing a 33% increase over base pay).
    • Cost Allocation: The total collection budget was distributed across 65% base video compensation, 21% performance bonuses, 11% recruitment bonuses, and 3% video-script verification tasks, achieving peak production of 1,225 videos per day from 72 workers.
  9. Knowl 9 — Multi-Stage Video Action Annotation and Localization Pipeline

    algorithm

    Annotation of crowdsourced videos is conducted using a staged verification and localization procedure:

    Input: Recorded video VV, generating script SS
    Output: Interacted object set OVO_V, action labels AV⊆{1,…,157}A_V \subseteq \{1, \dots, 157\}, temporal intervals TV={(tstart(a),tend(a))}a∈AVT_V = \{(t_{start}^{(a)}, t_{end}^{(a)})\}_{a \in A_V}
    1. Description Generation:
         Independent workers watch VV and write a free-text sentence description DD.
    2. Object Extraction and Verification:
         Extract candidate objects from SS and DD via NLP.
         Workers verify which extracted objects are actively interacted with in VV, yielding OVO_V.
    3. Candidate Action Shortlisting and Verification:
         Based on OVO_V, select a candidate subset of 4 to 5 action classes matching those object interactions.
         Workers verify the presence of each candidate action in VV, yielding verified set AVA_V.
    4. Temporal Localization:
         for each action a∈AVa \in A_V:
             Workers mark start timestamp tstart(a)t_{start}^{(a)} and end timestamp tend(a)t_{end}^{(a)} in VV.
    5. Return OVO_V, AVA_V, and TVT_V.

    This pipeline achieves an action label precision of 95.6% when validated against 19 consensus annotation iterations on a 50-video benchmark subset.

Coverage note — No substantial contributed material was omitted.

References

  1. 1.Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., Fei-Fei, L.: Imagenet: A large-scale hierarchical image database. In: Computer Vision and Pattern Recognition (CVPR), IEEE (2009) 248–255 1
  2. 2.Zhou, B., Lapedriza, A., Xiao, J., Torralba, A., Oliva, A.: Learning deep features for scene recognition using places database. In: Neural Information Processing Systems (NIPS). (2014) 487–495 1
  3. 3.Caba Heilbron, F., Escorcia, V., Ghanem, B., Carlos Niebles, J.: Activitynet: A large-scale video benchmark for human activity understanding. In: Computer Vision and Pattern Recognition (CVPR), IEEE (2015) 961–970 2, 4
  4. 4.Liu, J., Luo, J., Shah, M.: Recognizing realistic actions from videos in the wild. In: Computer Vision and Pattern Recognition (CVPR), IEEE (2009) 1996–2003 2, 3
  5. 5.Gorban, A., Idrees, H., Jiang, Y.G., Roshan Zamir, A., Laptev, I., Shah, M., Sukthankar, R.: THUMOS challenge: Action recognition with a large number of classes. http://www.thumos.info/ (2015) 2, 3, 4
  6. 6.Karpathy, A., Toderici, G., Shetty, S., Leung, T., Sukthankar, R., Fei-Fei, L.: Large-scale video classification with convolutional neural networks. In: Computer Vision and Pattern Recognition (CVPR), IEEE (2014) 1725–1732 2, 3, 4, 10
  7. 7.Kuehne, H., Jhuang, H., Garrote, E., Poggio, T., Serre, T.: HMDB: a large video database for human motion recognition. In: International Conference on Computer Vision (ICCV), IEEE (2011) 2556–2563 2, 3, 4
  8. 8.Soomro, K., Roshan Zamir, A., Shah, M.: UCF101: A dataset of 101 human actions classes from videos in the wild. In: CRCV-TR-12-01. (2012) 2, 3, 4
  9. 9.Laptev, I., Marsza lek, M., Schmid, C., Rozenfeld, B.: Learning realistic human actions from movies. In: Computer Vision and Pattern Recognition (CVPR), IEEE (2008) 1–8 2
  10. 10.Rodriguez, M.D., Ahmed, J., Shah, M.: Action mach a spatio-temporal maximum average correlation height filter for action recognition. In: Computer Vision and Pattern Recognition (CVPR), IEEE (2008) 1–8 2
  11. 11.Rohrbach, A., Rohrbach, M., Tandon, N., Schiele, B.: A dataset for movie description. In: Computer Vision and Pattern Recognition (CVPR), IEEE (2015) 2, 3, 4
  12. 12.Sch¨uldt, C., Laptev, I., Caputo, B.: Recognizing human actions: a local svm approach. In: International Conference on Pattern Recognition (ICPR). Volume 3., IEEE (2004) 32–36 2, 3
  13. 13.Gorelick, L., Blank, M., Shechtman, E., Irani, M., Basri, R.: Actions as space-time shapes. Transactions on Pattern Analysis and Machine Intelligence 29(12) (December 2007) 2247–2253 2
  14. 14.Rohrbach, M., Amin, S., Andriluka, M., Schiele, B.: A database for fine grained activity detection of cooking activities. In: Computer Vision and Pattern Recognition (CVPR), IEEE (2012) 1194–1201 2, 3, 4
  15. 15.Oh, S., Hoogs, A., Perera, A., Cuntoor, N., Chen, C.C., Lee, J.T., Mukherjee, S., Aggarwal, J., Lee, H., Davis, L., et al.: A large-scale benchmark dataset for event recognition in surveillance video. In: Computer Vision and Pattern Recognition (CVPR), IEEE (2011) 3153–3160 2
  16. 16.Kuehne, H., Arslan, A.B., Serre, T.: The language of actions: Recovering the syntax and semantics of goal-directed human activities. In: Computer Vision and Pattern Recognition (CVPR), IEEE (2014) 2, 3
  17. 17.Rohrbach, A., Rohrbach, M., Qiu, W., Friedrich, A., Pinkal, M., Schiele, B.: Coherent multi-sentence video description with variable level of detail. In: Pattern Recognition. Springer (2014) 184–195 2, 3
  18. 18.Marsza lek, M., Laptev, I., Schmid, C.: Actions in context. In: Computer Vision and Pattern Recognition (CVPR), IEEE (2009) 3
  19. 19.Ferrari, V., Mar´ın-Jim´enez, M., Zisserman, A.: 2d human pose estimation in tv shows. In: Statistical and Geometrical Approaches to Visual Motion Analysis. Springer (2009) 128–147 3
  20. 20.Chen, D.L., Dolan, W.B.: Collecting highly parallel data for paraphrase evaluation. In: Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies-Volume 1, Association for Computational Linguistics (2011) 190–200 3
  21. 21.Torabi, A., Pal, C., Larochelle, H., Courville, A.: Using descriptive video services to create a large data source for video annotation research. arXiv preprint arXiv:1503.01070 (2015) 3
  22. 22.Gupta, A., Davis, L.S.: Objects in action: An approach for combining action understanding and object perception. In: Computer Vision and Pattern Recognition (CVPR), IEEE (2007) 1–8 3
  23. 23.Ryoo, M.S., Aggarwal, J.K.: Spatio-temporal relationship match: Video structure comparison for recognition of complex human activities. In: International Conference on Computer Vision (ICCV), IEEE (2009) 1593–1600 3
  24. 24.Tuite, K., Snavely, N., Hsiao, D.y., Tabing, N., Popovic, Z.: Photocity: training experts at large-scale image acquisition through a competitive game. In: Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, ACM (2011) 1383–1392 4
  25. 25.Pirsiavash, H., Ramanan, D.: Detecting activities of daily living in first-person camera views. In: Computer Vision and Pattern Recognition (CVPR), IEEE (2012) 2847–2854 4
  26. 26.Iwashita, Y., Takamine, A., Kurazume, R., Ryoo, M.S.: First-person animal activity recognition from egocentric videos. In: International Conference on Pattern Recognition (ICPR), Stockholm, Sweden (August 2014) 4
  27. 27.Zitnick, C., Parikh, D.: Bringing semantics into focus using visual abstraction. In: Computer Vision and Pattern Recognition (CVPR), IEEE (2013) 3009–3016 4
  28. 28.Salton, G., Michael, J.: Mcgill. Introduction to modern information retrieval (1983) 24–51 5
  29. 29.Sigurdsson, G.A., Russakovsky, O., Farhadi, A., Laptev, I., Gupta, A.: Much ado about time: Exhaustive annotation of temporal data. arXiv preprint arXiv:1607.07429 (2016) 5, 6
  30. 30.Zipf, G.K.: The psycho-biology of language. (1935) 7
  31. 31.Simoncelli, E.P., Olshausen, B.A.: Natural image statistics and neural representation. Annual review of neuroscience 24(1) (2001) 1193–1216 7
  32. 32.Van der Maaten, L., Hinton, G.: Visualizing data using t-sne. Journal of Machine Learning Research 9(2579-2605) (2008) 85 7
  33. 33.Wang, H., Schmid, C.: Action recognition with improved trajectories. In: International Conference on Computer Vision (ICCV). (2013) 9, 11
  34. 34.Perronnin, F., S´anchez, J., Mensink, T.: Improving the fisher kernel for large-scale image classification. In: European Conference on Computer Vision (ECCV). (2010) 10
  35. 35.Simonyan, K., Zisserman, A.: Very deep convolutional networks for large-scale image recognition. In: International Conference on Learning Representations (ICLR). (2015) 10
  36. 36.Krizhevsky, A., Sutskever, I., Hinton, G.E.: Imagenet classification with deep convolutional neural networks. In: Neural Information Processing Systems (NIPS). (2012) 10
  37. 37.Simonyan, K., Zisserman, A.: Very deep convolutional networks for large-scale image recognition. CoRR abs/1409.1556 (2014) 10
  38. 38.Simonyan, K., Zisserman, A.: Two-stream convolutional networks for action recognition in videos. In: Neural Information Processing Systems (NIPS). (2014) 10
  39. 39.Tran, D., Bourdev, L., Fergus, R., Torresani, L., Paluri, M.: Learning spatiotemporal features with 3D convolutional networks. In: International Conference on Computer Vision (ICCV). (2015) 10
  40. 40.Chen, X., Fang, H., Lin, T., Vedantam, R., Gupta, S., Dollr, P., Zitnick, C.L.: Microsoft coco captions: Data collection and evaluation server. arXiv:1504.00325 (2015) 13
  41. 41.Devlin, J., Gupta, S., Girshick, R., Mitchell, M., Zitnick, C.L.: Exploring nearest neighbor approaches for image captioning. arXiv preprint arXiv:1505.04467 (2015) 14
  42. 42.Venugopalan, S., Rohrbach, M., Donahue, J., Mooney, R., Darrell, T., Saenko, K.: Sequence to sequence-video to text. In: International Conference on Computer Vision (ICCV). (2015) 4534–4542 14

Citation

MLA
Sigurdsson, G. A., et al. “Hollywood in Homes: Crowdsourcing Data Collection for Activity Understanding”. arXiv, 2016, http://arxiv.org/abs/1604.01753v3.
APA
Sigurdsson, G. A., Varol, G., Wang, X., Farhadi, A., Laptev, I., & Gupta, A. (2016). Hollywood in Homes: Crowdsourcing Data Collection for Activity Understanding. arXiv. http://arxiv.org/abs/1604.01753v3
Chicago
Sigurdsson, G. A., G. Varol, X. Wang, A. Farhadi, I. Laptev, and A. Gupta. 2016. “Hollywood in Homes: Crowdsourcing Data Collection for Activity Understanding”. arXiv. http://arxiv.org/abs/1604.01753v3.
Harvard
Sigurdsson, G.A. et al. (2016) “Hollywood in Homes: Crowdsourcing Data Collection for Activity Understanding”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1604.01753v3.
Vancouver
1. Sigurdsson GA, Varol G, Wang X, Farhadi A, Laptev I, Gupta A (2016) Hollywood in Homes: Crowdsourcing Data Collection for Activity Understanding. arXiv

BibTeX

@article{sigurdsson2016hollywood,
  title = {Hollywood in Homes: Crowdsourcing Data Collection for Activity Understanding},
  author = {Sigurdsson, Gunnar A. and Varol, Gül and Wang, Xiaolong and Farhadi, Ali and Laptev, Ivan and Gupta, Abhinav},
  year = {2016},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1604.01753v3},
  eprint = {1604.01753}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF