Hollywood in Homes: Crowdsourcing Data Collection for Activity Understanding
Gunnar A. SigurdssonGül VarolXiaolong WangAli FarhadiIvan LaptevAbhinav Gupta
Introduces a crowdsourced video collection methodology that captures realistic home activities across hundreds of participants, establishing the Charades dataset to advance human action recognition and video description generation.
Computer vision models increasingly support assistive technologies, smart home systems, and robotics, yet their training has relied heavily on curated web media, movies, and sports footage. These sources do not reflect the mundane, multi-step activities of daily living, while existing indoor datasets collected in controlled laboratory settings suffer from narrow environments and limited scalability. To address this gap, the article introduces a distributed framework to crowdsource realistic video creation at scale and demonstrates its utility by benchmarking leading activity recognition and video captioning algorithms.
To construct the resulting Charades benchmark, the researchers deployed a three-stage crowdsourcing workflow using Amazon Mechanical Turk across three continents. Online contributors generated realistic household scripts guided by a vocabulary of 40 objects and 30 actions across 15 indoor scene types, recorded themselves acting out these scripts in their own homes, and subsequently verified and temporally annotated the recordings. Financial incentives, such as sign-up and retention bonuses, reduced the base cost to approximately $1 per video and drove worker retention up by 34% and individual output by 109%. The resulting collection spans 9,848 videos with an average length of 30 seconds, 27,847 natural-language descriptions, and 66,500 temporally localized intervals covering 157 distinct action classes.
The evaluation yielded several critical performance findings. First, state-of-the-art action recognition models struggled significantly in realistic indoor environments; the best individual baseline achieved only 17.2% mean average precision, which rose to 18.6% when combining all models, well below performance levels typically observed on standard web or sports benchmarks. Second, the vast majority of classification errors stemmed from subtle differences among actions involving the same object, such as holding versus taking an item, whereas actions without specific object interactions achieved a much higher accuracy of 38.9%. Third, automated description models generated grammatically fluent sentences but frequently failed to identify the correct core activities, indicating a tendency to default to broad linguistic priors rather than precise visual evidence.
These findings indicate that existing computer vision systems carry major blind spots when applied to unstructured, real-world human environments. Relying on current architectures for fine-grained assistive tasks introduces operational risks of misinterpreting user intent and complex human-object interactions. The article establishes that scaling realistic data collection is economically feasible via crowdsourcing, but closing the performance gap will require algorithms designed specifically to track subtle object state changes and context rather than just global motion.
Organizations developing computer vision and robotic systems should incorporate realistic household benchmarks like Charades into their evaluation pipelines to measure deployment readiness accurately. Future development should prioritize fine-grained object interaction modeling and multi-modal grounding before deploying autonomous systems in home settings. While the dataset provides high label precision (95.6%) and broad geographic reach, its reliance on scripted acting within a controlled vocabulary may still introduce subtle behavioral biases compared to entirely unscripted life, requiring continued testing as richer naturalistic data becomes available.
- Paper: ActivityNet: A large-scale video benchmark for human activity understanding, Fabian Caba Heilbron et al. (2015). ActivityNet established crowdsourced collection and untrimmed multi-event video annotation benchmarks for complex daily activities, directly motivating Charades' focus on mundane household actions.
- Paper: UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild, Khurram Soomro et al. (2012). UCF101 provided the standard in-the-wild video action classification benchmark against which Hollywood in Homes positions its more diverse and realistic crowdsourced home environment collection.
- Paper: Two-Stream Convolutional Networks for Action Recognition in Videos, Karen Simonyan et al. (2014). Simonyan and Zisserman introduced the fundamental two-stream convolutional network framework used to establish standard action recognition baselines on video datasets like Charades.
- Paper: Collecting Highly Parallel Data for Paraphrase Evaluation, David L. Chen et al. (2011). This work pioneered crowdsourcing methods for eliciting unconstrained natural language descriptions of video clips, laying foundational methodologies for crowdsourced video-text collection.
- Paper: Learning Spatiotemporal Features with 3D Convolutional Networks, Du Tran et al. (2015). C3D established spatiotemporal 3D convolutional representations for video feature extraction, providing key architectural concepts for baseline action understanding evaluations.
- Paper: Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset, João Carreira et al. (2017). Carreira and Zisserman advance deep video action modeling beyond initial crowdsourced benchmarks by introducing Two-Stream Inflated 3D ConvNets (I3D) evaluated on complex action datasets.
- Paper: Dense-Captioning Events in Videos, Ranjay Krishna et al. (2017). This work extends video description generation by formulating dense-captioning architectures that jointly detect and describe temporally localized events in untrimmed video streams.
- Paper: Ego4D: Around the World in 3,000 Hours of Egocentric Video, Kristen Grauman et al. (2021). Ego4D builds upon crowdsourced daily-activity capture in home environments by scaling natural human interaction understanding to thousands of hours of worldwide egocentric video.
- Paper: VideoBERT: A Joint Model for Video and Language Representation Learning, Chen Sun et al. (2019). VideoBERT extends multimodal video-language learning from action-caption benchmarks by pre-training joint visual-linguistic transformer models on untrimmed everyday video recordings.
- Paper: HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips, Antoine Miech et al. (2019). HowTo100M scales video-and-language representation learning by leveraging vast collections of instructional daily-activity videos for text-video alignment and retrieval.
- Paper: Temporal Convolutional Networks for Action Segmentation and Detection, Colin Lea et al. (2016). Lea et al. evaluate temporal convolutional networks to model long-range action compositions and temporal boundaries in fine-grained daily activity videos.
- Paper: Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval, Max Bain et al. (2021). Frozen in Time presents an end-to-end dual-encoder architecture that unifies image and video description matching, advancing retrieval across benchmark video-text datasets.
