The “Something Something” Video Database for Learning and Evaluating Visual Common Sense
Raghav GoyalSamira Ebrahimi KahouVincent MichalskiJoanna MaterzyńskaSusanne WestphalHeuna KimValentin HaenelIngo FruendPeter YianilosMoritz Mueller-Freitag
Introduces the Something-Something video dataset containing over 100,000 crowd-sourced clips designed to benchmark and train neural networks on physical reasoning and visual common sense rather than superficial object recognition.
Current artificial intelligence systems perform remarkably well at classifying objects in still images, yet they frequently fail to reason about basic physical interactions and everyday common sense. Most existing video benchmarks focus on high-level, human-centered activities like sports or movies, where models can often guess the correct category from a single static frame rather than by tracking motion, forces, or object properties. As computer vision expands into practical domains like robotics and human-robot interaction, systems require visual models grounded in intuitive physics, such as understanding gravity, object permanence, and material deformability.
The article introduces the "something-something" video database and evaluates baseline deep learning architectures to establish whether neural networks can learn fine-grained physical concepts and intuitive common sense directly from short video demonstrations.
To build the database, the researchers deployed a crowdsourcing platform where 1,133 workers recorded short, synchronized video clips acting out structured caption templates, such as "Dropping [something] into [something]." Workers chose the physical objects and provided the specific nouns, yielding 108,499 video clips across 174 fine-grained classes with an average duration of 4.03 seconds. To prevent models from taking shortcuts—such as relying on background context or hand posture—the authors organized tasks into contrastive action groups that included subtle variations and pretended actions. Baseline computer vision architectures, including two-dimensional convolutional networks, recurrent networks, and three-dimensional convolutional networks, were tested on classification benchmarks ranging from 10 to 174 classes.
The baseline evaluations revealed that standard neural networks struggle significantly with fine-grained physical reasoning. On a hand-selected 10-class benchmark of simplified actions, the best combined model achieved an error rate of 44.9% (top-1) and 27.1% (top-2). When the task scaled to 40 classes, the error rate increased to 63.8% (top-1) and 50.7% (top-2). Across all 174 categories, a pre-trained three-dimensional convolutional model recorded an 88.5% top-1 error rate and a 70.3% top-5 error rate. Architectures incorporating three-dimensional convolutions generally outperformed purely two-dimensional image-based models, yet subtle distinctions—such as distinguishing genuine actions from pretended ones—remained difficult across all models.
These results indicate that solving real-world physical reasoning requires models to integrate temporal dynamics rather than rely on static visual features. For organizations building autonomous agents, robotic systems, or language-grounded vision tools, off-the-shelf image models present serious performance risks when physical actions must be accurately interpreted. Achieving reliable visual common sense demands architectures specifically designed to capture fine temporal details, spatial relations, and object state changes.
The authors recommend adopting a curriculum learning strategy, expanding data and model complexity incrementally as validation accuracy improves. Researchers and engineering teams should develop more sophisticated temporal network architectures rather than relying on standard frame-averaging baselines. For long-term development, the authors advocate expanding structured natural language captioning to bridge visual perception and broader language understanding.
The findings are bounded by the baseline nature of the evaluated models and the semantic ambiguities inherent in scaling crowdsourced language labels. However, the data collection methodology and baseline benchmarks provide high confidence that fine-grained video benchmarks are essential for exposing and addressing the current limits of machine common sense.
- Paper: UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild, Khurram Soomro et al. (2012). It introduces a foundational benchmark for action recognition in unconstrained web video, establishing the high-level activity paradigm that Something-Something seeks to transcend with fine-grained physical commonsense.
- Paper: Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset, João Carreira et al. (2017). It establishes large-scale video action pre-training architectures and benchmarks, illustrating the traditional high-level action recognition setup compared against Something-Something's template-based physical tasks.
- Paper: Two-Stream Convolutional Networks for Action Recognition in Videos, Karen Simonyan et al. (2014). It introduces the two-stream convolutional architecture combining appearance and motion cues, which serves as a core baseline methodology for video action recognition.
- Paper: Learning Spatiotemporal Features with 3D Convolutional Networks, Du Tran et al. (2015). It proposes 3D convolutional spatiotemporal feature learning directly from video, providing a foundational backbone used to benchmark models on video datasets.
- Paper: Temporal Segment Networks: Towards Good Practices for Deep Action Recognition, Limin Wang et al. (2016). It establishes long-range temporal segment modeling for deep action recognition, offering vital training strategies for video-level action understanding.
- Paper: Beyond short snippets: Deep networks for video classification, Joe Yue-Hei Ng et al. (2015). It demonstrates how to aggregate visual information across entire video durations using temporal pooling and LSTMs, directly motivating the temporal reasoning evaluated in Something-Something.
- Paper: MSR-VTT: A Large Video Description Dataset for Bridging Video and Language, Jun Xu et al. (2016). It provides a benchmark bridging video and natural language descriptions, framing the relationship between multi-modal captioning and temporal scene understanding.
- Paper: ImageNet Large Scale Visual Recognition Challenge, Olga Russakovsky et al. (2014). It outlines the foundational ImageNet classification paradigm and crowdsourcing methodology that visual common-sense video benchmarks build upon and extend into the temporal domain.
- Paper: ViViT: A Video Vision Transformer, Anurag Arnab et al. (2021). It evaluates pure spatio-temporal video vision transformers on benchmarks including Something-Something to demonstrate superior modeling of fine-grained physical interactions over 3D CNNs.
- Paper: Video Swin Transformer, Ze Liu et al. (2021). It advances physical interaction modeling by using spatio-temporal transformers to detect object state changes and pinpoint the precise moment of physical transformation in video.
- Paper: VideoBERT: A Joint Model for Video and Language Representation Learning, Chen Sun et al. (2019). It applies self-supervised transformer pretraining across video and language to learn joint multi-modal representations of actionable physical tasks.
- Paper: PIQA: Reasoning about Physical Commonsense in Natural Language, Yonatan Bisk et al. (2019). It extends physical commonsense reasoning evaluation into natural language question answering, testing whether models understand everyday object affordances and physical dynamics.
- Paper: Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis, Chaoyou Fu et al. (2025). It establishes a comprehensive multi-modal LLM evaluation benchmark targeting complex temporal and physical reasoning across extended real-world video scenarios.
- Paper: WorldSimBench: Towards Video Generation Models as World Simulators, Yiran Qin et al. (2025). It investigates whether video generation models can function as physical world simulators by evaluating their ability to simulate actionable physical interactions and real-world dynamics.
- Paper: GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering, Drew A. Hudson et al. (2019). It builds upon the goal of evaluating visual common sense by introducing compositionally balanced questions grounded in structured semantic scene graphs.
