The “Something Something” Video Database for Learning and Evaluating Visual Common Sense

Raghav GoyalSamira Ebrahimi KahouVincent MichalskiJoanna MaterzyńskaSusanne WestphalHeuna KimValentin HaenelIngo FruendPeter YianilosMoritz Mueller-Freitag

article2017ICCV2,056 citations

Introduces the Something-Something video dataset containing over 100,000 crowd-sourced clips designed to benchmark and train neural networks on physical reasoning and visual common sense rather than superficial object recognition.

Listen

Current artificial intelligence systems perform remarkably well at classifying objects in still images, yet they frequently fail to reason about basic physical interactions and everyday common sense. Most existing video benchmarks focus on high-level, human-centered activities like sports or movies, where models can often guess the correct category from a single static frame rather than by tracking motion, forces, or object properties. As computer vision expands into practical domains like robotics and human-robot interaction, systems require visual models grounded in intuitive physics, such as understanding gravity, object permanence, and material deformability.

The article introduces the "something-something" video database and evaluates baseline deep learning architectures to establish whether neural networks can learn fine-grained physical concepts and intuitive common sense directly from short video demonstrations.

To build the database, the researchers deployed a crowdsourcing platform where 1,133 workers recorded short, synchronized video clips acting out structured caption templates, such as "Dropping [something] into [something]." Workers chose the physical objects and provided the specific nouns, yielding 108,499 video clips across 174 fine-grained classes with an average duration of 4.03 seconds. To prevent models from taking shortcuts—such as relying on background context or hand posture—the authors organized tasks into contrastive action groups that included subtle variations and pretended actions. Baseline computer vision architectures, including two-dimensional convolutional networks, recurrent networks, and three-dimensional convolutional networks, were tested on classification benchmarks ranging from 10 to 174 classes.

The baseline evaluations revealed that standard neural networks struggle significantly with fine-grained physical reasoning. On a hand-selected 10-class benchmark of simplified actions, the best combined model achieved an error rate of 44.9% (top-1) and 27.1% (top-2). When the task scaled to 40 classes, the error rate increased to 63.8% (top-1) and 50.7% (top-2). Across all 174 categories, a pre-trained three-dimensional convolutional model recorded an 88.5% top-1 error rate and a 70.3% top-5 error rate. Architectures incorporating three-dimensional convolutions generally outperformed purely two-dimensional image-based models, yet subtle distinctions—such as distinguishing genuine actions from pretended ones—remained difficult across all models.

These results indicate that solving real-world physical reasoning requires models to integrate temporal dynamics rather than rely on static visual features. For organizations building autonomous agents, robotic systems, or language-grounded vision tools, off-the-shelf image models present serious performance risks when physical actions must be accurately interpreted. Achieving reliable visual common sense demands architectures specifically designed to capture fine temporal details, spatial relations, and object state changes.

The authors recommend adopting a curriculum learning strategy, expanding data and model complexity incrementally as validation accuracy improves. Researchers and engineering teams should develop more sophisticated temporal network architectures rather than relying on standard frame-averaging baselines. For long-term development, the authors advocate expanding structured natural language captioning to bridge visual perception and broader language understanding.

The findings are bounded by the baseline nature of the evaluated models and the semantic ambiguities inherent in scaling crowdsourced language labels. However, the data collection methodology and baseline benchmarks provide high confidence that fine-grained video benchmarks are essential for exposing and addressing the current limits of machine common sense.

arXiv: 1706.04261
Cover for The “Something Something” Video Database for Learning and Evaluating Visual Common Sense

Abstract

Neural networks trained on datasets such as ImageNet have led to major advances in visual object classification. One obstacle that prevents networks from reasoning more deeply about complex scenes and situations, and from integrating visual knowledge with natural language, like humans do, is their lack of common sense knowledge about the physical world. Videos, unlike still images, contain a wealth of detailed information about the physical world. However, most labelled video datasets represent high-level concepts rather than detailed physical aspects about actions and scenes. In this work, we describe our ongoing collection of the "something-something" database of video prediction tasks whose solutions require a common sense understanding of the depicted situation. The database currently contains more than 100,000 videos across 174 classes, which are defined as caption-templates. We also describe the challenges in crowd-sourcing this data at scale.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 Learning world models from video
  • 4 The “something-something” dataset
  • 4.1 Crowdsourced video recording
  • 4.2 Natural language labels and curriculum learning
  • 4.3 Non-uniform sampling of the Cartesian product of actions and objects
  • 4.4 Grouping and contrastive examples
  • 4.5 Data collection platform
  • 5 Baseline experiments
  • 5.1 Pre-processing
  • 5.2 Model specifications
  • 5.3 Results
  • 6 Discussion
  • References
  • A Dataset
  • A.1 10-selected classes
  • A.2 40-selected classes
  • A.3 All data - 175 classes
  • A.4 Data Collection Platform

Knowls

  1. Knowl 1 — The Something-Something Video Database

    definition

    The "Something-Something" database is a video dataset designed for learning and evaluating visual common sense and physical reasoning about human-object interactions. It contains 108,499108{,}499 short video clips ranging in duration from 22 to 66 seconds (mean duration of 4.034.03 seconds). The videos depict people performing simple, everyday physical manipulations on objects.

    The dataset defines 174174 fine-grained action classes formulated as natural language caption-templates containing placeholder slots for objects (for example, Dropping [something] into [something] or Moving [something] from right to left). Crowd workers record videos acting out these templates using objects of their choice and submit the specific noun phrases identifying the manipulated items. Across the dataset, crowd workers provided 23,13723{,}137 distinct object names (estimated to cover several thousand distinct physical objects). The dataset was produced by 1,1331{,}133 crowd workers, averaging 127.32127.32 workers per class and approximately 620620 video instances per class (ranging from a minimum of 7777 videos for Poking a hole into [some substance] to a maximum of 986986 videos for Holding [something]).

  2. Knowl 2 — Action Groups and Contrastive Actions for Dataset Debiasing

    model/method

    To prevent deep neural networks from exploiting visual shortcut biases—such as predicting action labels purely from static background context, object appearance in single frames, or coarse hand motion trajectories—action classes in the Something-Something dataset are structured into action groups containing subtle, contrastive distinctions:

    1. Fine-grained directional and spatial variations: Contrasting actions that differ only in spatial relation or direction, such as Putting [something] on top of [something], Putting [something] next to [something], and Putting [something] behind [something].

    2. Graded physical force and outcome: Actions that vary based on the physical result of an interaction, such as Poking [something] so lightly that it does not or almost does not move, Poking [something] so it slightly moves, and Poking [something] so that it falls over.

    3. Genuine versus pretend actions: Classes paired with corresponding deceptive actions, such as Poking [something] versus Pretending to poke something, or Putting [something] behind [something] versus Pretending to put something behind [something] (but not actually leaving it there).

    Distinguishing genuine from pretended manipulations requires a model to verify the physical presence, absence, and displacement of objects across time rather than relying on hand position heuristics alone.

  3. Knowl 3 — Crowd-Acting Video Collection Protocol

    model/method

    Rather than scraping unconstrained web videos and annotating them post-hoc, the "crowd acting" data collection paradigm presents crowd workers with pre-defined action templates containing object slots ([something]) and prompts them to perform and record corresponding video clips. The collection protocol operates as follows:

    1. A dedicated platform presents workers with balanced sets of action templates and action groups to choose from, dynamically tracking and constraining choices to maintain balanced representation across classes and ensure each class is enacted by many distinct workers.
    2. Workers can initiate task batches, gather everyday physical objects, and record videos across different environments, sessions, and lighting conditions.
    3. Upon uploading a recorded video, the placeholder slots in the template dynamically open text input fields where the worker inputs the descriptive noun phrases for the specific objects used.
    4. Uploads undergo automatic quality verification checks (e.g., video duration within [2,6][2, 6] seconds, file integrity, and uniqueness) followed by human operator review and feedback before acceptance and payment dispatch.
  4. Knowl 4 — Worker-Disjoint Dataset Partitioning

    experimental setup

    The Something-Something dataset is divided into training, validation, and test subsets according to an 8:1:18:1:1 ratio (80%80\% train, 10%10\% validation, 10%10\% test). The split is strictly worker-conditional: all videos submitted by any individual crowd worker are assigned exclusively to a single partition (train, validation, or test).

    This partitioning ensures that background environments, personal object collections, visual appearance of hands, and specific recording hardware associated with an individual contributor do not leak across splits, enforcing that models generalize to novel actors, unseen physical settings, and new object appearances.

  5. Knowl 5 — Video Preprocessing and Temporal Downsampling Pipeline

    experimental setup

    Input video clips are preprocessed and temporally augmented prior to model ingestion:

    1. Spatial normalization: Raw video frames are sampled at 24 frames per second (fps)24\text{ frames per second (fps)} and resized to a spatial resolution of 84×8484 \times 84 pixels (unless pre-trained architectures require their native input dimensions).
    2. Temporal anti-aliasing: The sequence of frames is lowpass filtered along the time dimension using a 1D Gaussian filter kernel with mean μ=0\mu = 0 and variance σ2=48 pixels\sigma^2 = 48\text{ pixels}. This temporal filtering attenuates frequencies above the Nyquist frequency corresponding to a target frame rate of 6 fps6\text{ fps}.
    3. Temporal data augmentation: During model training, temporal downsampling by a factor of 44 (yielding 6 fps6\text{ fps}) incorporates a random temporal offset r∈{0,1,2,3}r \in \{0, 1, 2, 3\}. For validation and testing, the temporal offset is fixed to 00.
  6. Knowl 6 — Baseline Video Encoders for Template Classification

    model/method

    Six video encoding architectures are evaluated for the task of classifying action templates from video clips:

    • 2D-CNN + Avg: A 16-layer VGG network trained from scratch on individual video frames. Frame-level feature representations are averaged across all frames to yield the final video descriptor.
    • Pre-2D-CNN + Avg: An ImageNet pre-trained VGG-16 network extracting features per frame, which are then average-pooled across the entire clip.
    • Pre-2D-CNN + LSTM: An ImageNet pre-trained VGG-16 network where frame features are fed sequentially into a Long Short-Term Memory (LSTM) network with a hidden state dimension of 256256. The final hidden state vector serves as the video encoding.
    • 3D-CNN + Stack: A spatiotemporal 3D convolutional network trained from scratch with 10241024-unit fully-connected layers. Videos are padded to a maximum length of 3636 frames, divided into non-overlapping clips of 99 frames, and processed to extract features that are stacked into a 40964096-dimensional vector (44 column blocks) with column masking to ignore padded frames.
    • Pre-3D-CNN + Avg: A 3D-CNN initialized on the Sports-1M dataset and fine-tuned at 8 fps8\text{ fps}. Feature representations are extracted from 1616-frame temporal windows with an 88-frame stride (yielding 55 overlapping columns per clip) and average-pooled across columns.
    • 2D+3D-CNN: An ensemble representation created by concatenating the feature encodings of the best-performing 2D-CNN model and the best-performing 3D-CNN model.
  7. Knowl 7 — Hierarchical Benchmark Subsets of Something-Something Classes

    model/method

    To evaluate baseline model performance across graded levels of task complexity, two hand-selected subsets were constructed alongside the full 174174-class dataset:

    1. 10-Class Benchmark Subset (28,19828{,}198 videos): Formed by selecting 4141 semantically clear, base classes and grouping them into 1010 overarching categories:

      • Dropping [something] (combining rock/feather falls, throwing down, etc.)
      • Moving [something] from right to left (combining pushing/pulling right-to-left)
      • Moving [something] from left to right (combining pushing/pulling left-to-right)
      • Picking [something] up (combining lifting, taking up, etc.)
      • Putting [something] (combining putting on surfaces, next to/behind objects, etc.)
      • Poking [something] (combining light pokes, collapsing/non-collapsing pokes, etc.)
      • Tearing [something] (combining tearing completely and tearing a little bit)
      • Pouring [something] (combining pouring in/out/onto and spilling)
      • Holding [something] (combining holding alone and holding in front of objects)
      • Showing [something] (almost no hand) (combining showing on top, behind, or next to objects)
    2. 40-Class Benchmark Subset (53,26753{,}267 videos): Composed of the 1010 grouped classes above plus 3030 additional common interaction classes (including subtle camera motions, unfastening/unfolding, plugging, and pretending actions).

    3. Full Dataset (108,499108{,}499 videos): The complete set of 174174 fine-grained action template classes.

  8. Knowl 8 — Error Rates for Baseline Action-Template Classification

    data/table

    The classification performance of video encoding baselines on the 10-class, 40-class, and 174-class subsets of the Something-Something dataset is shown below in terms of Top-1, Top-2, and Top-5 error rates (%):

    Method 10 classes 40 classes 174 classes
    top-1 top-2 top-1 top-2 top-1 top-2 top-5
    2D CNN + Avg 76.5 58.9 88.0 78.5 - - -
    Pre-2D CNN + Avg 54.7 39.0 79.2 70.0 - - -
    Pre-2D CNN + LSTM 52.3 34.1 77.8 68.0 - - -
    3D CNN + Stack 58.1 38.7 70.3 57.3 - - -
    Pre-3D CNN + Avg 47.5 29.2 66.2 52.7 88.5 81.5 70.0
    2D+3D-CNN 44.9 27.1 63.8 50.7 - - -

    The results demonstrate that 3D spatiotemporal convolutions consistently outperform 2D convolutional networks (even when 2D networks are pre-trained on ImageNet or paired with an LSTM), with feature concatenation (2D+3D-CNN) yielding the lowest error rates on the 10-class (44.9%44.9\% top-1) and 40-class (63.8%63.8\% top-1) subsets. On the full 174-class dataset, the fine-tuned 3D-CNN achieves an 88.5%88.5\% top-1 and 70.0%70.0\% top-5 error rate, illustrating the severe challenge posed by fine-grained physical and contrastive classes.

  9. Knowl 9 — Dominance of Spatiotemporal Representations over Static Frame Cues

    empirical result

    On the Something-Something benchmark, 2D convolutional architectures that rely on static frame representations and temporal feature pooling fail to distinguish actions with high accuracy. For instance, on the 10-class task, 2D CNN + Avg yields a top-1 error rate of 76.5%76.5\% (reduced to 54.7%54.7\% when using ImageNet pre-training), whereas spatiotemporal 3D models (Pre-3D CNN + Avg) achieve a top-1 error rate of 47.5%47.5\% and the combined 2D+3D model achieves 44.9%44.9\%.

    Because the dataset pairs visually identical objects and static scenes with distinct dynamic actions (such as moving left versus moving right, or real versus simulated interactions), models cannot succeed by simply identifying which object or environment is present; they must explicitly capture motion directionality, velocity, temporal order, and causal physical contact.

Coverage note — None was omitted; all key contributions including dataset design, crowdsourcing protocol, debiasing strategy, preprocessing, baseline models, benchmark subsets, and empirical classification results are represented.

References

  1. 1.P. Agrawal, A. Nair, P. Abbeel, J. Malik, and S. Levine. Learning to poke by poking: Experiential learning of intuitive physics. arXiv preprint arXiv:1606.07419, 2016.
  2. 2.Y. Bengio, J. Louradour, R. Collobert, and J. Weston. Curriculum learning. In ICML 2009, pages 41–48. ACM, 2009.
  3. 3.J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei. ImageNet: A Large-Scale Hierarchical Image Database. In CVPR 2009, 2009.
  4. 4.J. Donahue, Y. Jia, O. Vinyals, J. Hoffman, N. Zhang, E. Tzeng, and T. Darrell. Decaf: A deep convolutional activation feature for generic visual recognition. arXiv preprint arXiv:1310.1531, 2013.
  5. 5.B. G. Fabian Caba Heilbron, Victor Escorcia and J. C. Niebles. Activitynet: A large-scale video benchmark for human activity understanding. In CVPR 2015, pages 961–970, 2015.
  6. 6.A. Geiger, P. Lenz, C. Stiller, and R. Urtasun. Vision meets robotics: The kitti dataset. The International Journal of Robotics Research, 32(11):1231–1237, 2013.
  7. 7.R. Hartley and A. Zisserman. Multiple view geometry in computer vision. Cambridge university press, 2003.
  8. 8.K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. arXiv preprint arXiv:1512.03385, 2015.
  9. 9.D. Hofstadter, D. R. Hofstadter, and E. Sander. Surfaces and Essences. Basic Books, 2013.
  10. 10.Y.-G. Jiang, Z. Wu, J. Wang, X. Xue, and S.-F. Chang. Exploiting feature and class relationships in video categorization with regularized deep neural networks. arXiv preprint arXiv:1502.07209, 2015.
  11. 11.A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei. Large-scale video classification with convolutional neural networks. In CVPR 2014, pages 1725–1732, 2014.
  12. 12.R. Krishna, K. Hata, F. Ren, L. Fei-Fei, and J. C. Niebles. Dense-captioning events in videos. arXiv preprint arXiv:1705.00754, 2017.
  13. 13.A. Krizhevsky, I. Sutskever, and G. E. Hinton. Imagenet classification with deep convolutional neural networks. In Advances in neural information processing systems, pages 1097–1105, 2012.
  14. 14.G. Lakoff and M. Johnson. Metaphors we live by. University of Chicago Press, 1981.
  15. 15.A. Lerer, S. Gross, and R. Fergus. Learning physical intuition of block towers by example. arXiv preprint arXiv:1603.01312, 2016.
  16. 16.H. J. Levesque, E. Davis, and L. Morgenstern. The winograd schema challenge. In AAAI Spring Symposium: Logical Formalizations of Commonsense Reasoning, volume 46, page 47, 2011.
  17. 17.M. Marszalek, I. Laptev, and C. Schmid. Actions in context. In CVPR 2009, pages 2929–2936. IEEE, 2009.
  18. 18.R. Memisevic. Learning to relate images. IEEE transactions on pattern analysis and machine intelligence, 35(8):1829–1846, 2013.
  19. 19.V. Michalski, R. Memisevic, and K. Konda. Modeling deep temporal dependencies with recurrent grammar cells””. In Advances in neural information processing systems, pages 1925–1933, 2014.
  20. 20.A. Nguyen, A. Dosovitskiy, J. Yosinski, T. Brox, and J. Clune. Synthesizing the preferred inputs for neurons in neural networks via deep generator networks. arXiv preprint arXiv:1605.09304, 2016.
  21. 21.L. Pinto, D. Gandhi, Y. Han, Y.-L. Park, and A. Gupta. The curious robot: Learning visual representations via physical interactions. arXiv preprint arXiv:1604.01360, 2016.
  22. 22.M. Ranzato, A. Szlam, J. Bruna, M. Mathieu, R. Collobert, and S. Chopra. Video (language) modeling: a baseline for generative models of natural videos. arXiv preprint arXiv:1412.6604, 2014.
  23. 23.M. Regneri, M. Rohrbach, D. Wetzel, S. Thater, B. Schiele, and M. Pinkal. Grounding action descriptions in videos. Transactions of the Association for Computational Linguistics, 1:25–36, 2013.
  24. 24.M. D. Rodriguez, J. Ahmed, and M. Shah. Action mach a spatio-temporal maximum average correlation height filter for action recognition. In CVPR 2008, pages 1–8. IEEE, 2008.
  25. 25.A. Rohrbach, M. Rohrbach, N. Tandon, and B. Schiele. A dataset for movie description. In CVPR 2015, pages 3202–3212, 2015.
  26. 26.M. Rohrbach, S. Amin, M. Andriluka, and B. Schiele. A database for fine grained activity detection of cooking activities. In Computer Vision and Pattern Recognition (CVPR), 2012 IEEE Conference on, pages 1194–1201. IEEE, 2012.
  27. 27.C. Schuldt, I. Laptev, and B. Caputo. Recognizing human actions: a local svm approach. In ICPR 2004, volume 3, pages 32–36. IEEE, 2004.
  28. 28.A. Sharif Razavian, H. Azizpour, J. Sullivan, and S. Carlsson. Cnn features off-the-shelf: an astounding baseline for recognition. In CVPR Workshops, pages 806–813, 2014.
  29. 29.G. A. Sigurdsson, G. Varol, X. Wang, A. Farhadi, I. Laptev, and A. Gupta. Hollywood in homes: Crowdsourcing data collection for activity understanding. arXiv preprint arXiv:1604.01753, 2016.
  30. 30.K. Simonyan and A. Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014.
  31. 31.C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich. Going deeper with convolutions. In CVPR 2015, pages 1–9, 2015.
  32. 32.A. Torabi, C. Pal, H. Larochelle, and A. Courville. Using descriptive video services to create a large data source for video annotation research. arXiv preprint arXiv:1503.01070, 2015.
  33. 33.A. Torralba and A. A. Efros. Unbiased look at dataset bias. In Computer Vision and Pattern Recognition (CVPR), 2011 IEEE Conference on, pages 1521–1528. IEEE, 2011.
  34. 34.D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri. Learning spatiotemporal features with 3d convolutional networks. In ICCV 2015, pages 4489–4497, 2015.
  35. 35.C. Vondrick, H. Pirsiavash, and A. Torralba. Anticipating the future by watching unlabeled video. arXiv preprint arXiv:1504.08023, 2015.
  36. 36.Wikipedia. Affordance — wikipedia, the free encyclopedia, 2016. [Online; accessed 9-September-2016].
  37. 37.J. Wu. Computational perception of physical object properties. Master’s thesis, Massachusetts Institute of Technology, 2016.
  38. 38.J. Wu, J. J. Lim, H. Zhang, J. B. Tenenbaum, and W. T. Freeman. Physics 101: Learning physical object properties from unlabeled videos. In British Machine Vision Conference, 2016.
  39. 39.Z. Wu, Y. Fu, Y.-G. Jiang, and L. Sigal. Harnessing object and scene semantics for large-scale video understanding. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3112–3121, 2016.
  40. 40.J. Xu, T. Mei, T. Yao, and Y. Rui. Msr-vtt: A large video description dataset for bridging video and language.
  41. 41.M. Yatskar, L. Zettlemoyer, and A. Farhadi. Situation recognition: Visual semantic role labeling for image understanding. In CVPR, 2016.

Citation

MLA
Goyal, R., et al. “The "something Something" Video Database for Learning and Evaluating Visual Common Sense”. arXiv, 2017, http://arxiv.org/abs/1706.04261v2.
APA
Goyal, R., Kahou, S. E., Michalski, V., Materzyńska, J., Westphal, S., Kim, H., Haenel, V., Fruend, I., Yianilos, P., Mueller-Freitag, M., Hoppe, F., Thurau, C., Bax, I., & Memisevic, R. (2017). The "something something" video database for learning and evaluating visual common sense. arXiv. http://arxiv.org/abs/1706.04261v2
Chicago
Goyal, R., S. E. Kahou, V. Michalski, et al. 2017. “The "something Something" Video Database for Learning and Evaluating Visual Common Sense”. arXiv. http://arxiv.org/abs/1706.04261v2.
Harvard
Goyal, R. et al. (2017) “The "something something" video database for learning and evaluating visual common sense”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1706.04261v2.
Vancouver
1. Goyal R, Kahou SE, Michalski V, et al (2017) The "something something" video database for learning and evaluating visual common sense. arXiv

BibTeX

@article{goyal2017the,
  title = {The "something something" video database for learning and evaluating visual common sense},
  author = {Goyal, Raghav and Kahou, Samira Ebrahimi and Michalski, Vincent and Materzyńska, Joanna and Westphal, Susanne and Kim, Heuna and Haenel, Valentin and Fruend, Ingo and Yianilos, Peter and Mueller-Freitag, Moritz and Hoppe, Florian and Thurau, Christian and Bax, Ingo and Memisevic, Roland},
  year = {2017},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1706.04261v2},
  eprint = {1706.04261}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE