Building Machines that Learn and Think Like People
Josh Tenenbaum
Presents a framework that integrates probabilistic programming, program synthesis, and physics simulation engines with deep learning to reverse-engineer human commonsense reasoning and build artificial intelligence capable of learning and planning like young children.
Modern artificial intelligence has achieved major milestones primarily through sophisticated pattern recognition and data-intensive techniques like deep neural networks. Despite these successes, existing systems lack the general-purpose, flexible commonsense reasoning naturally exhibited even by one-year-old human infants. Current technologies struggle to model the physical and social dynamics of the real world, limiting their adaptability, problem-solving, and general reasoning capabilities.
The article outlines an approach to bridging this capability gap by modeling and reverse-engineering the foundational learning and thinking abilities humans demonstrate from early childhood. Rather than relying exclusively on recognizing patterns in large datasets, the described framework combines probabilistic programming, automated program generation, modern deep learning, and video game simulation engines to construct internal models of physical and social environments.
Three core insights emerge from this framework. First, human commonsense fundamentally depends on internal world models that allow individuals to explain observations, imagine unobserved scenarios, and plan actions to achieve specific goals. Second, combining probabilistic reasoning and programmatic simulations with deep learning provides a more versatile architecture than deep learning alone. Third, reverse-engineering infant cognitive development offers a clear, structured roadmap for engineering more adaptable, robust artificial intelligence systems.
These findings suggest that achieving safer, more dependable, and more resource-efficient autonomous systems requires shifting beyond purely data-driven pattern matching toward model-based reasoning. Adopting this integrated framework can reduce system brittleness and improve planning across dynamic environments. Organizations developing advanced machine learning applications should consider investing in hybrid architectures that incorporate probabilistic modeling and simulation alongside standard deep learning methods. Moving forward, continued research and empirical benchmarking are required to validate how effectively these combined cognitive toolkits scale to complex, industrial-grade operational challenges.
- Paper: Intelligence Without Reason, Rodney A. Brooks (1991). This seminal work establishes the foundational critique against relying solely on abstract, disembodied reasoning, providing the historical counterpoint that shapes modern debates on model-based versus reactive AI.
- Paper: The “Something Something” Video Database for Learning and Evaluating Visual Common Sense, Raghav Goyal et al. (2017). This paper introduces a critical benchmark for intuitive physics and visual common sense in video, directly highlighting the empirical failure of purely discriminative deep models that the source paper seeks to overcome.
- Paper: End-To-End Memory Networks, Sainbayar Sukhbaatar et al. (2015). This paper presents early architectural innovations for integrating explicit structured memory into neural networks, which directly precedes hybrid cognitive-inspired AI architectures.
- Paper: Relational inductive biases, deep learning, and graph networks, Peter W. Battaglia et al. (2018). This paper formalizes the relational inductive biases and graph networks necessary to implement the structured, entity-centric internal world models advocated by the source.
- Paper: Recurrent World Models Facilitate Policy Evolution, David Ha et al. (2018). This paper provides an explicit computational realization of internal simulation by training agents entirely inside recurrent predictive world models.
- Paper: Dream to Control: Learning Behaviors by Latent Imagination, Danijar Hafner et al. (2019). This work extends model-based reinforcement learning by optimizing policies purely through latent imagination within learned world models.
- Paper: Shortcut learning in deep neural networks, Robert Geirhos et al. (2020). This perspective analyzes the phenomenon of shortcut learning, diagnosing the precise systemic brittleness of purely pattern-matching networks that the source paper aims to solve.
- Paper: Zero-shot World Models Are Developmentally Efficient Learners, Khai Loong Aw et al. (2026). This study puts the source's developmental roadmap into practice by training zero-shot world models on infant egocentric video streams to bootstrap intuitive visual learning.
- Paper: Generative Agents: Interactive Simulacra of Human Behavior, Joon Sung Park et al. (2023). This paper applies the source's vision of generative, simulation-driven agents to model believable human behaviors and emergent social interactions in virtual environments.
- Paper: Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents, Wenlong Huang et al. (2022). This work operationalizes intuitive common-sense knowledge for planning by demonstrating how language models can generate actionable step-by-step goals for embodied agents.
- Paper: PaLM-E: An Embodied Multimodal Language Model, Danny Driess et al. (2023). This paper extends multimodal models into embodied physical interaction by grounding high-level reasoning directly in real-world sensor streams and robotic control.
- Paper: WorldSimBench: Towards Video Generation Models as World Simulators, Yiran Qin et al. (2025). This benchmark evaluates whether modern video generation models can act as actionable world simulators of physical interactions as envisioned by model-based AI.
- Paper: Efficient Rectification of Neuro-Symbolic Reasoning Inconsistencies by Abductive Reflection, Wen-Chao Hu et al. (2025). This research advances hybrid neuro-symbolic architectures by introducing abductive reflection to efficiently reconcile neural perceptions with symbolic logic.
