Building Machines that Learn and Think Like People

Josh Tenenbaum

article2018AAMAS2,261 citations

Presents a framework that integrates probabilistic programming, program synthesis, and physics simulation engines with deep learning to reverse-engineer human commonsense reasoning and build artificial intelligence capable of learning and planning like young children.

Listen

Modern artificial intelligence has achieved major milestones primarily through sophisticated pattern recognition and data-intensive techniques like deep neural networks. Despite these successes, existing systems lack the general-purpose, flexible commonsense reasoning naturally exhibited even by one-year-old human infants. Current technologies struggle to model the physical and social dynamics of the real world, limiting their adaptability, problem-solving, and general reasoning capabilities.

The article outlines an approach to bridging this capability gap by modeling and reverse-engineering the foundational learning and thinking abilities humans demonstrate from early childhood. Rather than relying exclusively on recognizing patterns in large datasets, the described framework combines probabilistic programming, automated program generation, modern deep learning, and video game simulation engines to construct internal models of physical and social environments.

Three core insights emerge from this framework. First, human commonsense fundamentally depends on internal world models that allow individuals to explain observations, imagine unobserved scenarios, and plan actions to achieve specific goals. Second, combining probabilistic reasoning and programmatic simulations with deep learning provides a more versatile architecture than deep learning alone. Third, reverse-engineering infant cognitive development offers a clear, structured roadmap for engineering more adaptable, robust artificial intelligence systems.

These findings suggest that achieving safer, more dependable, and more resource-efficient autonomous systems requires shifting beyond purely data-driven pattern matching toward model-based reasoning. Adopting this integrated framework can reduce system brittleness and improve planning across dynamic environments. Organizations developing advanced machine learning applications should consider investing in hybrid architectures that incorporate probabilistic modeling and simulation alongside standard deep learning methods. Moving forward, continued research and empirical benchmarking are required to validate how effectively these combined cognitive toolkits scale to complex, industrial-grade operational challenges.

arXiv: 1604.00289
  • Paper: Intelligence Without Reason, Rodney A. Brooks (1991). This seminal work establishes the foundational critique against relying solely on abstract, disembodied reasoning, providing the historical counterpoint that shapes modern debates on model-based versus reactive AI.
  • Paper: The “Something Something” Video Database for Learning and Evaluating Visual Common Sense, Raghav Goyal et al. (2017). This paper introduces a critical benchmark for intuitive physics and visual common sense in video, directly highlighting the empirical failure of purely discriminative deep models that the source paper seeks to overcome.
  • Paper: End-To-End Memory Networks, Sainbayar Sukhbaatar et al. (2015). This paper presents early architectural innovations for integrating explicit structured memory into neural networks, which directly precedes hybrid cognitive-inspired AI architectures.
Cover for Building Machines that Learn and Think Like People

Abstract

Recent successes in artificial intelligence and machine learning have been largely driven by methods for sophisticated pattern recognition, including deep neural networks and other data-intensive methods. But human intelligence is more than just pattern recognition. And no machine system yet built has anything like the flexible, general-purpose commonsense grasp of the world that we can see in even a one-year-old human infant. I will consider how we might capture the basic learning and thinking abilities humans possess from early childhood, as one route to building more human-like forms of machine learning and thinking.

At the heart of human common sense is our ability to model the physical and social environment around us: to explain and understand what we see, to imagine things we could see but haven’t yet, to solve problems and plan actions to make these things real, and to build new models as we learn more about the world. I will focus on our recent work reverse-engineering these capacities using methods from probabilistic programming, program induction and program synthesis, which together with deep learning methods and video game simulation engines, provide a toolkit for the joint enterprise of modeling human intelligence and making AI systems smarter in more human-like ways.

Table of Contents

  • Short Bio

Knowls

  1. Knowl 1 — Integrated Program Induction and Simulation Framework for Common Sense

    model/method

    To reverse-engineer human-like common sense and developmental learning, a unified computational framework integrates the following components:

    • Probabilistic programming, program induction, and program synthesis: Used to represent, infer, and construct explicit causal and symbolic models of the world.
    • Video game simulation engines: Used as internal generative physics and agent models to facilitate counterfactual imagination, forward simulation, and action planning.
    • Deep learning methods: Used for pattern recognition and processing high-dimensional perceptual inputs.

    This hybrid architecture aims to capture the core cognitive abilities evident from early human childhood: explaining perceptual observations, simulating potential outcomes, planning goal-directed actions, and learning structured domain theories.

  2. Knowl 2 — Generative Modeling Requirements for Commonsense Cognition

    assumption

    Human commonsense intelligence cannot be fully captured by data-intensive pattern recognition alone (such as standard deep neural network architectures). Human-like intelligence requires maintaining internal causal models of physical and social environments that support four foundational cognitive functions:

    1. Explanation: Inferring the hidden physical or mental causes underlying observed sensory data.
    2. Imagination: Simulating potential, unobserved, or counterfactual states of the environment.
    3. Planning and Problem Solving: Devising sequences of actions to achieve desired goal states based on forward model predictions.
    4. Model Building: Synthesizing and refining new structural models of domains during learning.

Coverage note — None omitted; the document is a one-page keynote presentation abstract presenting a high-level conceptual framework and research agenda.

Citation

MLA
Lake, B. M., et al. “Building Machines That Learn and Think Like People”. arXiv, 2016, http://arxiv.org/abs/1604.00289v3.
APA
Lake, B. M., Ullman, T. D., Tenenbaum, J. B., & Gershman, S. J. (2016). Building Machines That Learn and Think Like People. arXiv. http://arxiv.org/abs/1604.00289v3
Chicago
Lake, B. M., T. D. Ullman, J. B. Tenenbaum, and S. J. Gershman. 2016. “Building Machines That Learn and Think Like People”. arXiv. http://arxiv.org/abs/1604.00289v3.
Harvard
Lake, B.M. et al. (2016) “Building Machines That Learn and Think Like People”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1604.00289v3.
Vancouver
1. Lake BM, Ullman TD, Tenenbaum JB, Gershman SJ (2016) Building Machines That Learn and Think Like People. arXiv

BibTeX

@article{lake2016building,
  title = {Building Machines That Learn and Think Like People},
  author = {Lake, Brenden M. and Ullman, Tomer D. and Tenenbaum, Joshua B. and Gershman, Samuel J.},
  year = {2016},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1604.00289v3},
  eprint = {1604.00289}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF