CogBench: a large language model walks into a psychology lab
Julian Coda-FornoMarcel BinzJane X. WangEric Schulz
Introduces a cognitive psychology benchmark that phenotypes the decision-making behaviors of 40 large language models across ten metrics, revealing how scale, human feedback alignment, and specific prompting strategies directly shape model-based reasoning and risk tendencies.
Large language models have advanced rapidly and are increasingly deployed across critical industries, yet evaluating their capabilities remains a major challenge. Standard industry benchmarks focus almost exclusively on accuracy and task performance, treating these systems as black boxes. This narrow approach obscures how models actually make decisions and fails to assess underlying cognitive traits such as risk tolerance, exploration, learning styles, and self-awareness.
To address this gap, the article introduces CogBench, an open-access evaluation framework that adapts seven canonical experiments from cognitive psychology into ten behavioral and six performance metrics. The researchers tested 40 commercial and open-source models using prompt-based tasks—such as reward-learning games, multi-step planning scenarios, and risk simulations—without modifying the models through fine-tuning. They applied multilevel statistical modeling to analyze how model architecture, training techniques, and prompting strategies influence both performance and behavioral traits.
The investigation revealed several key insights into artificial decision-making. First, training models with human feedback significantly improved their alignment with human-like behavior (reducing behavioral distance by about 11.7%) and substantially increased metacognition, or the ability to accurately gauge confidence in decisions. Second, while larger parameter counts reliably improved overall task performance and model-based planning, fine-tuning models on computer code or simply expanding training data size did not yield meaningful improvements in these cognitive areas. Third, contrary to popular belief that proprietary systems are more cautious due to safety constraints, open-source models exhibited significantly less risk-taking behavior in simulated risk tasks. Finally, prompt-engineering techniques showed distinct specializations: step-by-step chain-of-thought prompting boosted probabilistic accuracy by roughly 9%, whereas abstract take-a-step-back prompting increased model-based planning behavior by nearly 119%.
These findings have direct operational and strategic implications for organizations developing or deploying artificial intelligence. High benchmark accuracy does not guarantee balanced decision-making; for example, many high-performing models succeed through pure exploitation while completely failing to explore alternative choices or properly balance prior assumptions against new evidence. Understanding these cognitive profiles helps decision-makers mitigate operational risks, prevent extreme risk-taking behaviors, and choose the most effective prompting strategies for complex reasoning workflows.
Organizations should adopt behavioral profiling alongside traditional accuracy metrics when auditing and selecting models for sensitive decision-support roles. However, leaders should interpret these findings with measured confidence due to current limitations. Many proprietary model architectures remain opaque, and psychological constructs originally designed for humans may not translate perfectly to artificial systems. Future work should focus on validating these behavioral metrics in real-world deployments, expanding the variety of cognitive tasks, and standardizing automated behavioral audits.
- Paper: Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models, BIG-bench authors (2022). Establishes foundational multi-task capability and human-comparison benchmarking methodologies across model scales that CogBench adapts into a cognitive psychology framework.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). Introduces chain-of-thought prompting for eliciting step-by-step reasoning, a key prompt-engineering intervention whose effects on probabilistic reasoning are directly evaluated in CogBench.
- Paper: Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them, Mirac Suzgun et al. (2022). Demonstrates the efficacy of chain-of-thought prompting on challenging reasoning tasks, providing foundational context for assessing prompt-dependent cognitive behaviors in LLMs.
- Paper: Understanding Social Reasoning in Language Models with Language Models, Kanishk Gandhi et al. (2023). Provides a paradigm for testing cognitive and social reasoning faculties like Theory of Mind in LLMs using controlled experimental scenarios.
- Paper: Llama 2: Open Foundation and Fine-Tuned Chat Models, Hugo Touvron et al. (2023). Details the application of reinforcement learning from human feedback (RLHF) and scale in open-source chat models, two key variables whose cognitive alignment effects are systematically measured in CogBench.
- Paper: Beyond Accuracy: Behavioral Testing of NLP Models with CheckList, Marco Túlio Ribeiro et al. (2020). Pioneers behavioral, capability-based testing beyond traditional accuracy metrics in NLP, motivating CogBench's behavioral phenotyping approach.
- Paper: Metacognition in LLMs: Foundations, Progress, and Opportunities, Gabrielle Kaili-May Liu et al. (2026). Extends the cognitive psychology perspective on LLMs by systematically reviewing metacognitive monitoring, confidence calibration, and self-assessment capabilities.
- Paper: Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Chain-of-Thought, Violet Xiang et al. (2025). Builds upon cognitive models of deliberate thought to formalize and train System 2 reasoning processes in LLMs via meta chain-of-thought search.
- Paper: Reasoning Models Generate Societies of Thought, Junsol Kim et al. (2026). Investigates the cognitive and conversational structures underlying reasoning models by analyzing internal deliberation traces and perspective shifts.
- Paper: Is this the real life? Is this just fantasy? The Misleading Success of Simulating Social Interactions With LLMs, Xuhui Zhou et al. (2024). Critically examines the limits and ecological validity of simulating complex human-like social interactions and behaviors using LLM agents.
- Paper: Dynamic Evaluation of Large Language Models by Meta Probing Agents, Kaijie Zhu et al. (2024). Applies psychometric principles to design dynamic probing protocols that assess model capabilities across varied evaluation contexts.
- Paper: The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models, Seungone Kim et al. (2025). Expands fine-grained, multi-dimensional capability evaluation across complex cognitive dimensions using model-based evaluators.
