CogBench: a large language model walks into a psychology lab

Julian Coda-FornoMarcel BinzJane X. WangEric Schulz

article2024ICML57 citations

Introduces a cognitive psychology benchmark that phenotypes the decision-making behaviors of 40 large language models across ten metrics, revealing how scale, human feedback alignment, and specific prompting strategies directly shape model-based reasoning and risk tendencies.

Listen

Large language models have advanced rapidly and are increasingly deployed across critical industries, yet evaluating their capabilities remains a major challenge. Standard industry benchmarks focus almost exclusively on accuracy and task performance, treating these systems as black boxes. This narrow approach obscures how models actually make decisions and fails to assess underlying cognitive traits such as risk tolerance, exploration, learning styles, and self-awareness.

To address this gap, the article introduces CogBench, an open-access evaluation framework that adapts seven canonical experiments from cognitive psychology into ten behavioral and six performance metrics. The researchers tested 40 commercial and open-source models using prompt-based tasks—such as reward-learning games, multi-step planning scenarios, and risk simulations—without modifying the models through fine-tuning. They applied multilevel statistical modeling to analyze how model architecture, training techniques, and prompting strategies influence both performance and behavioral traits.

The investigation revealed several key insights into artificial decision-making. First, training models with human feedback significantly improved their alignment with human-like behavior (reducing behavioral distance by about 11.7%) and substantially increased metacognition, or the ability to accurately gauge confidence in decisions. Second, while larger parameter counts reliably improved overall task performance and model-based planning, fine-tuning models on computer code or simply expanding training data size did not yield meaningful improvements in these cognitive areas. Third, contrary to popular belief that proprietary systems are more cautious due to safety constraints, open-source models exhibited significantly less risk-taking behavior in simulated risk tasks. Finally, prompt-engineering techniques showed distinct specializations: step-by-step chain-of-thought prompting boosted probabilistic accuracy by roughly 9%, whereas abstract take-a-step-back prompting increased model-based planning behavior by nearly 119%.

These findings have direct operational and strategic implications for organizations developing or deploying artificial intelligence. High benchmark accuracy does not guarantee balanced decision-making; for example, many high-performing models succeed through pure exploitation while completely failing to explore alternative choices or properly balance prior assumptions against new evidence. Understanding these cognitive profiles helps decision-makers mitigate operational risks, prevent extreme risk-taking behaviors, and choose the most effective prompting strategies for complex reasoning workflows.

Organizations should adopt behavioral profiling alongside traditional accuracy metrics when auditing and selecting models for sensitive decision-support roles. However, leaders should interpret these findings with measured confidence due to current limitations. Many proprietary model architectures remain opaque, and psychological constructs originally designed for humans may not translate perfectly to artificial systems. Future work should focus on validating these behavioral metrics in real-world deployments, expanding the variety of cognitive tasks, and standardizing automated behavioral audits.

Cover for CogBench: a large language model walks into a psychology lab

Abstract

Large language models (LLMs) have significantly advanced the field of artificial intelligence. Yet, evaluating them comprehensively remains challenging. We argue that this is partly due to the predominant focus on performance metrics in most benchmarks. This paper introduces CogBench, a benchmark that includes ten behavioral metrics derived from seven cognitive psychology experiments. This novel approach offers a toolkit for phenotyping LLMs’ behavior. We apply CogBench to 40 LLMs, yielding a rich and diverse dataset. We analyze this data using statistical multilevel modeling techniques, accounting for the nested dependencies among fine-tuned versions of specific LLMs. Our study highlights the crucial role of model size and reinforcement learning from human feedback (RLHF) in improving performance and aligning with human behavior. Interestingly, we find that open-source models are less risk-prone than proprietary models and that fine-tuning on code does not necessarily enhance LLMs’ behavior. Finally, we explore the effects of prompt-engineering techniques. We discover that chain-of-thought prompting improves probabilistic reasoning, while take-a-step-back prompting fosters model-based behaviors.

Table of Contents

  • 1. Introduction
  • 2. Related work
  • 3. Methods
  • 3.1. Prompting and summary of included models
  • 3.2. High-level summary of tasks
  • 3.3. Human data
  • 4. The cognitive phenotype of LLMs
  • 4.1. Performance summary
  • 4.2. Differences between behavioral and performance metrics
  • 5. Hypothesis-driven experiments
  • 5.1. Experimental procedure
  • 5.2. Results
  • Hypothesis 1: Does RLHF make LLMs more humanlike?
  • Hypothesis 3: Does an increase of parameters, training data, and the inclusion of code increase modelbasedness?
  • Hypothesis 4: Does RLHF enhance meta-cognition?
  • Hypothesis 5: Do open-source models take more risks?
  • 6. Impact of prompt-engineering
  • 7. Discussion
  • Acknowledgements
  • Impact statement
  • References
  • A. List of LLMs used
  • B. Comprehensive list & explanation of the cognitive experiments
  • B.1. Probabilistic reasoning (Dasgupta et al., 2020) - Prior & likelihood weighting
  • B.1.1. SUMMARY
  • B.1.2. METHODS
  • B.1.3. PROMPTS FOR LLMS
  • B.1.4. METRICS
  • B.2. Horizon task (Wilson et al., 2014) - Directed & random exploration
  • B.2.1. SUMMARY
  • B.2.2. METHODS
  • B.2.3. PROMPTS FOR LLMS
  • B.2.4. METRICS
  • B.3. Restless bandit task (Ershadmanesh et al., 2023) - Meta-cognition
  • B.3.1. SUMMARY
  • B.3.2. METHODS
  • B.3.3. PROMPTS FOR LLMS
  • B.3.4. METRICS
  • B.4. Experiment 2: Instrumental learning(Lefebvre et al., 2017) - Optimism bias & learning rate
  • B.4.1. SUMMARY
  • B.4.2. METHODS
  • B.4.3. PROMPTS FOR LLMS
  • B.4.4. METRICS
  • B.5. Two step task (Daw et al., 2011) - Model-basedness
  • B.5.1. SUMMARY
  • B.5.2. METHODS
  • B.5.3. PROMPTS FOR LLMS
  • B.5.4. METRICS
  • B.6. Temporal discounting (Ruggeri et al., 2022)
  • B.6.1. SUMMARY
  • B.6.2. METHODS
  • B.6.3. PROMPTS FOR LLMS
  • B.6.4. METRICS
  • B.7. Balloon Analogue Risk Task (BART) (Lejuez et al., 2002) - Risk
  • B.7.1. SUMMARY
  • B.7.2. METHODS
  • B.7.3. PROMPTS FOR LLMS
  • B.7.4. METRICS
  • C. Full benchmark results for rest of LLMs
  • D. Robustness analysis
  • E. Redundancy analysis
  • F. Prompt Engineering techniques
  • G. Regression package

Knowls

  1. Knowl 1 — CogBench Cognitive Evaluation Suite

    model/method

    CogBench is an evaluation benchmark for large language models (LLMs) based on seven experimental paradigms from cognitive psychology. It evaluates models along ten behavioral metrics and six task-performance metrics, utilizing in-context learning without fine-tuning (setting sampling temperature to T=0T = 0 for deterministic generation) and prompt-chaining history over trials.

    The benchmark consists of seven tasks:

    1. Probabilistic Reasoning: Measures belief updating from prior probabilities and evidence likelihoods, evaluating posterior accuracy alongside behavioral parameters for prior and likelihood weighting.
    2. Horizon Task: A two-armed bandit with fixed horizons (1 or 6 additional choices) that separates directed exploration from random exploration.
    3. Restless Bandit Task: A two-armed non-stationary bandit where reward distributions switch periodically, evaluating decision accuracy and metacognitive sensitivity (confidence alignment).
    4. Instrumental Learning: Interleaved two-armed bandit problems with symmetric and asymmetric reward contingencies, measuring overall learning rate and asymmetric updating (optimism bias).
    5. Two-Step Task: A two-stage Markov decision process that experimentally disentangles model-based reinforcement learning from model-free reinforcement learning.
    6. Temporal Discounting: Intertemporal choice scenarios across gain, loss, and magnitude conditions to measure preference for immediate versus delayed outcomes.
    7. Balloon Analogue Risk Task (BART): Sequential pump-or-bank choices under risk of balloon explosion, quantifying risk-taking behavior as average inflation attempts.
  2. Knowl 2 — Hierarchical Multilevel Modeling for LLM Phenotyping

    model/method

    To account for nested dependencies among LLMs (such as base models and their fine-tuned or conversational variants), LLM behaviors are analyzed using a two-level hierarchical linear regression model.

    Level 1 (Within-Group Model): Models the relationship between a cognitive metric YijY_{ij} and pp specific LLM features XijkX_{ijk} within each model family group jj:

    Yij=β0j+∑k=1pβkjXijk+ϵijY_{ij} = \beta_{0j} + \sum_{k=1}^p \beta_{kj} X_{ijk} + \epsilon_{ij}

    where YijY_{ij} is the outcome metric for LLM ii in group jj, β0j\beta_{0j} is the group-specific intercept, βkj\beta_{kj} is the slope for feature kk in group jj, XijkX_{ijk} is the kk-th standardized feature predictor (e.g., parameter count, context length, dataset size), and ϵij∼N(0,σ2)\epsilon_{ij} \sim \mathcal{N}(0, \sigma^2) is the residual error.

    Level 2 (Between-Group Model): Models the group-specific intercepts β0j\beta_{0j} and slopes βkj\beta_{kj} as functions of group-level predictors ZjZ_j (such as whether a model is a chat-tuned variant):

    β0j=γ00+γ01Zj+u0j\beta_{0j} = \gamma_{00} + \gamma_{01} Z_j + u_{0j}

    βkj=γk0+γk1Zj+ukj\beta_{kj} = \gamma_{k0} + \gamma_{k1} Z_j + u_{kj}

    where γ00\gamma_{00} and γk0\gamma_{k0} are the fixed average intercept and slope across groups, γ01\gamma_{01} and γk1\gamma_{k1} are the effects of the group-level predictor ZjZ_j, and u0j,ukju_{0j}, u_{kj} are group-level random effects.

  3. Knowl 3 — Log-Odds Estimation of Prior and Likelihood Weightings in Probabilistic Reasoning

    model/method

    In the probabilistic reasoning task, belief updating from prior probabilities and likelihood evidence is modeled using a generalized Bayesian updating formulation with subjective weighting exponents β1\beta_1 (prior weighting) and β2\beta_2 (likelihood weighting):

    P(A∣B)∝P(B∣A)β2⋅P(A)β1P(A \mid B) \propto P(B \mid A)^{\beta_2} \cdot P(A)^{\beta_1}

    For an experiment presenting prior odds between urns FF and JJ given observed ball evidence, the parameters are estimated via least-squares linear regression on the log-odds of the reported subjective probability P(Urn F∣Ball)∈(0,1)P(\text{Urn } F \mid \text{Ball}) \in (0, 1):

    log⁡(P(Urn F∣Ball)1−P(Urn F∣Ball))=β0+β1log⁡(P(Urn F)1−P(Urn F))+β2log⁡(P(Ball∣Urn F)P(Ball∣Urn J))\log\left(\frac{P(\text{Urn } F \mid \text{Ball})}{1 - P(\text{Urn } F \mid \text{Ball})}\right) = \beta_0 + \beta_1 \log\left(\frac{P(\text{Urn } F)}{1 - P(\text{Urn } F)}\right) + \beta_2 \log\left(\frac{P(\text{Ball} \mid \text{Urn } F)}{P(\text{Ball} \mid \text{Urn } J)}\right)

    where β0\beta_0 is the intercept term. Values of β1,β2<1\beta_1, \beta_2 < 1 reflect system neglect (underweighting of priors or evidence), while β1,β2=1\beta_1, \beta_2 = 1 corresponds to optimal Bayesian updating. Posterior accuracy is computed as 1−∣P(Urn F∣Ball)−PBayes∣1 - |P(\text{Urn } F \mid \text{Ball}) - P_{\text{Bayes}}|, where PBayesP_{\text{Bayes}} is the true Bayesian posterior probability.

  4. Knowl 4 — Quantifying Directed and Random Exploration via the Horizon Task

    model/method

    In the Horizon task, an agent completes four forced-choice bandit trials followed by either 1 additional choice (Horizon 1, pure exploitation baseline) or 6 additional choices (Horizon 6, exploration condition). Forced trials provide either unequal information (1 observation from one arm, 3 from the other) or equal information (2 observations from each arm).

    Behavioral exploration strategies are quantified via linear regression on the first free choice using three regressors:

    1. x1x_1: Difference in observed average rewards between options.
    2. x2x_2: Task horizon (00 for Horizon 1, 11 for Horizon 6).
    3. x3=x1×x2x_3 = x_1 \times x_2: Interaction between reward difference and horizon.
    • Directed Exploration: Defined as the regression coefficient for x2x_2 in the unequal information condition, measuring the agent's propensity to choose the less frequently observed option when future choice opportunities exist.
    • Random Exploration: Defined as the regression coefficient for the interaction term x3x_3 in the equal information condition, measuring how much the agent increases choice stochasticity (flattening the policy slope with respect to reward difference) when moving from Horizon 1 to Horizon 6.
  5. Knowl 5 — Metacognitive Sensitivity via Adjusted Quadratic Scoring Rule

    model/method

    In the restless bandit task, an agent chooses between two slot machines with non-stationary reward distributions (N(60,8)N(60, 8) vs. N(40,8)N(40, 8), switching every 18–22 trials) and reports a subjective confidence rating c∈[0,1]c \in [0, 1] after each choice.

    Metacognitive sensitivity is quantified using the adjusted Quadratic Scoring Rule (QSR):

    QSR=1−(accuracy−scaled confidence)2\text{QSR} = 1 - (\text{accuracy} - \text{scaled confidence})^2

    where accuracy∈{0,1}\text{accuracy} \in \{0, 1\} indicates whether the chosen arm was the optimal (higher expected reward) option on that trial, and scaled confidence∈[0,1]\text{scaled confidence} \in [0, 1] normalizes reported confidence relative to the agent's full range of reported values:

    scaled confidence=confidence−min⁡(confidence)max⁡(confidence)−min⁡(confidence)\text{scaled confidence} = \frac{\text{confidence} - \min(\text{confidence})}{\max(\text{confidence}) - \min(\text{confidence})}

    Higher QSR values indicate that the agent accurately calibrates confidence to actual decision correctness.

  6. Knowl 6 — Asymmetric Reinforcement Learning and Optimism Bias Estimation

    model/method

    In the instrumental learning task, agents complete 96 trials across four interleaved two-armed bandits with reward probabilities P∈{0.25,0.75}P \in \{0.25, 0.75\}. Value updating is modeled via Rescorla-Wagner reinforcement learning with Softmax action selection. Expected values update according to prediction errors δt=Rt−Vt(at)\delta_t = R_t - V_t(a_t):

    ΔV(at)=α⋅δt\Delta V(a_t) = \alpha \cdot \delta_t

    where α∈[0,1]\alpha \in [0, 1] is the overall learning rate estimated by minimizing negative log-likelihood of observed actions.

    To measure optimism bias, the asymmetric RW±\text{RW}\pm model fits separate learning rates for positive and negative prediction errors:

    ΔV(at)={α+⋅δtif δt>0α−⋅δtif δt≤0\Delta V(a_t) = \begin{cases} \alpha^+ \cdot \delta_t & \text{if } \delta_t > 0 \\ \alpha^- \cdot \delta_t & \text{if } \delta_t \le 0 \end{cases}

    Optimism bias is defined as the difference:

    Optimism Bias=α+−α−\text{Optimism Bias} = \alpha^+ - \alpha^-

    A positive value reflects a higher rate of belief revision following positive outcomes than negative outcomes.

  7. Knowl 7 — Model-Basedness Quantification in the Two-Step Markov Decision Task

    model/method

    The two-step task evaluates sequential decision-making in a two-stage MDP. A first-stage action transitions to one of two second-stage states with a fixed common transition probability of 70%70\% and a rare transition probability of 30%30\%. Second-stage choices probabilistically deliver rewards with drifting values.

    To disentangle model-based from model-free reinforcement learning, first-stage stay probabilities (the probability of repeating the same first-stage action on trial t+1t+1) are regressed on three regressors:

    1. x1x_1: Reward outcome on trial tt (+1+1 for reward, 00 for no reward).
    2. x2x_2: Transition type on trial tt (+1+1 for common, 00 for rare).
    3. x3=x1×x2x_3 = x_1 \times x_2: Interaction between reward and transition type.

    Model-free agents repeat actions solely based on reward (x1x_1), regardless of transition type. Model-based agents utilize task transition structure, yielding a positive interaction coefficient β3\beta_3 (increasing stay probability after rewarded common transitions or unrewarded rare transitions). The regression slope β3\beta_3 serves as the metric for model-basedness.

  8. Knowl 8 — Effects of Model Scale and RLHF on LLM Cognitive Phenotypes

    empirical result

    Multilevel regressions across 40 LLMs reveal specific relationships between model attributes and cognitive traits:

    • RLHF and Human-Likeness: RLHF significantly reduces the Euclidean (L2L_2-norm) distance between LLM behavioral profiles and average human behavioral vectors by 11.7%11.7\%, increasing human-likeness approximately 2×2\times in UMAP projection space.
    • RLHF and Metacognition: RLHF exhibits a strong positive effect on metacognitive sensitivity (beta=0.461±0.15,z=5.9,p<0.001\\beta = 0.461 \pm 0.15, z = 5.9, p < 0.001).
    • Parameter Scale: Model parameter count positively predicts overall benchmark performance across all seven tasks (beta=0.277±0.39,z=14.1,p<0.001\\beta = 0.277 \pm 0.39, z = 14.1, p < 0.001) and specifically increases model-basedness (beta=0.481±0.22,z=4.2,p<0.001\\beta = 0.481 \pm 0.22, z = 4.2, p < 0.001).
  9. Knowl 9 — Effects of Model Openness and Training Modality on Cognitive Phenotypes

    empirical result

    Multilevel regression analysis across 40 LLMs disconfirms common hypotheses regarding open-source models and code fine-tuning:

    • Risk-Taking: Open-source models take significantly fewer risks on the Balloon Analogue Risk Task compared to proprietary models (β=−0.612±0.11,z=−11.4,p<0.001\beta = -0.612 \pm 0.11, z = -11.4, p < 0.001). Holding other model attributes equal, proprietary models (which typically contain hidden system prompts and safety pre-prompts) take more risks.
    • Code Fine-Tuning and Dataset Scale: Fine-tuning on code datasets and increasing total training dataset volume showed no statistically significant effect on either general cognitive task performance or the emergence of model-based reasoning.
  10. Knowl 10 — Differential Effects of Chain-of-Thought and Step-Back Prompting on Cognitive Reasoning

    empirical result

    Evaluating Chain-of-Thought (CoT) and Take-a-Step-Back (SB) prompting across five models (GPT-4, PaLM-2 text-bison@002, Claude-1, Claude-2, and LLaMA-2-70) demonstrates task-specific behavioral advantages (aggregated using inverse-variance weighting):

    • Probabilistic Reasoning: CoT increases posterior accuracy by +9.01%+9.01\% over unprompted baselines, compared to +3.10%+3.10\% for SB prompting, indicating that step-by-step arithmetic decomposition is more effective for Bayesian belief updating.
    • Model-Basedness: SB prompting increases model-basedness scores on the two-step task by +118.59%+118.59\% over baseline, compared to +64.59%+64.59\% for CoT, demonstrating that abstracting general principles and task dynamics is more effective for model-based planning.
  11. Knowl 11 — Systemic Biases in LLM Cognitive Phenotypes

    empirical result

    Evaluation of established LLMs across CogBench reveals several consistent deviations from human behavior:

    • Exploitation Without Exploration: Most models achieve near- or super-human performance on the Horizon task without engaging in human-like directed or random exploration, solving the explore-exploit trade-off purely via exploitation.
    • Prior Conservatism and Optimism Bias: LLMs consistently place higher weight on priors than on likelihood observations (high prior weighting β1\beta_1, low likelihood weighting β2\beta_2), indicating strong prior biases that resist revision. Furthermore, models exhibit substantial optimism bias (learning rates for positive prediction errors α+\alpha^+ exceed negative prediction error rates α−\alpha^-).
    • Bimodal Risk Distribution: On the BART task, LLMs exhibit extreme variance, tending to either never inflate (zero risk) or inflate repeatedly until the balloon explodes.
  12. Knowl 12 — Non-Redundancy and Dimensionality of Cognitive Psychology Metrics in LLMs

    empirical result

    Principal Component Analysis (PCA) on 40 LLMs demonstrates that capturing 95%95\% of total variance requires all 10 behavioral dimensions (explained variance ratios: [0.21,0.18,0.18,0.11,0.08,0.06,0.06,0.06,0.04,0.03][0.21, 0.18, 0.18, 0.11, 0.08, 0.06, 0.06, 0.06, 0.04, 0.03]) and all 6 performance dimensions ([0.52,0.17,0.13,0.09,0.05,0.04][0.52, 0.17, 0.13, 0.09, 0.05, 0.04]).

    Furthermore, the mean pairwise correlations between metrics within the three bandit tasks (restless bandit, instrumental learning, and horizon task) were indistinguishable from correlations across unrelated tasks for both behavioral metrics (0.00±0.270.00 \pm 0.27 vs. 0.05±0.200.05 \pm 0.20) and performance metrics (0.45±0.180.45 \pm 0.18 vs. 0.41±0.160.41 \pm 0.16), confirming that the behavioral metrics isolate distinct cognitive constructs rather than redundant task formats.

Coverage note — Omitted verbatim prompt strings for sub-questions and prompt variants across all tasks, exact per-model numerical benchmark tables in Appendices C and D, and the Python code listing for calling `statsmodels`.

References

  1. 1.Akata, E., Schulz, L., Coda-Forno, J., Oh, S. J., Bethge, M., and Schulz, E. Playing repeated games with large language models, 2023.
  2. 2.Almazrouei, E., Alobeidli, H., Alshamsi, A., Cappelli, A., Cojocaru, R., Debbah, M., Goffinet, E., Heslow, D., Launay, J., Malartic, Q., Noune, B., Pannier, B., and Penedo, G. Falcon-40B: an open large language model with state-of-the-art performance. 2023.
  3. 3.Anthropic. Claude 2. Blog post, 2023. URL https://www.anthropic.com/news/claude-2. Accessed: 2024-01-19.
  4. 4.Binz, M. and Schulz, E. Using cognitive psychology to understand gpt-3. Proceedings of the National Academy of Sciences, 120(6):e2218523120, 2023.
  5. 5.Binz, M., Alaniz, S., Roskies, A., Aczel, B., Bergstrom, C. T., Allen, C., Schad, D., Wulff, D., West, J. D., Zhang, Q., et al. How should the advent of large language models affect the practice of science? arXiv preprint arXiv:2312.03759, 2023.
  6. 6.Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258, 2021.
  7. 7.Brändle, F., Binz, M., and Schulz, E. Exploration beyond bandits. The drive for knowledge: The science of human information seeking, pp. 147–168, 2021.
  8. 8.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. Advances in neural information processing systems, 33: 1877–1901, 2020.
  9. 9.Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., et al. Sparks of artificial general intelligence: Early experiments with gpt-4. arXiv preprint arXiv:2303.12712, 2023.
  10. 10.Burnell, R., Schellaert, W., Burden, J., Ullman, T. D., Martinez-Plumed, F., Tenenbaum, J. B., Rutar, D., Cheke, L. G., Sohl-Dickstein, J., Mitchell, M., et al. Rethink reporting of evaluation results in ai. Science, 380(6641): 136–138, 2023.
  11. 11.Buschoff, L. M. S., Akata, E., Bethge, M., and Schulz, E. Visual cognition in multimodal large language models, 2024.
  12. 12.Carpenter, J., Sherman, M. T., Kievit, R. A., Seth, A. K., Lau, H., and Fleming, S. M. Domain-general enhancements of metacognitive ability through adaptive training. Journal of Experimental Psychology: General, 148(1):51, 2019.
  13. 13.Cavagnaro, D. R., Aranovich, G. J., McClure, S. M., Pitt, M. A., and Myung, J. I. On the functional form of temporal discounting: An optimized adaptive test. Journal of Risk and Uncertainty, 52:233–254, 2016.
  14. 14.Chen, M., Tworek, J., Jun, H., Yuan, Q., de Oliveira Pinto, H. P., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., Ray, A., Puri, R., Krueger, G., Petrov, M., Khlaaf, H., Sastry, G., Mishkin, P., Chan, B., Gray, S., Ryder, N., Pavlov, M., Power, A., Kaiser, L., Bavarian, M., Winter, C., Tillet, P., Such, F. P., Cummings, D., Plappert, M., Chantzis, F., Barnes, E., Herbert-Voss, A., Guss, W. H., Nichol, A., Paino, A., Tezak, N., Tang, J., Babuschkin, I., Balaji, S., Jain, S., Saunders, W., Hesse, C., Carr, A. N., Leike, J., Achiam, J., Misra, V., Morikawa, E., Radford, A., Knight, M., Brundage, M., Murati, M., Mayer, K., Welinder, P., McGrew, B., Amodei, D., McCandlish, S., Sutskever, I., and Zaremba, W. Evaluating large language models trained on code, 2021.
  15. 15.Chen, Y., Liu, T. X., Shan, Y., and Zhong, S. The emergence of economic rationality of gpt. Proceedings of the National Academy of Sciences, 120(51):e2316205120, 2023. doi: 10.1073/pnas. 2316205120. URL https://www.pnas.org/doi/abs/10.1073/pnas.2316205120.
  16. 16.Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D. Deep reinforcement learning from human preferences. Advances in neural information processing systems, 30, 2017.
  17. 17.Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J. Training verifiers to solve math word problems, 2021.
  18. 18.Coda-Forno, J., Witte, K., Jagadish, A. K., Binz, M., Akata, Z., and Schulz, E. Inducing anxiety in large language models increases exploration and bias. arXiv preprint arXiv:2304.11111, 2023.
  19. 19.Coda-Forno, J., Binz, M., Akata, Z., Botvinick, M., Wang, J., and Schulz, E. Meta-in-context learning in large language models. Advances in Neural Information Processing Systems, 36, 2024.
  20. 20.Collins, K. M., Wong, C., Feng, J., Wei, M., and Tenenbaum, J. B. Structured, flexible, and robust: benchmarking and improving large language models towards more human-like behavior in out-of-distribution reasoning tasks. arXiv preprint arXiv:2205.05718, 2022.
  21. 21.Dasgupta, I., Schulz, E., Tenenbaum, J. B., and Gershman, S. J. A theory of learning to infer. Psychological review, 127(3):412, 2020.
  22. 22.Dasgupta, I., Lampinen, A. K., Chan, S. C., Creswell, A., Kumaran, D., McClelland, J. L., and Hill, F. Language models show human-like content effects on reasoning. arXiv preprint arXiv:2207.07051, 2022.
  23. 23.Daw, N. D., Gershman, S. J., Seymour, B., Dayan, P., and Dolan, R. J. Model-based influences on humans’ choices and striatal prediction errors. Neuron, 69(6):1204–1215, 2011.
  24. 24.Ershadmanesh, S., Gholamzadeh, A., Desender, K., and Dayan, P. Meta-cognitive efficiency in learned value-based choice. In 2023 Conference on Cognitive Computational Neuroscience, pp. 29–32, 2023. doi: 10.32470/CCN.2023.1570-0. URL https://hdl.handle.net/21.11116/0000-000D-5BC7-D.
  25. 25.Fleming, S. M. and Lau, H. C. How to measure metacognition. Frontiers in Human Neuroscience, 8, 2014. ISSN 1662-5161. doi: 10.3389/fnhum.2014.00443. URL https://www.frontiersin.org/articles/10.3389/fnhum.2014.00443.
  26. 26.Gershman, S. J. Deconstructing the human algorithms for exploration. Cognition, 173:34–42, 2018.
  27. 27.Google. Palm 2 technical report. arXiv preprint arXiv:2305.10403, 2023.
  28. 28.Hagendorff, T., Fabi, S., and Kosinski, M. Human-like intuitive behavior and reasoning biases emerged in large language models but disappeared in chatgpt. Nature Computational Science, 3(10):833–838, 2023.
  29. 29.Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J. Measuring massive multitask language understanding, 2021.
  30. 30.Joshi, M., Choi, E., Weld, D. S., and Zettlemoyer, L. Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension, 2017.
  31. 31.Kambhampati, S., Valmeekam, K., Guan, L., Stechly, K., Verma, M., Bhambri, S., Saldyt, L., and Murthy, A. Llms can’t plan, but can help planning in llm-modulo frameworks, 2024.
  32. 32.Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020.
  33. 33.Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., and Iwasawa, Y. Large language models are zero-shot reasoners. Advances in neural information processing systems, 35: 22199–22213, 2022.
  34. 34.Kool, W., Gershman, S. J., and Cushman, F. A. Cost-benefit arbitration between multiple reinforcement-learning systems. Psychological science, 28(9):1321–1333, 2017.
  35. 35.LAION. Towards a transparent ai future: The call for less regulatory hurdles on open-source ai in europe. Available at: https://laion.ai/blog/transparent-ai/, 2024. Accessed: January 19, 2024.
  36. 36.Lampinen, A. K., Dasgupta, I., Chan, S. C., Matthewson, K., Tessler, M. H., Creswell, A., McClelland, J. L., Wang, J. X., and Hill, F. Can language models learn from explanations in context? arXiv preprint arXiv:2204.02329, 2022.
  37. 37.Lefebvre, G., Lebreton, M., Meyniel, F., Bourgeois-Gironde, S., and Palminteri, S. Behavioural and neural characterization of optimistic reinforcement learning. Nature Human Behaviour, 1(4):0067, 2017.
  38. 38.Lejuez, C. W. et al. Evaluation of a behavioral measure of risk taking: the balloon analogue risk task (bart). Journal of experimental psychology. Applied, 8(2):75–84, 2002. doi: 10.1037//1076-898x.8.2.75.
  39. 39.Liu, X., Wang, J., Sun, J., Yuan, X., Dong, G., Di, P., Wang, W., and Wang, D. Prompting frameworks for large language models: A survey, 2023.
  40. 40.Massey, C. and Wu, G. Detecting regime shifts: The causes of under-and overreaction. Management Science, 51(6): 932–947, 2005.
  41. 41.McCoy, R. T., Yao, S., Friedman, D., Hardy, M., and Griffiths, T. L. Embers of autoregression: Understanding large language models through the problem they are trained to solve. arXiv preprint arXiv:2309.13638, 2023.
  42. 42.McCurdy, L. Y., Maniscalco, B., Metcalfe, J., Liu, K. Y., De Lange, F. P., and Lau, H. Anatomical coupling between distinct metacognitive systems for memory and visual perception. Journal of Neuroscience, 33(5):1897–1906, 2013.
  43. 43.McInnes, L., Healy, J., and Melville, J. Umap: Uniform manifold approximation and projection for dimension reduction, 2020.
  44. 44.Montague, P. R., Dolan, R. J., Friston, K. J., and Dayan, P. Computational psychiatry. Trends in cognitive sciences, 16(1):72–80, 2012.
  45. 45.MosaicML. Introducing mpt-30b: Raising the bar for open-source foundation models. Blog post, 2023. URL www.mosaicml.com/blog/mpt-30b. Accessed: 2023-06-22.
  46. 46.Nardo, C. The waluigi effect (mega-post). Available at: https://www.lesswrong.com/posts/D7PumeYTDPfBTp3i7/the-waluigi-effect-mega-post, 2024. Accessed: January 19, 2024.
  47. 47.OpenAI. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023.
  48. 48.Ouyang, S., Zhang, J. M., Harman, M., and Wang, M. Llm is like a box of chocolates: the non-determinism of chatgpt in code generation. arXiv preprint arXiv:2308.02828, 2023.
  49. 49.Palminteri, S. and Lebreton, M. The computational roots of positivity and confirmation biases in reinforcement learning. Trends in Cognitive Sciences, 2022.
  50. 50.Patzelt, E. H., Hartley, C. A., and Gershman, S. J. Computational phenotyping: using models to understand individual differences in personality, development, and mental illness. Personality Neuroscience, 1:e18, 2018.
  51. 51.Rescorla, R. A. Classical conditioning ii: current research and theory. pp. 64, 1972.
  52. 52.Ruggeri, K., Panin, A., Vdovic, M., Većkalov, B., Abdul-Salaam, N., Achterberg, J., Akil, C., Amatya, J., Amatya, K., Andersen, T. L., et al. The globalizability of temporal discounting. Nature Human Behaviour, 6(10):1386–1397, 2022.
  53. 53.Salewski, L., Alaniz, S., Rio-Torto, I., Schulz, E., and Akata, Z. In-context impersonation reveals large language models’ strengths and biases. arXiv preprint arXiv:2305.14930, 2023.
  54. 54.Schaeffer, R., Miranda, B., and Koyejo, S. Are emergent abilities of large language models a mirage? arXiv preprint arXiv:2304.15004, 2023.
  55. 55.Schurr, R., Reznik, D., Hillman, H., Bhui, R., and Gershman, S. J. Dynamic computational phenotyping of human cognition. 2023.
  56. 56.Shekhar, M. and Rahnev, D. Sources of metacognitive inefficiency. Trends in Cognitive Sciences, 25(1):12–23, 2021.
  57. 57.Sprague, Z., Ye, X., Bostrom, K., Chaudhuri, S., and Durrett, G. Musr: Testing the limits of chain-of-thought with multistep soft reasoning, 2024.
  58. 58.Srivastava, A. and authors. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models, 2023.
  59. 59.Tamkin, A., Brundage, M., Clark, J., and Ganguli, D. Understanding the capabilities, limitations, and societal impact of large language models. arXiv preprint arXiv:2102.02503, 2021.
  60. 60.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023.
  61. 61.Ullman, T. Large language models fail on trivial alterations to theory-of-mind tasks. arXiv preprint arXiv:2302.08399, 2023.
  62. 62.Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., et al. Emergent abilities of large language models. arXiv preprint arXiv:2206.07682, 2022.
  63. 63.Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q., and Zhou, D. Chain-of-thought prompting elicits reasoning in large language models, 2023.
  64. 64.Wilson, R. C., Geana, A., White, J. M., Ludvig, E. A., and Cohen, J. D. Humans use directed and random exploration to solve the explore–exploit dilemma. Journal of Experimental Psychology: General, 143(6):2074, 2014.
  65. 65.Yax, N., Anllo, H., and Palminteri, S. Studying and improving reasoning in humans and machines. arXiv preprint arXiv:2309.12485, 2023.
  66. 66.Zheng, H. S., Mishra, S., Chen, X., Cheng, H.-T., Chi, E. H., Le, Q. V., and Zhou, D. Take a step back: Evoking reasoning via abstraction in large language models, 2023a.
  67. 67.Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., and Stoica, I. Judging llm-as-a-judge with mt-bench and chatbot arena. arXiv preprint arXiv:2306.05685, 2023b.

Citation

MLA
Coda-Forno, J., et al. “CogBench: A Large Language Model Walks into a Psychology Lab”. arXiv, 2024, http://arxiv.org/abs/2402.18225v1.
APA
Coda-Forno, J., Binz, M., Wang, J. X., & Schulz, E. (2024). CogBench: a large language model walks into a psychology lab. arXiv. http://arxiv.org/abs/2402.18225v1
Chicago
Coda-Forno, J., M. Binz, J. X. Wang, and E. Schulz. 2024. “CogBench: A Large Language Model Walks into a Psychology Lab”. arXiv. http://arxiv.org/abs/2402.18225v1.
Harvard
Coda-Forno, J. et al. (2024) “CogBench: a large language model walks into a psychology lab”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.18225v1.
Vancouver
1. Coda-Forno J, Binz M, Wang JX, Schulz E (2024) CogBench: a large language model walks into a psychology lab. arXiv

BibTeX

@article{codaforno2024cogbench,
  title = {CogBench: a large language model walks into a psychology lab},
  author = {Coda-Forno, Julian and Binz, Marcel and Wang, Jane X. and Schulz, Eric},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.18225v1},
  eprint = {2402.18225}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/