When Agents Say One Thing and Do Another: Validating Elicited Beliefs from LLMs
Khurram YaminJingjing TangSantiago Cortes-GomezAmit SharmaEric HorvitzBryan Wilder
Develops a decision-theoretic framework to verify whether large language models make choices consistent with their stated probabilistic beliefs, establishing testable conditions to audit agent rationality without assuming an underlying utility function.
As artificial intelligence systems are increasingly deployed to assist in high-stakes domains like clinical medicine, decision-makers often rely on the probability estimates reported by large language models to understand the rationale behind their recommendations. However, it remains fundamentally unclear whether the stated probabilities of a model truly reflect the internal beliefs driving its choices, or if they represent disconnected text outputs. The article evaluates whether elicited probabilistic beliefs from language models can be formally validated as genuine decision-guiding subjective probabilities, proposing a mathematically grounded, black-box framework to test consistency between stated beliefs and actions.
The researchers developed a decision-theoretic framework based on perturbed utility maximization, which accommodates stochastic choices and unobserved preferences without requiring assumptions about a model's specific utility function. Under this setup, valid beliefs must satisfy two testable conditions: conditional independence (the belief acts as a sufficient statistic, meaning the chosen action provides no additional predictive information about the true outcome once the belief is known) and cyclic monotonicity (actions systematically shift toward higher-payoff options as reported probabilities rise). The framework was evaluated across four medical diagnosis domains—electrocardiogram-based structural heart disease, survey-based diabetes indicators, and two pediatric Bayesian networks for fever and infant crying—using four frontier and open-source models: GPT-5 (High Reasoning and Minimal Reasoning), DeepSeek-R1, and Llama-4 Scout.
The evaluation produced four key findings. First, all tested models formally failed the conditional independence test across every dataset, demonstrating that models consistently reveal more predictive information in their actions than they verbalize in their stated beliefs. Second, this information gap varied substantially across models and domains: for example, incorporating the model's action reduced predictive error for structural heart disease by an average of 16.33% and for the Llama model by an average of 15.19%, whereas higher-reasoning frontier models showed much smaller discrepancies (averaging around 5% error reduction). Third, post-hoc probability calibration techniques, such as isotonic regression, failed to resolve this issue and frequently worsened residual dependence, showing that the problem stems from structural representational mismatches rather than simple probability miscalibration. Fourth, most models maintained reasonable behavioral consistency, with choices remaining largely monotone relative to stated risk levels, even while failing separate probabilistic coherence checks like the law of iterated expectation.
These findings have immediate practical implications for risk, safety, and governance in automated decision systems. Stated probabilities from language models cannot be accepted as complete or faithful explanations of why an automated system took a specific action. Relying naively on verbalized confidence for triage, auditing, or compliance creates safety risks because an agent may possess and act upon critical diagnostic information that remains uncommunicated to human supervisors. Nevertheless, because higher-performing reasoning models exhibit relatively small belief-action gaps, verbalized probabilities remain practically useful approximations for interpreting model behavior if proper validation is applied.
Organizations deploying decision-support agents should not treat verbalized confidence as a direct substitute for rigorous behavioral auditing. Decision-makers should implement task- and model-specific black-box validation suites that test both probability reports and actual choice policies against ground truth. Furthermore, model developers should prioritize architectural and training methods that align internal representations with externalized explanations rather than relying on standard post-hoc calibration.
The findings are subject to several boundary conditions, as the empirical validation was restricted to stylized binary diagnostic settings, specific prompting structures, and four clinical domains. Additionally, black-box testing cannot rule out strategic misreporting or assess internal model representations directly. Despite these scope limitations, the mathematical proofs and robust bootstrap confidence intervals provide high confidence that stated beliefs alone are currently insufficient to guarantee faithful decision explanations.
- Paper: Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting, Miles Turpin et al. (2023). It provides foundational evidence that verbalized reasoning in language models can systematically diverge from true internal decision drivers, motivating the need for decision-theoretic belief validation.
- Paper: Language Models (Mostly) Know What They Know, Saurav Kadavath et al. (2022). It establishes key baselines for how language models express self-knowledge and probabilistic calibration, which the source formally interrogates against revealed choices.
- Paper: Discovering Latent Knowledge in Language Models Without Supervision, Collin Burns et al. (2023). It introduces unsupervised methods for recovering latent beliefs directly from model representations, establishing an essential reference point for comparing internal beliefs against verbalized judgments.
- Paper: Evaluating the World Model Implicit in a Generative Model, Keyon Vafa et al. (2024). It formalizes theoretical criteria for testing the structural coherence of implicit representations in generative models, providing conceptual precedent for testing behavioral consistency.
- Paper: Faithfulness Tests for Natural Language Explanations, Pepa Atanasova et al. (2023). It develops diagnostic tests for whether model-generated explanations faithfully reflect internal decision processes, offering essential grounding on faithfulness evaluation.
- Paper: Say What You Mean! Large Language Models Speak Too Positively about Negative Commonsense Knowledge, Jiangjie Chen et al. (2023). It highlights empirical discrepancies between language models' recognized knowledge and their explicit statements, illustrating the core challenge of validating elicited outputs.
- Paper: Metacognition in LLMs: Foundations, Progress, and Opportunities, Gabrielle Kaili-May Liu et al. (2026). It synthesizes broader frameworks and benchmarks for metacognition and confidence elicitation, extending the source's findings on belief-action consistency to the wider landscape of self-monitoring in LLMs.
- Paper: Reasoning Models Generate Societies of Thought, Junsol Kim et al. (2026). It analyzes multi-perspective internal deliberations in reasoning models, providing a concrete setting to apply decision-theoretic tests of belief and action coherence.
