When Agents Say One Thing and Do Another: Validating Elicited Beliefs from LLMs

Khurram YaminJingjing TangSantiago Cortes-GomezAmit SharmaEric HorvitzBryan Wilder

article2026arXiv6 citations

Develops a decision-theoretic framework to verify whether large language models make choices consistent with their stated probabilistic beliefs, establishing testable conditions to audit agent rationality without assuming an underlying utility function.

Listen

As artificial intelligence systems are increasingly deployed to assist in high-stakes domains like clinical medicine, decision-makers often rely on the probability estimates reported by large language models to understand the rationale behind their recommendations. However, it remains fundamentally unclear whether the stated probabilities of a model truly reflect the internal beliefs driving its choices, or if they represent disconnected text outputs. The article evaluates whether elicited probabilistic beliefs from language models can be formally validated as genuine decision-guiding subjective probabilities, proposing a mathematically grounded, black-box framework to test consistency between stated beliefs and actions.

The researchers developed a decision-theoretic framework based on perturbed utility maximization, which accommodates stochastic choices and unobserved preferences without requiring assumptions about a model's specific utility function. Under this setup, valid beliefs must satisfy two testable conditions: conditional independence (the belief acts as a sufficient statistic, meaning the chosen action provides no additional predictive information about the true outcome once the belief is known) and cyclic monotonicity (actions systematically shift toward higher-payoff options as reported probabilities rise). The framework was evaluated across four medical diagnosis domains—electrocardiogram-based structural heart disease, survey-based diabetes indicators, and two pediatric Bayesian networks for fever and infant crying—using four frontier and open-source models: GPT-5 (High Reasoning and Minimal Reasoning), DeepSeek-R1, and Llama-4 Scout.

The evaluation produced four key findings. First, all tested models formally failed the conditional independence test across every dataset, demonstrating that models consistently reveal more predictive information in their actions than they verbalize in their stated beliefs. Second, this information gap varied substantially across models and domains: for example, incorporating the model's action reduced predictive error for structural heart disease by an average of 16.33% and for the Llama model by an average of 15.19%, whereas higher-reasoning frontier models showed much smaller discrepancies (averaging around 5% error reduction). Third, post-hoc probability calibration techniques, such as isotonic regression, failed to resolve this issue and frequently worsened residual dependence, showing that the problem stems from structural representational mismatches rather than simple probability miscalibration. Fourth, most models maintained reasonable behavioral consistency, with choices remaining largely monotone relative to stated risk levels, even while failing separate probabilistic coherence checks like the law of iterated expectation.

These findings have immediate practical implications for risk, safety, and governance in automated decision systems. Stated probabilities from language models cannot be accepted as complete or faithful explanations of why an automated system took a specific action. Relying naively on verbalized confidence for triage, auditing, or compliance creates safety risks because an agent may possess and act upon critical diagnostic information that remains uncommunicated to human supervisors. Nevertheless, because higher-performing reasoning models exhibit relatively small belief-action gaps, verbalized probabilities remain practically useful approximations for interpreting model behavior if proper validation is applied.

Organizations deploying decision-support agents should not treat verbalized confidence as a direct substitute for rigorous behavioral auditing. Decision-makers should implement task- and model-specific black-box validation suites that test both probability reports and actual choice policies against ground truth. Furthermore, model developers should prioritize architectural and training methods that align internal representations with externalized explanations rather than relying on standard post-hoc calibration.

The findings are subject to several boundary conditions, as the empirical validation was restricted to stylized binary diagnostic settings, specific prompting structures, and four clinical domains. Additionally, black-box testing cannot rule out strategic misreporting or assess internal model representations directly. Despite these scope limitations, the mathematical proofs and robust bootstrap confidence intervals provide high confidence that stated beliefs alone are currently insufficient to guarantee faithful decision explanations.

arXiv: 2602.06286
  • Paper: Metacognition in LLMs: Foundations, Progress, and Opportunities, Gabrielle Kaili-May Liu et al. (2026). It synthesizes broader frameworks and benchmarks for metacognition and confidence elicitation, extending the source's findings on belief-action consistency to the wider landscape of self-monitoring in LLMs.
  • Paper: Reasoning Models Generate Societies of Thought, Junsol Kim et al. (2026). It analyzes multi-perspective internal deliberations in reasoning models, providing a concrete setting to apply decision-theoretic tests of belief and action coherence.
Cover for When Agents Say One Thing and Do Another: Validating Elicited Beliefs from LLMs

Abstract

Large language models (LLMs) are increasingly deployed in high-stakes settings where good decisions require forming beliefs over the probability of unknown outcomes. However, it is unclear whether LLMs act as if they hold coherent beliefs when making decisions, or if so, how we could validate models' reports of such beliefs. We propose a decision-theoretic framework that elicits both probability judgments and decisions from an agent and tests their mutual consistency. Formally, our methods characterize whether it is possible for the actions to be produced by a ``near-rational" decision maker who holds the elicited probability as their true belief. We show that, perhaps surprisingly, this formalization implies empirically testable conditions even without any assumption about the agent's utility function. Applying our framework to stylized clinical diagnosis tasks, we find that models' reported beliefs are demonstrably imperfect summaries of the information revealed in their decisions, but that the discrepancies are small for the strongest models.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Methodology
  • 3.1 Decision-Theoretic Framework
  • 3.2 Testable Implications for Belief Elicitation
  • 4 Experimental Setup
  • 5 Results
  • 6 Discussion and Conclusions
  • References
  • A Proofs
  • A.1 Proof of Proposition
  • A.1.1 Conditional independence under prospect-theoretic risk aversion
  • A.2 Proof of Proposition
  • A.3 Proof of Proposition
  • A.4 Proof of Proposition
  • A.5 Proof of Proposition
  • B Prompts
  • B.1 Probability Elicitation Prompts
  • B.1.1 Standard Probability Elicitation (Prompt π0\pi_{0})
  • B.1.2 MSE Scoring Rule Prompt (πMSE\pi_{\text{MSE}})
  • B.1.3 Absolute Loss Scoring Rule Prompt (πABS\pi_{\text{ABS}})
  • B.1.4 Bayesian Reasoning Prompt (πBayes\pi_{\text{Bayes}})
  • B.2 Decision Elicitation Prompts
  • B.2.1 Decision Prompt A: Without Loss Function
  • B.3 Internal Consistency (Law of Iterated Expectation) Prompts
  • B.3.1 Next-State Distribution Prompt: PE​(z∈Bj∣x)P_{E}(z\in B_{j}\mid x)
  • B.3.2 Conditional Probability Prompt: PE​(θ∣x,z∈Bj)P_{E}(\theta\mid x,z\in B_{j})
  • B.4 Evidence-to-Language Conversion
  • B.5 Clinical Questions by Dataset
  • C Prompting Analysis
  • C.1 Prompt Consistency
  • D Heart Disease Dataset
  • D.0.1 Data Source
  • D.0.2 Target Variable
  • D.0.3 Covariate Variables
  • D.0.4 Bayesian Network Structure
  • D.0.5 Ground-Truth Probability Computation
  • D.0.6 Clinical Phrasing Examples
  • E Diabetes Dataset
  • E.0.1 Data Source
  • E.0.2 Target Variable
  • E.0.3 Covariate Variables
  • E.0.4 Bayesian Network Structure
  • E.0.5 Ground-Truth Probability Computation
  • E.0.6 Clinical Phrasing Examples
  • F Parameter Details
  • G Effect of Isotonic Calibration on Conditional Independence Tests
  • H Compute

Knowls

  1. Knowl 1 — Perturbed Utility Maximizer

    definition

    In a decision problem, an environment generates a latent state of the world θ∈Θ\theta \in \Theta according to prior distribution P⋆(θ)P^\star(\theta) and observations x∈Xx \in \mathcal{X} according to likelihood P⋆(x∣θ)P^\star(x \mid \theta). Upon observing xx, an agent forms a subjective posterior belief distribution PS(θ∣x)∈Δ(Θ)P_S(\theta \mid x) \in \Delta(\Theta) over states.

    Let A\mathcal{A} be a finite set of JJ available actions, and let Δ(A)\Delta(\mathcal{A}) denote the probability simplex over actions. An agent is defined as a perturbed utility maximizer if there exists:

    1. A utility function u:A×Θ→Ru : \mathcal{A} \times \Theta \to \mathbb{R} that is non-degenerate in relative action utilities, meaning there exist observations x,x′x, x' in the support of X\mathcal{X} and distinct actions a,a′∈Aa, a' \in \mathcal{A} such that ∑θ∈ΘPS(θ∣x)[u(a,θ)−u(a′,θ)]≠∑θ∈ΘPS(θ∣x′)[u(a,θ)−u(a′,θ)],\sum_{\theta \in \Theta} P_S(\theta \mid x) [u(a, \theta) - u(a', \theta)] \neq \sum_{\theta \in \Theta} P_S(\theta \mid x') [u(a, \theta) - u(a', \theta)],
    2. A strictly convex regularizer C:Δ(A)→RC : \Delta(\mathcal{A}) \to \mathbb{R},

    such that the agent's conditional choice-probability vector q(x)=(q1(x),…,qJ(x))∈Δ(A)q(x) = (q_1(x), \dots, q_J(x)) \in \Delta(\mathcal{A}) satisfies q(x)∈arg⁡max⁡λ∈Δ(A)(∑a∈Aλa(∑θ∈ΘPS(θ∣x)u(a,θ))−C(λ)).q(x) \in \arg\max_{\lambda \in \Delta(\mathcal{A})} \left( \sum_{a \in \mathcal{A}} \lambda_a \left( \sum_{\theta \in \Theta} P_S(\theta \mid x) u(a, \theta) \right) - C(\lambda) \right).

    The realized action AA is drawn according to Pr⁡(A=ai∣x,θ)=qai(x)\Pr(A = a_i \mid x, \theta) = q_{a_i}(x). This class encompasses discrete choice models including the multinomial logit model (where C(λ)=∑aλalog⁡λaC(\lambda) = \sum_a \lambda_a \log \lambda_a), random utility models (u(a,θ)+ϵau(a, \theta) + \epsilon_a with random noise ϵa\epsilon_a), KL-regularized policy optimization (C(λ)=DKL(λ∥q0)C(\lambda) = D_{\mathrm{KL}}(\lambda \parallel q_0)), and models of rational inattention.

  2. Knowl 2 — Joint Characterization of Rationalizable Elicited Beliefs

    theoretical result

    Let P\mathbb{P} denote an observed joint distribution over elicited beliefs PE(θ∣x)P_E(\theta \mid x), true states θ∈Θ\theta \in \Theta, and chosen actions a∈Aa \in \mathcal{A}. Let q(PE)∈Δ(A)q(P_E) \in \Delta(\mathcal{A}) denote the induced conditional choice rule (Pr⁡P(a=a′∣PE(θ∣x)=PE))a′∈A\left(\Pr_{\mathbb{P}}(a = a' \mid P_E(\theta \mid x) = P_E)\right)_{a' \in \mathcal{A}}, and let va(PE)=∑θ∈ΘPE(θ∣x)u(a,θ)v_a(P_E) = \sum_{\theta \in \Theta} P_E(\theta \mid x) u(a, \theta) denote the expected-utility index for action aa.

    There exists a utility function u:A×Θ→Ru : \mathcal{A} \times \Theta \to \mathbb{R} non-degenerate in relative action utilities and an agent satisfying the perturbed utility maximization framework (with a convex regularizer C:Δ(A)→RC : \Delta(\mathcal{A}) \to \mathbb{R}) whose subjective belief equals the elicited belief (PS(θ∣x)=PE(θ∣x)P_S(\theta \mid x) = P_E(\theta \mid x)) and whose induced joint distribution over (PE(θ∣x),θ,a)(P_E(\theta \mid x), \theta, a) equals P\mathbb{P}, if and only if:

    1. Conditional Independence (Belief Sufficiency): The chosen action is conditionally independent of the true state given the elicited belief: a⊥ ⁣ ⁣ ⁣⊥θ∣PE(θ∣x).a \perp\!\!\!\perp \theta \mid P_E(\theta \mid x).
    2. Cyclic Monotonicity: The conditional choice rule q(PE)q(P_E) is cyclically monotone with respect to the utility index v(PE)v(P_E).
  3. Knowl 3 — Conditional Independence under Truthful Reporting

    theoretical result

    For any agent that is a perturbed utility maximizer with subjective posterior PS(θ∣x)P_S(\theta \mid x) and choice rule q(x)∈Δ(A)q(x) \in \Delta(\mathcal{A}), the chosen action aa and the true outcome θ\theta are conditionally independent given the subjective belief: a⊥ ⁣ ⁣ ⁣⊥θ∣PS(θ∣x).a \perp\!\!\!\perp \theta \mid P_S(\theta \mid x).

    Conversely, if an agent verbalizes elicited beliefs PE(θ∣x)P_E(\theta \mid x) such that a⊥̸ ⁣ ⁣ ⁣⊥θ∣PE(θ∣x),a \not\perp\!\!\!\perp \theta \mid P_E(\theta \mid x), then there exists no perturbed utility maximizer (under any utility function uu and strictly convex regularizer CC) whose true subjective belief is PS=PEP_S = P_E. This property holds because the agent's expected utility depends on the context xx solely through the subjective belief; conditioning on that belief leaves no residual predictive information about θ\theta in the action aa.

    This conditional independence property also holds under non-expected-utility models such as prospect-theoretic risk aversion, provided the evaluation of each action depends on the context xx only through the subjective belief.

  4. Knowl 4 — Binary-State Characterization of Cyclic Monotonicity

    theoretical result

    Let the state space be binary, θ∈{0,1}\theta \in \{0, 1\}, and let p=PE(θ=1∣x)∈[0,1]p = P_E(\theta = 1 \mid x) \in [0, 1] be the elicited probability of the positive state. For an action set A={1,…,J}\mathcal{A} = \{1, \dots, J\}, define the utility-difference vector d=(d1,…,dJ)∈RJd = (d_1, \dots, d_J) \in \mathbb{R}^J by da:=u(a,1)−u(a,0),d_a := u(a, 1) - u(a, 0), which represents the relative advantage of taking action aa in state 11 versus state 00. Let q(p)=(q1(p),…,qJ(p))q(p) = (q_1(p), \dots, q_J(p)) be the conditional choice-probability vector at belief pp.

    The choice rule q(p)q(p) is cyclically monotone with respect to the expected-utility index vector v(p)=(v1(p),…,vJ(p))v(p) = (v_1(p), \dots, v_J(p)) if and only if the expected utility-difference index is weakly monotonically increasing in pp. That is, for every pair of beliefs p<p′p < p': ∑a=1Jqa(p)da≤∑a=1Jqa(p′)da.\sum_{a=1}^J q_a(p) d_a \le \sum_{a=1}^J q_a(p') d_a.

    Furthermore, for any agent satisfying perturbed utility maximization with a strictly convex regularizer CC, this inequality is strict whenever q(p)≠q(p′)q(p) \neq q(p').

  5. Knowl 5 — Signed-Margin Linear Program for Monotonicity Testing

    model/method

    To empirically test whether an agent's choice probabilities shift monotonically with elicited beliefs in a binary state space θ∈{0,1}\theta \in \{0, 1\}, elicited probabilities p(x)=PE(θ=1∣x)p(x) = P_E(\theta = 1 \mid x) are partitioned into KK quantile bins with centers pˉ1≤⋯≤pˉK\bar{p}_1 \le \dots \le \bar{p}_K. For each bin kk, the vector of empirical action shares is q^k=(q^1k,…,q^Jk)∈Δ(A)\hat{q}_k = (\hat{q}_{1k}, \dots, \hat{q}_{Jk}) \in \Delta(\mathcal{A}).

    Define the fitted index for bin kk given a utility-difference vector d∈[0,1]Jd \in [0, 1]^J as m^k(d):=∑a=1Jq^akda\hat{m}_k(d) := \sum_{a=1}^J \hat{q}_{ak} d_a. Let S:={k∈{1,…,K−1}:q^k≠q^k+1}S := \{k \in \{1, \dots, K-1\} : \hat{q}_k \neq \hat{q}_{k+1}\} denote the set of adjacent bins with differing empirical choice shares. For each ordered pair of distinct actions (a⋆,b⋆)(a^\star, b^\star) with a⋆≠b⋆a^\star \neq b^\star, the signed-margin linear program is defined as:

    \text{SignedLP}(a^\star, b^\star) : \quad \max_{d, \gamma} \quad & \gamma \\ \text{s.t.} \quad & \hat{m}_{k+1}(d) - \hat{m}_k(d) \ge \gamma, \quad \forall k \in S, \\ & d_{a^\star} = 0, \quad d_{b^\star} = 1, \\ & d \in [0, 1]^J, \quad \gamma \in \mathbb{R}. \end{aligned}$$ The overall signed monotonicity statistic is: $$\hat{\gamma} := \max_{a^\star \neq b^\star} \text{val}(\text{SignedLP}(a^\star, b^\star)).$$ Properties of $\hat{\gamma}$: 1. $\hat{\gamma} \in [-1, 1]$. 2. $\hat{\gamma} > 0$ if and only if there exists a non-constant $d \in \mathbb{R}^J$ such that $\hat{m}_1(d) \le \dots \le \hat{m}_K(d)$ with strict inequality $\hat{m}_k(d) < \hat{m}_{k+1}(d)$ for all $k \in S$ (strict monotonicity with uniform margin $\hat{\gamma}$). 3. $\hat{\gamma} \ge 0$ if and only if there exists a non-constant $d \in \mathbb{R}^J$ such that $\hat{m}_1(d) \le \dots \le \hat{m}_K(d)$ (weak monotonicity). 4. $\hat{\gamma} < 0$ implies every normalized $d$ produces a negative increment for at least one adjacent pair in $S$, violating weak monotonicity.
  6. Knowl 6 — Law of Iterated Expectation Internal Consistency Metric

    model/method

    To evaluate internal probabilistic coherence independently of actions, beliefs are tested for compliance with the Law of Iterated Expectation (LIE) across an auxiliary context variable zz partitioned into subsets {B1,…,Bk}\{B_1, \dots, B_k\}.

    For a base clinical context xx (with zz withheld), the model is queried for:

    1. The unconditional belief: PE(θ∣x)P_E(\theta \mid x),
    2. The conditional beliefs: PE(θ∣x,z∈Bj)P_E(\theta \mid x, z \in B_j) for each partition element BjB_j,
    3. The transition distribution: PE(z∈Bj∣x)P_E(z \in B_j \mid x).

    The Law of Iterated Expectation absolute discrepancy is computed as: ΔLIE(x)=∣PE(θ∣x)−∑j=1kPE(θ∣x,z∈Bj)PE(z∈Bj∣x)∣.\Delta_{\mathrm{LIE}}(x) = \left| P_E(\theta \mid x) - \sum_{j=1}^k P_E(\theta \mid x, z \in B_j) P_E(z \in B_j \mid x) \right|.

    To prevent distortion in cases where the elicited base probability p(x)=PE(θ=1∣x)p(x) = P_E(\theta = 1 \mid x) is near zero, results are evaluated using the median normalized ratio ΔLIE(x)/p(x)\Delta_{\mathrm{LIE}}(x) / p(x).

  7. Knowl 7 — Cross-Task Framing Stability Metric

    model/method

    Under subjective probability theory (prize independence), an agent's reported belief about an outcome θ\theta given context xx should remain invariant to how the downstream task is framed. To measure cross-prompt stability, elicited probabilities are compared across pairs of prompts sharing identical context (x,θ)(x, \theta) but using different task framings: the default diagnostic prompt π0\pi_0 versus a Mean Squared Error prompt πMSE\pi_{\mathrm{MSE}} that explicitly informs the model its estimates will be evaluated under the proper Brier scoring rule.

    For nn contexts evaluated across rr repetitions, let p(xi;π,j)p(x_i; \pi, j) denote the elicited probability for context xix_i under prompt π\pi on repetition jj. Cross-task inconsistency is defined as the root mean squared error (RMSE): ΔMSE=1nr∑i=1n∑j=1r(p(xi;πMSE,j)−p(xi;π0,j))2.\Delta_{\mathrm{MSE}} = \sqrt{\frac{1}{nr} \sum_{i=1}^n \sum_{j=1}^r \left( p(x_i; \pi_{\mathrm{MSE}}, j) - p(x_i; \pi_0, j) \right)^2}.

  8. Knowl 8 — Empirical Violation of Belief Sufficiency in LLMs

    empirical result

    Belief sufficiency (A⊥ ⁣ ⁣ ⁣⊥θ∣pA \perp\!\!\!\perp \theta \mid p) was evaluated on four clinical diagnostic tasks (Structural Heart Disease, Diabetes, Pediatric Fever, Infant Crying) across GPT-5 Thinking High Reasoning, GPT-5 Minimal Reasoning, DeepSeek-R1 (671B), and Llama-4 Scout (17B) with 200 cases and 5 repetitions.

    Two statistical tests were applied:

    1. kk-Nearest Neighbors Conditional Mutual Information (kNN CMI, k=3k=3): All 16 model-dataset pairs rejected the null hypothesis H0:I(A;θ∣p)=0H_0 : I(A; \theta \mid p) = 0 with 95% bootstrap confidence intervals strictly above zero (ranging from 0.0193 to 0.4223).
    2. Out-of-Sample Predictive MSE Reduction: A Random Forest predicting true state θ\theta from elicited belief alone (θ∼p\theta \sim p) was compared to an augmented model including both action and belief (θ∼(A,p)\theta \sim (A, p)) using grouped cross-validation to prevent leakage across repetitions.
    Dataset / Model kNN CMI [95% CI] RF % Imp. [95% CI] Group Avg RF % Imp. [95% CI]
    Heart–GPT-Min 0.1454 [0.1119, 0.1789] 18.39 [0.99, 31.66] GPT-Minimal 6.15 [-1.22, 13.52]
    Heart–GPT-High 0.0753 [0.0422, 0.1085] 9.93 [2.91, 17.75] GPT-High 4.89 [1.25, 8.53]
    Heart–Llama 0.0718 [0.0365, 0.1070] 14.60 [6.15, 22.64] Llama 15.19 [6.11, 24.27]
    Heart–DeepSeek 0.0675 [0.0354, 0.0997] 22.41 [9.64, 32.98] DeepSeek-R1 5.11 [-4.83, 15.06]
    Cry–GPT-Min 0.2232 [0.1874, 0.2589] 6.27 [-0.40, 11.43] Heart Disease 16.33 [11.81, 20.85]
    Cry–GPT-High 0.1901 [0.1390, 0.2412] 6.92 [2.22, 10.67] Cry 5.21 [-0.03, 10.45]
    Cry–Llama 0.4223 [0.3764, 0.4681] 11.12 [2.39, 20.25] Fever 2.07 [0.30, 3.85]
    Cry–DeepSeek 0.1745 [0.1265, 0.2225] -3.48 [-6.17, -0.62] Diabetes 7.73 [-4.93, 20.38]
    Fever–GPT-Min 0.1446 [0.1128, 0.1765] -0.06 [-0.24, 0.14]
    Fever–GPT-High 0.0944 [0.0564, 0.1323] 1.85 [0.32, 3.27]
    Fever–Llama 0.3289 [0.2893, 0.3686] 4.95 [-1.31, 10.65]
    Fever–DeepSeek 0.2060 [0.1663, 0.2456] 1.54 [-3.65, 6.52]
    Diabetes–GPT-Min 0.0193 [0.0129, 0.0258] 0.00 [-0.00, 0.00]
    Diabetes–GPT-High 0.0461 [0.0182, 0.0740] 0.85 [-0.26, 1.91]
    Diabetes–Llama 0.2695 [0.2357, 0.3033] 30.07 [20.56, 38.96]
    Diabetes–DeepSeek 0.0351 [0.0169, 0.0533] -0.03 [-0.10, 0.01]

    These results demonstrate that LLMs possess decision-relevant predictive information about true patient outcomes that is expressed in their choice of action but omitted from their verbalized probability estimates.

  9. Knowl 9 — Empirical Evaluation of Action Monotonicity across LLM Beliefs

    empirical result

    The signed-margin cyclic monotonicity statistic γ^\hat{\gamma} was estimated across four clinical domains with elicited probabilities partitioned into five quantile bins.

    Dataset / Model γ^\hat{\gamma} [95% CI] Dataset / Model γ^\hat{\gamma} [95% CI]
    Heart–GPT-Min 0.104 [0.012, 0.168] Fever–GPT-Min 0.053 [0.000, 0.082]
    Heart–GPT-High 0.059 [0.000, 0.160] Fever–GPT-High 0.103 [0.001, 0.194]
    Heart–Llama 0.042 [0.005, 0.068] Fever–Llama -0.011 [-0.018, -0.005]
    Heart–DeepSeek 0.192 [0.117, 0.205] Fever–DeepSeek 0.148 [0.078, 0.175]
    Cry–GPT-Min 0.050 [0.001, 0.117] Diabetes–GPT-Min 0.000 [0.000, 0.032]
    Cry–GPT-High 0.059 [0.003, 0.120] Diabetes–GPT-High 0.031 [0.000, 0.049]
    Cry–Llama 0.266 [0.217, 0.317] Diabetes–Llama 0.000 [0.000, 0.003]
    Cry–DeepSeek 0.097 [0.034, 0.142] Diabetes–DeepSeek 0.010 [0.000, 0.028]
    Model Average γ^\hat{\gamma} Domain Average γ^\hat{\gamma}
    GPT-Min 0.052 Heart 0.099
    GPT-High 0.063 Cry 0.118
    Llama 0.074 Fever 0.073
    DeepSeek-R1 0.112 Diabetes 0.010

    Key empirical findings:

    1. In 9 of 16 settings, the 95% confidence interval for γ^\hat{\gamma} was strictly positive, demonstrating strict cyclic monotonicity.
    2. Six settings had non-negative intervals containing zero, consistent with weak monotonicity.
    3. Fever–Llama exhibited a statistically significant violation of weak monotonicity (γ^=−0.011\hat{\gamma} = -0.011, 95% CI [−0.018,−0.005][-0.018, -0.005]).
    4. DeepSeek-R1 achieved the highest average signed margin across models (0.112), while Diabetes had the weakest domain alignment (average γ^=0.010\hat{\gamma} = 0.010).
  10. Knowl 10 — Inefficacy of Post-Hoc Isotonic Calibration for Restoring Belief Sufficiency

    empirical result

    To test whether belief insufficiency (A⊥̸ ⁣ ⁣ ⁣⊥θ∣pA \not\perp\!\!\!\perp \theta \mid p) arises from simple marginal miscalibration, raw elicited probabilities pp were transformed via monotonic isotonic regression into calibrated probabilities pisop_{\mathrm{iso}}, and the conditional mutual information I(A;θ∣piso)I(A; \theta \mid p_{\mathrm{iso}}) was re-estimated using a kk-nearest neighbors estimator (k=3k=3).

    CMI (Raw pp) CMI (Isotonic pisop_{\mathrm{iso}})
    Dataset / Model CMI 95% CI CMI 95% CI
    Heart–GPT-Min 0.1454 [0.1119, 0.1789] 0.3088 [0.2711, 0.3464]
    Heart–GPT-High 0.0753 [0.0422, 0.1085] 0.3295 [0.2967, 0.3624]
    Heart–Llama 0.0718 [0.0365, 0.1070] 0.1011 [0.0650, 0.1373]
    Heart–DeepSeek 0.0675 [0.0354, 0.0997] 0.1948 [0.1604, 0.2291]
    Cry–GPT-Min 0.2232 [0.1874, 0.2589] 0.2589 [0.2247, 0.2931]
    Cry–GPT-High 0.1901 [0.1390, 0.2412] 0.2002 [0.1499, 0.2505]
    Cry–Llama 0.4223 [0.3764, 0.4681] 0.5857 [0.5340, 0.6375]
    Cry–DeepSeek 0.1745 [0.1265, 0.2225] 0.2110 [0.1673, 0.2547]
    Fever–GPT-Min 0.1446 [0.1128, 0.1765] 0.1918 [0.1593, 0.2242]
    Fever–GPT-High 0.0944 [0.0564, 0.1323] 0.1112 [0.0718, 0.1505]
    Fever–Llama 0.3289 [0.2893, 0.3686] 0.5584 [0.5167, 0.6001]
    Fever–DeepSeek 0.2060 [0.1663, 0.2456] 0.2407 [0.1947, 0.2867]
    Diab–GPT-Min 0.0193 [0.0129, 0.0258] 0.0300 [0.0222, 0.0378]
    Diab–GPT-High 0.0461 [0.0182, 0.0740] 0.0340 [0.0013, 0.0667]
    Diab–Llama 0.2695 [0.2357, 0.3033] 0.4521 [0.4044, 0.4998]
    Diab–DeepSeek 0.0351 [0.0169, 0.0533] 0.0353 [0.0162, 0.0544]

    Across all models and datasets, I(A;θ∣piso)I(A; \theta \mid p_{\mathrm{iso}}) remained strictly positive, and in many instances increased in magnitude relative to raw beliefs. This indicates that belief insufficiency is caused by structural discrepancies between verbalized probabilities and internal decision representations rather than marginal calibration error.

  11. Knowl 11 — Decoupling of Probabilistic Coherence and Decision Consistency

    empirical result

    Internal probabilistic coherence was evaluated by measuring the median normalized Law of Iterated Expectation error, ΔLIE(x)/p(x)\Delta_{\mathrm{LIE}}(x) / p(x), and compared against a cross-validated Random Forest baseline trained directly on ground truth conditional distributions P⋆(x)P^\star(x) and P⋆(x,z)P^\star(x, z).

    Key observations:

    1. All LLMs (GPT-5 Thinking High, GPT-5 Minimal, DeepSeek-R1 671B, Llama-4 Scout 17B) exhibited normalized LIE errors substantially higher than the data-driven Random Forest baseline across all clinical domains (Heart Disease, Pediatric Cry, Pediatric Fever, Diabetes).
    2. No single LLM consistently dominated across all datasets in probabilistic coherence.
    3. Performance on decision-consistency tests (e.g., cyclic monotonicity and low cross-task variation) did not correlate with lower LIE error.

    This demonstrates that probabilistic coherence (conformance to probability axioms) and decision-theoretic coherence (beliefs functioning as sufficient statistics for choices) represent distinct, decoupled dimensions of language model behavior.

  12. Knowl 12 — Observational Equivalence and Falsifiability Limits in Black-Box Belief Elicitation

    limitation

    The black-box decision-theoretic validation framework operates without access to internal model representations and without assuming a specific utility function. Consequently, passing the conditional independence (A⊥ ⁣ ⁣ ⁣⊥θ∣PEA \perp\!\!\!\perp \theta \mid P_E) and cyclic monotonicity tests only establishes that the agent is observationally equivalent to some near-rational decision maker holding PEP_E as a subjective probability on the evaluated data distribution.

    This constitutes a failure to falsify the claim that the model acts on its reported beliefs, rather than proof that the model internally represents or computes those beliefs. In particular, black-box testing cannot detect deliberate misreporting, deceptive alignment, or strategically faked belief verbalization by a sufficiently sophisticated model.

Coverage note — None omitted; all core theoretical definitions, necessary/sufficient conditions, linear programming formulation, internal and cross-task consistency metrics, empirical benchmark results, calibration analysis, and conceptual limitations are fully represented.

References

  1. 1.Ankur Ankan and Johannes Textor. pgmpy: A python toolkit for bayesian networks. Journal of Machine Learning Research, 25(265):1–8, 2024. URL http://jmlr.org/papers/v25/23-0487.html.
  2. 2.Francis Anscombe and Robert Aumann. A definition of subjective probability. The Annals of Mathematical Statistics, 34, 03 1963. doi: 10.1214/aoms/1177704255.
  3. 3.Amos Azaria and Tom Mitchell. The internal state of an llm knows when it’s lying. In Findings of the Association for Computational Linguistics: EMNLP 2023, pp. 967–976, 2023.
  4. 4.Leonard Bereska and Efstratios Gavves. Mechanistic interpretability for ai safety–a review, 2024. URL https://arxiv. org/abs/2404.14082.
  5. 5.Marcel Binz and Eric Schulz. Using cognitive psychology to understand gpt-3. Proceedings of the National Academy of Sciences, 120(6):e2218523120, 2023.
  6. 6.Mariana Blanco, Dirk Engelmann, Alexander K Koch, and Hans-Theo Normann. Belief elicitation in experiments: is there a hedging problem? Experimental economics, 13(4):412–438, 2010.
  7. 7.CDC. Cdc diabetes health indicators, 2017. URL https://archive.ics.uci.edu/dataset/891.
  8. 8.Gary Charness, Uri Gneezy, and Vlastimil Rasocha. Experimental methods: Eliciting beliefs. Journal of Economic Behavior & Organization, 189:234–256, 2021.
  9. 9.Vanessa Cheung, Maximilian Maier, and Falk Lieder. Large language models show amplified cognitive biases in moral decision-making. Proceedings of the National Academy of Sciences, 122(25):e2412015122, 2025.
  10. 10.Rachel TA Croson. Thinking like a game theorist: factors affecting the frequency of equilibrium play. Journal of economic behavior & organization, 41(3):299–314, 2000.
  11. 11.Andre F Cruz, Moritz Hardt, and Celestine Mendler-Dunner. Evaluating language models as risk scores. Advances in Neural Information Processing Systems, 37:97378–97407, 2024.
  12. 12.Jessica Maria Echterhoff, Yao Liu, Abeer Alessa, Julian McAuley, and Zexue He. Cognitive bias in decision-making with llms. In Findings of the association for computational linguistics: EMNLP 2024, pp. 12640–12653, 2024.
  13. 13.Pierre Elias and Joshua Finer. Echonext: A dataset for detecting echocardiogram-confirmed structural heart disease from ecgs, Sep 2025. URL https://physionet.org/content/echonext/1.1.0/.
  14. 14.Mogens Fosgerau, Emerson Melo, Andre De Palma, and Matthew Shum. Discrete choice and rational inattention: A general equivalence result. International economic review, 61(4):1569–1589, 2020.
  15. 15.Gabriel Freedman and Francesca Toni. Exploring the potential for large language models to demonstrate rational probabilistic beliefs. arXiv preprint arXiv:2504.13644, 2025.
  16. 16.Drew Fudenberg, Ryota Iijima, and Tomasz Strzalecki. Stochastic choice and revealed perturbed utility. Econometrica, 83(6):2371–2409, 2015.
  17. 17.Farieda Gaber, Maqsood Shaik, Fabio Allega, Agnes Julia Bilecz, Felix Busch, Kelsey Goon, Vedran Franke, and Altuna Akalin. Evaluating large language model workflows in clinical decision support for triage and referral and diagnosis, May 2025. URL https://www.nature.com/articles/s41746-025-01684-1.
  18. 18.Tilmann Gneiting and Adrian E Raftery. Strictly proper scoring rules, prediction, and estimation. Journal of the American statistical Association, 102(477):359–378, 2007.
  19. 19.Maxime Griot, Coralie Hemptinne, Jean Vanderdonckt, and Demet Yuksel. Large language models lack essential metacognition for reliable medical reasoning. Nature Communications, 16, 01 2025. doi: 10.1038/s41467-024-55628-6.
  20. 20.Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Peiyi Wang, Qihao Zhu, Runxin Xu, Ruoyu Zhang, Shirong Ma, Xiao Bi, Xiaokang Zhang, Xingkai Yu, Yu Wu, Z. F. Wu, Zhibin Gou, Zhihong Shao, Zhuoshu Li, Ziyi Gao, Aixin Liu, Bing Xue, Bingxuan Wang, Bochao Wu, Bei Feng, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chong Ruan, Damai Dai, Deli Chen, Dongjie Ji, Erhang Li, Fangyun Lin, Fucong Dai, Fuli Luo, Guangbo Hao, Guanting Chen, Guowei Li, H. Zhang, Hanwei Xu, Honghui Ding, Huazuo Gao, Hui Qu, Hui Li, Jianzhong Guo, Jiashi Li, Jingchang Chen, Jingyang Yuan, Jinhao Tu, Junjie Qiu, Junlong Li, J. L. Cai, Jiaqi Ni, Jian Liang, Jin Chen, Kai Dong, Kai Hu, Kaichao You, Kaige Gao, Kang Guan, Kexin Huang, Kuai Yu, Lean Wang, Lecong Zhang, Liang Zhao, Litong Wang, Liyue Zhang, Lei Xu, Leyi Xia, Mingchuan Zhang, Minghua Zhang, Minghui Tang, Mingxu Zhou, Meng Li, Miaojun Wang, Mingming Li, Ning Tian, Panpan Huang, Peng Zhang, Qiancheng Wang, Qinyu Chen, Qiushi Du, Ruiqi Ge, Ruisong Zhang, Ruizhe Pan, Runji Wang, R. J. Chen, R. L. Jin, Ruyi Chen, Shanghao Lu, Shangyan Zhou, Shanhuang Chen, Shengfeng Ye, Shiyu Wang, Shuiping Yu, Shunfeng Zhou, Shuting Pan, S. S. Li, Shuang Zhou, Shaoqing Wu, Tao Yun, Tian Pei, Tianyu Sun, T. Wang, Wangding Zeng, Wen Liu, Wenfeng Liang, Wenjun Gao, Wenqin Yu, Wentao Zhang, W. L. Xiao, Wei An, Xiaodong Liu, Xiaohan Wang, Xiaokang Chen, Xiaotao Nie, Xin Cheng, Xin Liu, Xin Xie, Xingchao Liu, Xinyu Yang, Xinyuan Li, Xuecheng Su, Xuheng Lin, X. Q. Li, Xiangyue Jin, Xiaojin Shen, Xiaosha Chen, Xiaowen Sun, Xiaoxiang Wang, Xinnan Song, Xinyi Zhou, Xianzu Wang, Xinxia Shan, Y. K. Li, Y. Q. Wang, Y. X. Wei, Yang Zhang, Yanhong Xu, Yao Li, Yao Zhao, Yaofeng Sun, Yaohui Wang, Yi Yu, Yichao Zhang, Yifan Shi, Yiliang Xiong, Ying He, Yishi Piao, Yisong Wang, Yixuan Tan, Yiyang Ma, Yiyuan Liu, Yongqiang Guo, Yuan Ou, Yuduan Wang, Yue Gong, Yuheng Zou, Yujia He, Yunfan Xiong, Yuxiang Luo, Yuxiang You, Yuxuan Liu, Yuyang Zhou, Y. X. Zhu, Yanping Huang, Yaohui Li, Yi Zheng, Yuchen Zhu, Yunxian Ma, Ying Tang, Yukun Zha, Yuting Yan, Z. Z. Ren, Zehui Ren, Zhangli Sha, Zhe Fu, Zhean Xu, Zhenda Xie, Zhengyan Zhang, Zhewen Hao, Zhicheng Ma, Zhigang Yan, Zhiyu Wu, Zihui Gu, Zijia Zhu, Zijun Liu, Zilin Li, Ziwei Xie, Ziyang Song, Zizheng Pan, Zhen Huang, Zhipeng Xu, Zhongyu Zhang, and Zhen Zhang. Deepseek-r1 incentivizes reasoning in llms through reinforcement learning. Nature, 645 (8081):633–638, September 2025. ISSN 1476-4687. doi: 10.1038/s41586-025-09422-z. URL http://dx.doi.org/10.1038/s41586-025-09422-z.
  21. 21.Paul Hager, Friederike Jungmann, Robbie Holland, Kunal Bhagat, Inga Hubrecht, Manuel Knauer, Jakob Vielhauer, Marcus Makowski, Rickmer Braren, Georgios Kaissis, and Daniel Rueckert. Evaluation and mitigation of the limitations of large language models in clinical decision-making. Nature Medicine, 30:2613–2622, 07 2024. doi: 10.1038/s41591-024-03097-1.
  22. 22.Daniel A Herrmann and Benjamin A Levinstein. Standards for belief representations in llms. Minds and Machines, 35(1):5, 2024.
  23. 23.Josef Hofbauer and William H Sandholm. On the global convergence of stochastic fictitious play. Econometrica, 70(6):2265–2294, 2002.
  24. 24.Shreya Johri, Jaehwan Jeong, Benjamin A Tran, Daniel I Schlessinger, Shannon Wongvibulsin, Leandra A Barnes, Hong-Yu Zhou, Zhuo Ran Cai, Eliezer M Van Allen, David Kim, et al. An evaluation framework for clinical use of large language models in patient interaction tasks. Nature medicine, 31(1):77–86, 2025.
  25. 25.Daniel Kahneman and Amos Tversky. Prospect theory: An analysis of decision under risk. Econometrica, 47(2):263–291, 1979. ISSN 00129682, 14680262. URL http://www.jstor.org/stable/1914185.
  26. 26.Stella Li, Vidhisha Balachandran, Shangbin Feng, Jonathan Ilgen, Emma Pierson, Pang Wei W Koh, and Yulia Tsvetkov. Mediq: Question-asking llms and a benchmark for reliable interactive clinical reasoning. Advances in Neural Information Processing Systems, 37:28858–28888, 2024.
  27. 27.Ryan Liu, Jiayi Geng, Joshua C Peterson, Ilia Sucholutsky, and Thomas L Griffiths. Large language models assume people are more rational than we really are. arXiv preprint arXiv:2406.17055, 2024.
  28. 28.Samuel Marks and Max Tegmark. The geometry of truth: Emergent linear structure in large language model representations of true/false datasets. arXiv preprint arXiv:2310.06824, 2023.
  29. 29.Daniel McFadden. Conditional logit analysis of qualitative choice behavior. Frontiers in Econometrics, pp. 105–142, 1973.
  30. 30.Daniel McFadden and Marcel K Richter. Stochastic rationality and revealed stochastic preference. Preferences, Uncertainty, and Optimality, Essays in Honor of Leo Hurwicz, Westview Press: Boulder, CO, pp. 161–186, 1990.
  31. 31.MetaAI. Introducing llama 4: Advancing multimodal intelligence. https://ai.meta.com/blog/llama-4-multimodal-intelligence/, 2024.
  32. 32.Yaw Nyarko and Andrew Schotter. An experimental study of belief learning using elicited beliefs. Econometrica, 70(3):971–1005, 2002.
  33. 33.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. Training language models to follow instructions with human feedback, 2022. URL https://arxiv.org/abs/2203.02155.
  34. 34.Arka Pal, Teo Kitanovski, Arthur Liang, Akilesh Potti, and Micah Goldblum. Incoherent beliefs & inconsistent actions in large language models, 2025. URL https://arxiv.org/abs/2511.13240.
  35. 35.Pedro Rey-Biel. Equilibrium play and best response to (stated) beliefs in normal form games. Games and Economic Behavior, 65(2):572–585, 2009.
  36. 36.R Tyrrell Rockafellar. Convex analysis, volume 28. Princeton university press, 1997.
  37. 37.David Ronayne, Roberto Veneziani, and William R Zame. Do decision makers have subjective probabilities? an experimental test. SSRN, 2022.
  38. 38.Leonard J Savage. Elicitation of personal probabilities and expectations. Journal of the American Statistical Association, 66(336):783–801, 1971.
  39. 39.Leonard J Savage. The foundations of statistics. Courier Corporation, 1972.
  40. 40.Karl H Schlag, James Tremewan, and Joel J Van der Weele. A penny for your thoughts: A survey of methods for eliciting beliefs. Experimental Economics, 18(3):457–490, 2015.
  41. 41.Lee Sharkey, Bilal Chughtai, Joshua Batson, Jack Lindsey, Jeff Wu, Lucius Bushnaq, Nicholas Goldowsky-Dill, Stefan Heimersheim, Alejandro Ortega, Joseph Bloom, et al. Open problems in mechanistic interpretability. arXiv preprint arXiv:2501.16496, 2025.
  42. 42.Aaditya Singh, Adam Fry, Adam Perelman, Adam Tart, Adi Ganesh, Ahmed El-Kishky, Aidan McLaughlin, Aiden Low, AJ Ostrow, Akhila Ananthram, Akshay Nathan, Alan Luo, Alec Helyar, Aleksander Madry, Aleksandr Efremov, Aleksandra Spyra, Alex Baker-Whitcomb, Alex Beutel, Alex Karpenko, Alex Makelov, Alex Neitz, Alex Wei, Alexandra Barr, Alexandre Kirchmeyer, Alexey Ivanov, Alexi Christakis, Alistair Gillespie, Allison Tam, Ally Bennett, Alvin Wan, Alyssa Huang, Amy McDonald Sandjideh, Amy Yang, Ananya Kumar, Andre Saraiva, Andrea Vallone, Andrei Gheorghe, Andres Garcia Garcia, Andrew Braunstein, Andrew Liu, Andrew Schmidt, Andrey Mereskin, Andrey Mishchenko, Andy Applebaum, Andy Rogerson, Ann Rajan, Annie Wei, Anoop Kotha, Anubha Srivastava, Anushree Agrawal, Arun Vijayvergiya, Ashley Tyra, Ashvin Nair, Avi Nayak, Ben Eggers, Bessie Ji, Beth Hoover, Bill Chen, Blair Chen, Boaz Barak, Borys Minaiev, Botao Hao, Bowen Baker, Brad Lightcap, Brandon McKinzie, Brandon Wang, Brendan Quinn, Brian Fioca, Brian Hsu, Brian Yang, Brian Yu, Brian Zhang, Brittany Brenner, Callie Riggins Zetino, Cameron Raymond, Camillo Lugaresi, Carolina Paz, Cary Hudson, Cedric Whitney, Chak Li, Charles Chen, Charlotte Cole, Chelsea Voss, Chen Ding, Chen Shen, Chengdu Huang, Chris Colby, Chris Hallacy, Chris Koch, Chris Lu, Christina Kaplan, Christina Kim, CJ Minott-Henriques, Cliff Frey, Cody Yu, Coley Czarnecki, Colin Reid, Colin Wei, Cory Decareaux, Cristina Scheau, Cyril Zhang, Cyrus Forbes, Da Tang, Dakota Goldberg, Dan Roberts, Dana Palmie, Daniel Kappler, Daniel Levine, Daniel Wright, Dave Leo, David Lin, David Robinson, Declan Grabb, Derek Chen, Derek Lim, Derek Salama, Dibya Bhattacharjee, Dimitris Tsipras, Dinghua Li, Dingli Yu, DJ Strouse, Drew Williams, Dylan Hunn, Ed Bayes, Edwin Arbus, Ekin Akyurek, Elaine Ya Le, Elana Widmann, Eli Yani, Elizabeth Proehl, Enis Sert, Enoch Cheung, Eri Schwartz, Eric Han, Eric Jiang, Eric Mitchell, Eric Sigler, Eric Wallace, Erik Ritter, Erin Kavanaugh, Evan Mays, Evgenii Nikishin, Fangyuan Li, Felipe Petroski Such, Filipe de Avila Belbute Peres, Filippo Raso, Florent Bekerman, Foivos Tsimpourlas, Fotis Chantzis, Francis Song, Francis Zhang, Gaby Raila, Garrett McGrath, Gary Briggs, Gary Yang, Giambattista Parascandolo, Gildas Chabot, Grace Kim, Grace Zhao, Gregory Valiant, Guillaume Leclerc, Hadi Salman, Hanson Wang, Hao Sheng, Haoming Jiang, Haoyu Wang, Haozhun Jin, Harshit Sikchi, Heather Schmidt, Henry Aspegren, Honglin Chen, Huida Qiu, Hunter Lightman, Ian Covert, Ian Kivlichan, Ian Silber, Ian Sohl, Ibrahim Hammoud, Ignasi Clavera, Ikai Lan, Ilge Akkaya, Ilya Kostrikov, Irina Kofman, Isak Etinger, Ishaan Singal, Jackie Hehir, Jacob Huh, Jacqueline Pan, Jake Wilczynski, Jakub Pachocki, James Lee, James Quinn, Jamie Kiros, Janvi Kalra, Jasmyn Samaroo, Jason Wang, Jason Wolfe, Jay Chen, Jay Wang, Jean Harb, Jeffrey Han, Jeffrey Wang, Jennifer Zhao, Jeremy Chen, Jerene Yang, Jerry Tworek, Jesse Chand, Jessica Landon, Jessica Liang, Ji Lin, Jiancheng Liu, Jianfeng Wang, Jie Tang, Jihan Yin, Joanne Jang, Joel Morris, Joey Flynn, Johannes Ferstad, Johannes Heidecke, John Fishbein, John Hallman, Jonah Grant, Jonathan Chien, Jonathan Gordon, Jongsoo Park, Jordan Liss, Jos Kraaijeveld, Joseph Guay, Joseph Mo, Josh Lawson, Josh McGrath, Joshua Vendrow, Joy Jiao, Julian Lee, Julie Steele, Julie Wang, Junhua Mao, Kai Chen, Kai Hayashi, Kai Xiao, Kamyar Salahi, Kan Wu, Karan Sekhri, Karan Sharma, Karan Singhal, Karen Li, Kenny Nguyen, Keren Gu-Lemberg, Kevin King, Kevin Liu, Kevin Stone, Kevin Yu, Kristen Ying, Kristian Georgiev, Kristie Lim, Kushal Tirumala, Kyle Miller, Lama Ahmad, Larry Lv, Laura Clare, Laurance Fauconnet, Lauren Itow, Lauren Yang, Laurentia Romaniuk, Leah Anise, Lee Byron, Leher Pathak, Leon Maksin, Leyan Lo, Leyton Ho, Li Jing, Liang Wu, Liang Xiong, Lien Mamitsuka, Lin Yang, Lindsay McCallum, Lindsey Held, Liz Bourgeois, Logan Engstrom, Lorenz Kuhn, Louis Feuvrier, Lu Zhang, Lucas Switzer, Lukas Kondraciuk, Lukasz Kaiser, Manas Joglekar, Mandeep Singh, Mandip Shah, Manuka Stratta, Marcus Williams, Mark Chen, Mark Sun, Marselus Cayton, Martin Li, Marvin Zhang, Marwan Aljubeh, Matt Nichols, Matthew Haines, Max Schwarzer, Mayank Gupta, Meghan Shah, Melody Huang, Meng Dong, Mengqing Wang, Mia Glaese, Micah Carroll, Michael Lampe, Michael Malek, Michael Sharman, Michael Zhang, Michele Wang, Michelle Pokrass, Mihai Florian, Mikhail Pavlov, Miles Wang, Ming Chen, Mingxuan Wang, Minnia Feng, Mo Bavarian, Molly Lin, Moose Abdool, Mostafa Rohaninejad, Nacho Soto, Natalie Staudacher, Natan LaFontaine, Nathan Marwell, Nelson Liu, Nick Preston, Nick Turley, Nicklas Ansman, Nicole Blades, Nikil Pancha, Nikita Mikhaylin, Niko Felix, Nikunj Handa, Nishant Rai, Nitish Keskar, Noam Brown, Ofir Nachum, Oleg Boiko, Oleg Murk, Olivia Watkins, Oona Gleeson, Pamela Mishkin, Patryk Lesiewicz, Paul Baltescu, Pavel Belov, Peter Zhokhov, Philip Pronin, Phillip Guo, Phoebe Thacker, Qi Liu, Qiming Yuan, Qinghua Liu, Rachel Dias, Rachel Puckett, Rahul Arora, Ravi Teja Mullapudi, Raz Gaon, Reah Miyara, Rennie Song, Rishabh Aggarwal, RJ Marsan, Robel Yemiru, Robert Xiong, Rohan Kshirsagar, Rohan Nuttall, Roman Tsiupa, Ronen Eldan, Rose Wang, Roshan James, Roy Ziv, Rui Shu, Ruslan Nigmatullin, Saachi Jain, Saam Talaie, Sam Altman, Sam Arnesen, Sam Toizer, Sam Toyer, Samuel Miserendino, Sandhini Agarwal, Sarah Yoo, Savannah Heon, Scott Ethersmith, Sean Grove, Sean Taylor, Sebastien Bubeck, Sever Banesiu, Shaokyi Amdo, Shengjia Zhao, Sherwin Wu, Shibani Santurkar, Shiyu Zhao, Shraman Ray Chaudhuri, Shreyas Krishnaswamy, Shuaiqi, Xia, Shuyang Cheng, Shyamal Anadkat, Simon Posada Fishman, Simon Tobin, Siyuan Fu, Somay Jain, Song Mei, Sonya Egoian, Spencer Kim, Spug Golden, SQ Mah, Steph Lin, Stephen Imm, Steve Sharpe, Steve Yadlowsky, Sulman Choudhry, Sungwon Eum, Suvansh Sanjeev, Tabarak Khan, Tal Stramer, Tao Wang, Tao Xin, Tarun Gogineni, Taya Christianson, Ted Sanders, Tejal Patwardhan, Thomas Degry, Thomas Shadwell, Tianfu Fu, Tianshi Gao, Timur Garipov, Tina Sriskandarajah, Toki Sherbakov, Tomer Kaftan, Tomo Hiratsuka, Tongzhou Wang, Tony Song, Tony Zhao, Troy Peterson, Val Kharitonov, Victoria Chernova, Vineet Kosaraju, Vishal Kuo, Vitchyr Pong, Vivek Verma, Vlad Petrov, Wanning Jiang, Weixing Zhang, Wenda Zhou, Wenlei Xie, Wenting Zhan, Wes McCabe, Will DePue, Will Ellsworth, Wulfie Bain, Wyatt Thompson, Xiangning Chen, Xiangyu Qi, Xin Xiang, Xinwei Shi, Yann Dubois, Yaodong Yu, Yara Khakbaz, Yifan Wu, Yilei Qian, Yin Tat Lee, Yinbo Chen, Yizhen Zhang, Yizhong Xiong, Yonglong Tian, Young Cha, Yu Bai, Yu Yang, Yuan Yuan, Yuanzhi Li, Yufeng Zhang, Yuguang Yang, Yujia Jin, Yun Jiang, Yunyun Wang, Yushi Wang, Yutian Liu, Zach Stubenvoll, Zehao Dou, Zheng Wu, and Zhigang Wang. Openai gpt-5 system card, 2025. URL https://arxiv.org/abs/2601.03267.
  43. 43.K. Singhal, T. Tu, J. Gottweis, R. Sayres, E. Wulczyn, M. Amin, L. Hou, K. Clark, S. Pfohl, H. Cole-Lewis, D. Neal, Q. Rashid, M. Schaekermann, A. Wang, D. Dash, J. Chen, N. Shah, S. Lachgar, P. Mansfield, S. Prakash, B. Green, E. Dominowska, B. Aguera y Arcas, N. Tomasev, Y. Liu, R. Wong, C. Semturs, S. Mahdavi, J. Barral, D. Webster, G. Corrado, Y. Matias, S. Azizi, A. Karthikesalingam, and V. Natarajan. Toward expert-level medical question answering with large language models. Nature Medicine, 31:943–950, January 2025. doi: 10.1038/s41591-024-03423-7. URL https://doi.org/10.1038/s41591-024-03423-7.
  44. 44.Kenneth E Train. Discrete Choice Methods with Simulation. Cambridge University Press, 2nd edition, 2009.
  45. 45.Dennis Ulmer, Martin Gubri, Hwaran Lee, Sangdoo Yun, and Seong Oh. Calibrating large language models using their generations only. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 15440–15459, 2024.
  46. 46.Bo Waggoner and Yiling Chen. Output agreement mechanisms and common knowledge. In Second AAAI Conference on Human Computation and Crowdsourcing, HCOMP, 2014.
  47. 47.Cheng Wang, Gyuri Szarvas, Georges Balazs, Pavel Danchenko, and Patrick Ernst. Calibrating verbalized probabilities for large language models. arXiv preprint arXiv:2410.06707, 2024.
  48. 48.Christopher Y. K. Williams, Jaskaran Bains, Tianyu Tang, Kishan Patel, Alexa N. Lucas, Fiona Chen, Brenda Y. Miao, Atul J. Butte, and Aaron E. Kornblith. Evaluating large language models for drafting emergency department encounter summaries. PLOS Digital Health, 4(6):1–14, 06 2025. doi: 10.1371/journal.pdig.0000899. URL https://doi.org/10.1371/journal.pdig.0000899.
  49. 49.Jian-Qiao Zhu and Thomas L Griffiths. Eliciting the priors of large language models using iterated in-context learning. arXiv preprint arXiv:2406.01860, 2024.
  50. 50.Jian-Qiao Zhu and Thomas L. Griffiths. Incoherent probability judgments in large language models, 2025. URL https://arxiv.org/abs/2401.16646.

Citation

MLA
Yamin, K., et al. “When Agents Say One Thing and Do Another: Validating Elicited Beliefs from LLMs”. arXiv, 2026, http://arxiv.org/abs/2602.06286v3.
APA
Yamin, K., Tang, J., Cortes-Gomez, S., Sharma, A., Horvitz, E., & Wilder, B. (2026). When Agents Say One Thing and Do Another: Validating Elicited Beliefs from LLMs. arXiv. http://arxiv.org/abs/2602.06286v3
Chicago
Yamin, K., J. Tang, S. Cortes-Gomez, A. Sharma, E. Horvitz, and B. Wilder. 2026. “When Agents Say One Thing and Do Another: Validating Elicited Beliefs from LLMs”. arXiv. http://arxiv.org/abs/2602.06286v3.
Harvard
Yamin, K. et al. (2026) “When Agents Say One Thing and Do Another: Validating Elicited Beliefs from LLMs”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2602.06286v3.
Vancouver
1. Yamin K, Tang J, Cortes-Gomez S, Sharma A, Horvitz E, Wilder B (2026) When Agents Say One Thing and Do Another: Validating Elicited Beliefs from LLMs. arXiv

BibTeX

@article{yamin2026when,
  title = {When Agents Say One Thing and Do Another: Validating Elicited Beliefs from LLMs},
  author = {Yamin, Khurram and Tang, Jingjing and Cortes-Gomez, Santiago and Sharma, Amit and Horvitz, Eric and Wilder, Bryan},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2602.06286v3},
  eprint = {2602.06286}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/