Can Revealed Preferences Clarify LLM Alignment and Steering?

Khurram YaminJingjing TangEric HorvitzBryan Wilder

article2026arXiv4 citations

Introduces an empirical method using revealed preferences to recover the latent cost functions governing language model decisions, demonstrating that frontier models often fail to faithfully report or adjust their decision tradeoffs in high-stakes medical tasks.

Listen

As large language models are increasingly deployed to support high-stakes choices under uncertainty, evaluating their alignment requires understanding not only their factual accuracy but also how they weigh tradeoffs among competing risks. In critical applications such as clinical decision support, failures often arise from implicitly misweighting false positives, false negatives, or the decision to defer to human review. The article develops a formal, decision-theoretic framework based on revealed preferences to rigorously evaluate whether large language models make decisions as if pursuing coherent goals, whether they can accurately verbalize those goals, and whether prompt instructions can reliably steer their decision policies toward user-specified objectives.

To establish this framework, the article introduces a statistical pipeline that elicits a model's numeric beliefs regarding unknown factors alongside its selected decisions for the same scenario, and then fits a discrete-choice multinomial logit model using maximum likelihood estimation to recover the implied cost ratios that best rationalize observed behavior. The study evaluates frontier and open-source models—specifically GPT-5 (Thinking High and Minimal configurations), DeepSeek-R1 (671B), and Llama-4 Scout (17B)—across four clinical diagnosis domains: structural heart disease, diabetes, pediatric fever, and infant crying. The authors evaluate decision consistency, test whether stated cost preferences match implied behavior, and conduct counterfactual simulations to measure the effects of steering interventions.

Key findings show that while models exhibit a substantial degree of internal decision-making consistency, their self-reports and steerability suffer from major limitations. First, models fail to faithfully verbalize their operative objectives: self-reported cost tradeoffs match realized choices poorly (between 24.2% and 60.4% consistency across models), whereas revealed-preference estimates rationalize baseline decisions far better (61.3% to 82.2%). Second, attempting to steer models by explicitly specifying cost functions yields highly erratic outcomes; while implied preferences often shift toward the target, models only land within 80% to 120% of the target in 14.6% to 31.2% of settings, frequently undershooting, overshooting (up to 45.8% for Llama), or moving in the wrong direction (up to 27.1% for DeepSeek-R1). Third, counterfactual analyses confirm that cost-function steering effects are tightly explained by changes in the revealed cost parameters (Pearson correlation of r = 0.91), validating the utility-model framework. Finally, supplying explicit probability values in prompts improves decision consistency (up to 100% in GPT-5 High) but unintentionally alters the model's underlying cost weighting, making models substantially less willing to defer.

These findings have direct operational and safety implications for decision-support deployments. Relying on an AI system’s stated reasoning or self-reported priorities introduces substantial risk because its actual behavior follows an unstated and often skewed decision rule. Furthermore, simple prompt instructions cannot be trusted to reliably enforce organizational risk tolerances, as models either misinterpret explicit penalties or produce off-target preference shifts. However, because the revealed-preference framework reliably predicts post-steering loss reductions, decision-theoretic modeling provides a rigorous, black-box diagnostic tool to measure misalignment and evaluate model behavior prior to real-world integration.

Organizations considering the deployment of language models for high-stakes decision workflows should avoid relying on verbalized self-reports or ad-hoc prompting to control decision policies. Instead, teams should implement empirical revealed-preference benchmarking to quantify operative error tradeoffs and verify alignment against ground-truth benchmarks. Before adopting these systems in practice, further testing across non-medical operational domains and broader multi-class decision spaces is recommended. Because this approach operates externally on observable input-output behavior, it does not uncover internal neural mechanisms, but extensive sensitivity analyses confirm high confidence and robustness against noisy probability elicitation.

arXiv: 2605.08556
Cover for Can Revealed Preferences Clarify LLM Alignment and Steering?

Abstract

LLMs are increasingly used to make or support high-stakes decisions under uncertainty, where alignment depends not only on factual accuracy but on how models weigh tradeoffs between different outcomes. We present an empirical pipeline for estimating the implied preferences that an LLM's observed choices optimize: we elicit the model's probability distribution over unknowns along with the choice it would make for the decision task and then fit a discrete choice model to recover the cost function that best rationalizes the model's decisions. We show how this revealed-preference description allows rigorous evaluation of whether models behave in a consistently goal-directed way, whether they can verbalize a description of their objectives which matches their revealed decision policy, and whether prompting can reliably steer those policies to implement a user-specified cost function. We apply this evaluation across four medical diagnosis domains and multiple frontier and open-source models. We find that while many models have a nontrivial degree of internal coherence, they also have significant weaknesses in faithfully reporting or adopting preferences in response to user direction.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Methodology
  • 3.1 Decision-Theoretic Framework
  • 3.2 Loss Function Estimation
  • 3.3 Counterfactual Evaluation of Steering Interventions
  • 4 Experimental Setup
  • 5 Results
  • 5.1 Deriving Implied Cost Functions
  • 5.2 Predicting Steering Benefits with Counterfactuals
  • 6 Discussion
  • References
  • A Belief Validity
  • A.1 Belief Validity Comparison
  • A.2 Downstream Analysis of Alternative Beliefs
  • A.3 Sensitivity of Loss Functions to Belief Noise
  • A.4 Averaging Beliefs
  • B Implied Loss Function Consistency By Domains
  • C Prompts
  • C.1 Probability Elicitation Prompts
  • C.1.1 Standard Probability Elicitation (Prompt π0\pi_{0})
  • C.1.2 Loss-Conditioned Probability Elicitation (Prompt πLF\pi_{\mathrm{LF}})
  • C.2 Decision Elicitation Prompts
  • C.2.1 Decision Prompt A: Without Loss Function
  • C.2.2 Decision Prompt B: With Loss Function
  • C.2.3 Decision Prompt C: With Ground-Truth Probability
  • C.2.4 Decision Prompt E: With Elicited Probability
  • C.3 Self-Reported Loss-Function Prompts
  • C.3.1 Global Self-Reported Loss Function
  • C.3.2 Case-Specific Self-Reported Loss Function
  • C.4 Evidence-to-Language Conversion
  • C.5 Clinical Questions by Dataset
  • D Datasets
  • D.1 Heart Disease
  • D.2 Diabetes
  • E Compute

Knowls

  1. Knowl 1 — Discrete Choice Framework for Estimating LLM Revealed Preferences

    model/method

    A decision-theoretic framework estimates the implicit utility function that rationalizes a large language model's observed choices under uncertainty.

    Let an environment generate an unobserved binary state θ∈Θ={0,1}\theta \in \Theta = \{0, 1\} and an observation xx, inducing a ground-truth posterior P⋆(θ∣x)P^\star(\theta \mid x). The decision maker forms a subjective posterior PS(θ∣x)P_S(\theta \mid x), proxied by an elicited numeric posterior probability pE(x):=PE(θ=1∣x)p_E(x) := P_E(\theta = 1 \mid x), and selects an action a∈A={1,0,defer}a \in \mathcal{A} = \{1, 0, \text{defer}\} representing diagnosing disease, diagnosing non-disease, or deferring decision-making.

    The loss function is parameterized as:

    ℓc(a,θ)=cFP⋅I[a=1,θ=0]+cFN⋅I[a=0,θ=1]+cdefer⋅I[a=defer]\ell_c(a, \theta) = c_{FP} \cdot \mathbb{I}[a = 1, \theta = 0] + c_{FN} \cdot \mathbb{I}[a = 0, \theta = 1] + c_{\text{defer}} \cdot \mathbb{I}[a = \text{defer}]

    where c=(cFP,cFN,cdefer)c = (c_{FP}, c_{FN}, c_{\text{defer}}) captures false-positive, false-negative, and deferral costs, and I[⋅]\mathbb{I}[\cdot] is the indicator function.

    Under a random utility model with Type-1 extreme value (Gumbel) unobserved shock εa\varepsilon_a and noise scale parameter β=1\beta = 1, the agent's expected loss for action aa given observation xx is ℓˉc(a,x)=Eθ∼PS(⋅∣x)[ℓc(a,θ)]\bar{\ell}_c(a, x) = \mathbb{E}_{\theta \sim P_S(\cdot \mid x)}[\ell_c(a, \theta)]. Choice probabilities follow a multinomial logit distribution:

    Pr⁡(a∣x;c,β=1)=exp⁡(−ℓˉc(a,x))∑a′∈Aexp⁡(−ℓˉc(a′,x))\Pr(a \mid x; c, \beta = 1) = \frac{\exp(-\bar{\ell}_c(a, x))}{\sum_{a' \in \mathcal{A}} \exp(-\bar{\ell}_c(a', x))}

    Given nn elicited beliefs and choices {(xi,ai)}i=1n\{(x_i, a_i)\}_{i=1}^n, the cost parameters c^\hat{c} are estimated via maximum likelihood:

    c^=arg⁡max⁡c∑i=1nlog⁡Pr⁡(ai∣xi;c,β=1)\hat{c} = \arg\max_c \sum_{i=1}^n \log \Pr(a_i \mid x_i; c, \beta = 1)

    optimized using L-BFGS-B. Because absolute scale is not identified when β\beta is fixed, the procedure recovers the relative cost ratios cFN/cFPc_{FN}/c_{FP} and cdefer/cFPc_{\text{defer}}/c_{FP}.

  2. Knowl 2 — Implied Loss-Function Consistency Metric

    definition

    The Implied Loss-Function Consistency (ILFC) metric measures the percentage of decision instances where a language model's observed choice matches the optimal action implied by combining a specified belief distribution with an inferred or reported loss function.

    For a given belief vector p~=(p~1,…,p~n)\tilde{p} = (\tilde{p}_1, \dots, \tilde{p}_n) and cost vector c^\hat{c}, the optimal decision aiopt(p~,c^)a_i^{\text{opt}}(\tilde{p}, \hat{c}) for instance ii is defined as the action minimizing expected loss:

    aiopt(p~,c^)∈arg⁡min⁡a∈AEθ∼p~i[ℓc^(a,θ)]a_i^{\text{opt}}(\tilde{p}, \hat{c}) \in \arg\min_{a \in \mathcal{A}} \mathbb{E}_{\theta \sim \tilde{p}_i}[\ell_{\hat{c}}(a, \theta)]

    where A\mathcal{A} is the action space and ℓc^(a,θ)\ell_{\hat{c}}(a, \theta) is the loss function parameterized by c^\hat{c}.

    The ILFC score across nn test cases is defined as:

    ILFC(p~,c^)=100n∑i=1nI[ai=aiopt(p~,c^)]\text{ILFC}(\tilde{p}, \hat{c}) = \frac{100}{n} \sum_{i=1}^n \mathbb{I}[a_i = a_i^{\text{opt}}(\tilde{p}, \hat{c})]

    where aia_i is the actual action chosen by the model in instance ii, and I[⋅]\mathbb{I}[\cdot] is the indicator function. Higher ILFC values indicate that the model's behavior is more consistent with rational expected loss minimization under the given belief and cost parameters.

  3. Knowl 3 — Counterfactual Decomposition for Steering Interventions

    model/method

    A counterfactual evaluation framework predicts the performance impact of steering interventions by isolating whether prompt changes alter factual beliefs, preference parameters, or both.

    Let c(k)c^{(k)} denote a benchmark target cost vector, c^\hat{c} the baseline implied cost vector estimated via maximum likelihood on baseline decisions, pE=(pE,1,…,pE,n)p_E = (p_{E,1}, \dots, p_{E,n}) the elicited baseline beliefs, and p⋆=(p1⋆,…,pn⋆)p^\star = (p^\star_1, \dots, p^\star_n) the ground-truth posterior probabilities. For any belief-cost pair (c,p)(c, p), let a(p,c)a(p, c) denote the simulated optimal decision vector where a(p,c)i=arg⁡min⁡α∈AEθ∼pi[ℓc(α,θ)]a(p, c)_i = \arg\min_{\alpha \in \mathcal{A}} \mathbb{E}_{\theta \sim p_i}[\ell_c(\alpha, \theta)]. The realized loss under benchmark cost c(k)c^{(k)} for action vector aa and true states θ=(θ1,…,θn)\theta = (\theta_1, \dots, \theta_n) is Lk(a)=∑i=1nℓc(k)(ai,θi)L_k(a) = \sum_{i=1}^n \ell_{c^{(k)}}(a_i, \theta_i).

    The counterfactual predicted percent loss reduction from transitioning from belief-cost state (c,p)(c, p) to (c′,p′)(c', p') is:

    Δ^k((c,p)→(c′,p′))=100⋅Lk(a(p,c))−Lk(a(p′,c′))Lk(a(p,c))\hat{\Delta}_k((c, p) \to (c', p')) = 100 \cdot \frac{L_k(a(p, c)) - L_k(a(p', c'))}{L_k(a(p, c))}

    The actual realized prompting effect is computed from actual model baseline decisions abasea_{\text{base}} and post-prompt decisions asteereda_{\text{steered}}:

    Δk=100⋅Lk(abase)−Lk(asteered)Lk(abase)\Delta_k = 100 \cdot \frac{L_k(a_{\text{base}}) - L_k(a_{\text{steered}})}{L_k(a_{\text{base}})}

    Two counterfactual predictions are evaluated:

    1. Target Prediction: Assumes the prompt updates only the targeted component while holding the other fixed: Δ^k((c^,pE)→(c(k),pE))\hat{\Delta}_k((\hat{c}, p_E) \to (c^{(k)}, p_E)) for cost prompting, and Δ^k((c^,pE)→(c^,p⋆))\hat{\Delta}_k((\hat{c}, p_E) \to (\hat{c}, p^\star)) for probability prompting.
    2. Steered Prediction: Evaluates the cost vector c^steered\hat{c}_{\text{steered}} re-estimated from post-prompt decisions: Δ^k((c^,pE)→(c^steered(k),pE))\hat{\Delta}_k((\hat{c}, p_E) \to (\hat{c}_{\text{steered}}^{(k)}, p_E)) for cost prompting, and Δ^k((c^,pE)→(c^p⋆,p⋆))\hat{\Delta}_k((\hat{c}, p_E) \to (\hat{c}_{p^\star}, p^\star)) for probability prompting.
  4. Knowl 4 — Behavioral Consistency of Revealed versus Self-Reported Preferences

    data/table

    Across four frontier and open-source models (GPT-5 Minimal, GPT-5 Thinking High, Llama-4 Scout 17B, and DeepSeek-R1 671B), self-reported cost functions fail to explain models' actual choices, while revealed-preference loss functions estimated via discrete choice modeling achieve substantially higher behavioral consistency.

    Table 2 reports the Implied Loss-Function Consistency (ILFC in %) across prompt conditions. Verbalized cost ratios (Global and Case-Specific Self-Report) match decisions poorly (24.2% to 60.4%), demonstrating that stated objectives are unreliable descriptions of operative decision policies. By contrast, baseline revealed preferences explain 61.3% to 82.2% of decisions. Supplying explicit probabilities (p⋆p^\star or elicited pEp_E) raises consistency up to 98.8%–100.0% for reasoning models (GPT-5 High and DeepSeek-R1).

    Model Global Self-Rep. Case-Specific Self-Rep. Baseline True pp Elicit. pp Cost Function
    GPT5-Minimal 24.2% 28.0% 82.2% 89.8% 86.9% 77.4%
    GPT5-High 58.0% 57.9% 74.6% 99.2% 100.0% 77.3%
    Llama 54.8% 60.4% 61.3% 90.6% 85.0% 76.7%
    DeepSeek-R1 41.2% 32.4% 73.1% 98.8% 94.5% 73.7%
  5. Knowl 5 — Steerability of Implied Utility Ratios Under Cost-Function Prompting

    data/table

    Prompting models with explicit benchmark cost tuples (cFP,cFN,cdefer)(c_{FP}, c_{FN}, c_{\text{defer}}) yields only partial and heterogeneous shifts in their implied preference ratios (cFN/cFPc_{FN}/c_{FP} and cdefer/cFPc_{\text{defer}}/c_{FP}). Progress is classified as "On Target" if the implied ratio shifts 80%–120% of the distance from baseline to target, "Undershot" if it moves in the correct direction by <80%, "Overshot" if it moves in the correct direction by >120%, and "Wrong" if it moves away from the target.

    Table 1 summarizes steering performance across all benchmark configurations:

    • DeepSeek-R1 is the most prone to moving in the wrong direction (27.1%).
    • GPT-5 Minimal predominantly undershoots the target cost ratios (47.9%).
    • GPT-5 Thinking High achieves the highest on-target rate (31.2%).
    • Llama-4 Scout 17B is the most prone to overshooting targets in the correct direction (45.8%).
    Model Wrong Direction Undershot On Target (80–120%) Overshot
    DeepSeek-R1 27.1% 29.2% 20.8% 22.9%
    Llama-4 Scout 14.6% 14.6% 25.0% 45.8%
    GPT5-High 12.5% 33.3% 31.2% 22.9%
    GPT5-Minimal 25.0% 47.9% 14.6% 12.5%
  6. Knowl 6 — Counterfactual Prediction of Cost and Probability Steering Effects

    empirical result

    Counterfactual predictions validate the belief-preference decomposition and identify off-target effects of prompt interventions:

    1. Cost-Function Steering: The realized loss reduction Δk(cost)\Delta_k(\text{cost}) is positively correlated with the target counterfactual prediction Δ^k((c^,pE)→(c(k),pE))\hat{\Delta}_k((\hat{c}, p_E) \to (c^{(k)}, p_E)) (Pearson r=0.557r = 0.557, p=3.87×10−9p = 3.87 \times 10^{-9}). When evaluated against the steered counterfactual prediction Δ^k((c^,pE)→(c^steered(k),pE))\hat{\Delta}_k((\hat{c}, p_E) \to (\hat{c}_{\text{steered}}^{(k)}, p_E)) using the post-prompt implied cost vector, the correlation increases to r=0.908r = 0.908 (p=3.094×10−37p = 3.094 \times 10^{-37}). This indicates that steering residuals stem from models imperfectly adopting target costs rather than unrationalizable behavior.

    2. Probabilistic Steering: The realized loss reduction Δk(prob)\Delta_k(\text{prob}) is weakly negatively correlated with the target counterfactual prediction Δ^k((c^,pE)→(c^,p⋆))\hat{\Delta}_k((\hat{c}, p_E) \to (\hat{c}, p^\star)) (Pearson r=−0.256r = -0.256, p=0.01p = 0.01). When evaluated against the steered counterfactual prediction Δ^k((c^,pE)→(c^p⋆,p⋆))\hat{\Delta}_k((\hat{c}, p_E) \to (\hat{c}_{p^\star}, p^\star)), the correlation becomes strongly positive (Pearson r=0.748r = 0.748, p=2.00×10−18p = 2.00 \times 10^{-18}). This reveals that supplying explicit probabilities does not merely update beliefs, but causes an off-target preference shift by substantially increasing the implicit relative cost of deferral (cdefer/cFPc_{\text{defer}}/c_{FP}) and forcing more decisive actions.

  7. Knowl 7 — Experimental Setup on Clinical Decision Benchmarks

    experimental setup

    The revealed-preference pipeline is instantiated across four stylized clinical diagnostic decision tasks under uncertainty:

    1. Structural Heart Disease: Derived from Columbia University Medical Center ECG and echocardiogram records (>100,000>100,000 patient encounters). Structured as a 20-node, 121-edge Bayesian network modeling demographics, 4 ECG measurements, and 11 echocardiographic flags. The target binary outcome is moderate-or-greater structural heart disease.
    2. Diabetes: Derived from the CDC Behavioral Risk Factor Surveillance System (BRFSS) survey. Structured as a 22-node, 77-edge Bayesian network spanning demographics, cardiometabolic variables, lifestyle behaviors, and healthcare access. The binary target is diabetes/prediabetes.
    3. Pediatric Fever and Infant Crying: Two clinician-specified pediatric Bayesian networks modeling symptoms (e.g., jaundice, lethargy, feeding difficulties) to predict fever threshold exceedance (≥99∘F\ge 99^\circ\text{F} oral / ≥100∘F\ge 100^\circ\text{F} rectal) and infant colic, respectively.

    Evaluated models comprise GPT-5 Thinking High, GPT-5 Minimal, DeepSeek-R1 (671B), and Llama-4 Scout (17B). Ground-truth posteriors p⋆(x)p^\star(x) are computed via variable elimination using pgmpy. For each case context xx, models are queried for subjective beliefs pE(x)p_E(x), baseline diagnostic actions a∈{1,0,defer}a \in \{1, 0, \text{defer}\}, actions under prompted beliefs (pE(x)p_E(x) or p⋆(x)p^\star(x)), actions under specified benchmark cost functions c(k)c^{(k)}, and self-reported cost ratios.

  8. Knowl 8 — Robustness of Inferred Cost Ratios to Elicitation Noise and Belief Framing

    empirical result

    Sensitivity analyses demonstrate that maximum likelihood estimates of revealed preference cost ratios (cFN/cFPc_{FN}/c_{FP} and cdefer/cFPc_{\text{defer}}/c_{FP}) are robust to belief perturbations and elicitation variants:

    1. Gaussian Belief Noise: Adding independent Gaussian noise N(0,σ2)\mathcal{N}(0, \sigma^2) to elicited probabilities with σ=0.05\sigma = 0.05 (causing an average relative belief shift of ≈10%\approx 10\%) changes the median estimated baseline cost ratios by at most 1.0%–2.5%1.0\%\text{--}2.5\% relative to unperturbed estimates.
    2. Repeated Belief Averaging: Averaging 5 independent belief elicitation replicates per instance alters the baseline estimated cost ratios by <2%<2\% (median absolute changes: 0.03290.0329 for cFN/cFPc_{FN}/c_{FP}, 0.00350.0035 for cdefer/cFPc_{\text{defer}}/c_{FP}; median percent changes: 1.88%1.88\% and 1.62%1.62\%, respectively).
    3. Belief Conditioning: Eliciting probabilities from prompts containing explicit loss functions worsens belief sufficiency compared to standard elicitation (higher log-loss improvement when ground truth θ\theta is added to a CatBoost model of actions A∼pA \sim p: signed difference of +11.3%+11.3\% for DeepSeek and +8.1%+8.1\% for GPT-5 High), confirming the validity of standard probability elicitation as the primary belief proxy.
  9. Knowl 9 — Methodological and Scope Limitations of Revealed Preference Probing

    limitation

    The paper identifies several inherent limitations in its framework and empirical evaluation:

    1. Black-Box Behavioral Scope: The discrete choice formulation models only the observable input-output decision behavior of language models. It does not identify or inspect internal mechanistic representations, attention patterns, or transformer circuits that generate the observed decisions.
    2. Subjective Belief Proxy Assumption: Estimating loss parameters relies on using explicitly elicited verbal probabilities pE(x)p_E(x) as a proxy for the model's true subjective belief distribution PS(θ∣x)P_S(\theta \mid x). Any discrepancy between verbalized probabilities and the internal probabilistic constructs driving action generation can introduce estimation bias.
    3. Stylized Decision Setting: The experimental case study is restricted to binary state classifications (Θ={0,1}\Theta = \{0, 1\}) with ternary action spaces (A={1,0,defer}\mathcal{A} = \{1, 0, \text{defer}\}) and assumes a multinomial logit random utility error distribution with fixed noise scale β=1\beta = 1, which may not fully represent more complex, multi-state, or continuous real-world decision tasks.

Coverage note — None was omitted; the knowls cover the full discrete-choice revealed preference formulation, consistency metric, counterfactual steering evaluation framework, empirical benchmark setup, consistency and steerability tables, predictive counterfactual results, belief sensitivity analyses, and stated limitations.

References

  1. 1.Ankan, A. and Textor, J. pgmpy: A python toolkit for bayesian networks. Journal of Machine Learning Research, 25(265):1–8, 2024. URL http://jmlr.org/papers/v25/23-0487.html.
  2. 2.Binz, M. and Schulz, E. Using cognitive psychology to understand gpt-3. Proceedings of the National Academy of Sciences, 120(6):e2218523120, 2023.
  3. 3.Burnell, R., Schellaert, W., Burden, J., Ullman, T., Plumed, F., Tenenbaum, J., Rutar, D., Cheke, L., Sohl-Dickstein, J., Mitchell, M., Kiela, D., Shanahan, M., Voorhees, E., Cohn, A., Leibo, J., and Hernandez-Orallo, J. Rethink reporting of evaluation results in ai. Science (New York, N.Y.), 380: 136–138, 04 2023. doi: 10.1126/science.adf6369.
  4. 4.Byrd, R. H., Lu, P., Nocedal, J., and Zhu, C. A limited memory algorithm for bound constrained optimization. SIAM Journal on Scientific Computing, 16(5):1190–1208, 1995. doi: 10.1137/ 0916069. URL https://doi.org/10.1137/0916069.
  5. 5.Casper, S., Davies, X., Shi, C., Gilbert, T. K., Scheurer, J., Rando, J., Freedman, R., Korbak, T., Lindner, D., Freire, P., Wang, T., Marks, S., Segerie, C.-R., Carroll, M., Peng, A., Christoffersen, P., Damani, M., Slocum, S., Anwar, U., Siththaranjan, A., Nadeau, M., Michaud, E. J., Pfau, J., Krasheninnikov, D., Chen, X., Langosco, L., Hase, P., Bıyık, E., Dragan, A., Krueger, D., Sadigh, D., and Hadfield-Menell, D. Open problems and fundamental limitations of reinforcement learning from human feedback, 2023. URL https://arxiv.org/abs/2307.15217.
  6. 6.Centers for Disease Control and Prevention. Cdc diabetes health indicators, 2017. URL https://archive.ics.uci.edu/dataset/891.
  7. 7.Elias, P. and Finer, J. Echonext: A dataset for detecting echocardiogram-confirmed structural heart disease from ecgs, Sep 2025. URL https://physionet.org/content/echonext/1.1.0/.
  8. 8.Gaber, F., Shaik, M., Allega, F., Bilecz, A. J., Busch, F., Goon, K., Franke, V., and Akalin, A. Evaluating large language model workflows in clinical decision support for triage and referral and diagnosis, May 2025. URL https://www.nature.com/articles/s41746-025-01684-1.
  9. 9.Griot, M., Hemptinne, C., Vanderdonckt, J., and Yuksel, D. Large language models lack essential metacognition for reliable medical reasoning. Nature Communications, 16, 01 2025. doi: 10.1038/s41467-024-55628-6.
  10. 10.Guo, D., Yang, D., Zhang, H., Song, J., Wang, P., Zhu, Q., Xu, R., Zhang, R., Ma, S., Bi, X., Zhang, X., Yu, X., Wu, Y., Wu, Z. F., Gou, Z., Shao, Z., Li, Z., Gao, Z., Liu, A., Xue, B., Wang, B., Wu, B., Feng, B., Lu, C., Zhao, C., Deng, C., Ruan, C., Dai, D., Chen, D., Ji, D., Li, E., Lin, F., Dai, F., Luo, F., Hao, G., Chen, G., Li, G., Zhang, H., Xu, H., Ding, H., Gao, H., Qu, H., Li, H., Guo, J., Li, J., Chen, J., Yuan, J., Tu, J., Qiu, J., Li, J., Cai, J. L., Ni, J., Liang, J., Chen, J., Dong, K., Hu, K., You, K., Gao, K., Guan, K., Huang, K., Yu, K., Wang, L., Zhang, L., Zhao, L., Wang, L., Zhang, L., Xu, L., Xia, L., Zhang, M., Zhang, M., Tang, M., Zhou, M., Li, M., Wang, M., Li, M., Tian, N., Huang, P., Zhang, P., Wang, Q., Chen, Q., Du, Q., Ge, R., Zhang, R., Pan, R., Wang, R., Chen, R. J., Jin, R. L., Chen, R., Lu, S., Zhou, S., Chen, S., Ye, S., Wang, S., Yu, S., Zhou, S., Pan, S., Li, S. S., Zhou, S., Wu, S., Yun, T., Pei, T., Sun, T., Wang, T., Zeng, W., Liu, W., Liang, W., Gao, W., Yu, W., Zhang, W., Xiao, W. L., An, W., Liu, X., Wang, X., Chen, X., Nie, X., Cheng, X., Liu, X., Xie, X., Liu, X., Yang, X., Li, X., Su, X., Lin, X., Li, X. Q., Jin, X., Shen, X., Chen, X., Sun, X., Wang, X., Song, X., Zhou, X., Wang, X., Shan, X., Li, Y. K., Wang, Y. Q., Wei, Y. X., Zhang, Y., Xu, Y., Li, Y., Zhao, Y., Sun, Y., Wang, Y., Yu, Y., Zhang, Y., Shi, Y., Xiong, Y., He, Y., Piao, Y., Wang, Y., Tan, Y., Ma, Y., Liu, Y., Guo, Y., Ou, Y., Wang, Y., Gong, Y., Zou, Y., He, Y., Xiong, Y., Luo, Y., You, Y., Liu, Y., Zhou, Y., Zhu, Y. X., Huang, Y., Li, Y., Zheng, Y., Zhu, Y., Ma, Y., Tang, Y., Zha, Y., Yan, Y., Ren, Z. Z., Ren, Z., Sha, Z., Fu, Z., Xu, Z., Xie, Z., Zhang, Z., Hao, Z., Ma, Z., Yan, Z., Wu, Z., Gu, Z., Zhu, Z., Liu, Z., Li, Z., Xie, Z., Song, Z., Pan, Z., Huang, Z., Xu, Z., Zhang, Z., and Zhang, Z. Deepseek-r1 incentivizes reasoning in llms through reinforcement learning. Nature, 645(8081):633–638, September 2025. ISSN 1476-4687. doi: 10.1038/s41586-025-09422-z. URL http://dx.doi.org/10.1038/s41586-025-09422-z.
  11. 11.Hager, P., Jungmann, F., Holland, R., Bhagat, K., Hubrecht, I., Knauer, M., Vielhauer, J., Makowski, M., Braren, R., Kaissis, G., and Rueckert, D. Evaluation and mitigation of the limitations of large language models in clinical decision-making. Nature Medicine, 30:2613–2622, 07 2024. doi: 10.1038/s41591-024-03097-1.
  12. 12.Hensher, D. A., Rose, J. M., and Greene, W. H. Applied Choice Analysis. Cambridge University Press, 2 edition, 2015.
  13. 13.Jia, J., Yuan, Z., Pan, J., McNamara, P. E., and Chen, D. Decision-making behavior evaluation framework for llms under uncertain context, 2024. URL https://arxiv.org/abs/2406.05972.
  14. 14.Khalaf, H., Verdun, C. M., Oesterling, A., Lakkaraju, H., and du Pin Calmon, F. Inference-time reward hacking in large language models, 2025. URL https://arxiv.org/abs/2506.19248.
  15. 15.Kim, J., Podlasek, A., Shidara, K., Liu, F., Alaa, A., and Bernardo, D. Limitations of large language models in clinical problem-solving arising from inflexible reasoning, Nov 2025. URL https://www.nature.com/articles/s41598-025-22940-0.
  16. 16.Liu, A., Ghate, K., Diab, M., Fried, D., Kasirzadeh, A., and Kleiman-Weiner, M. Generative value conflicts reveal llm priorities, 2026. URL https://arxiv.org/abs/2509.25369.
  17. 17.Liu, R., Geng, J., Peterson, J. C., Sucholutsky, I., and Griffiths, T. L. Large language models assume people are more rational than we really are. arXiv preprint arXiv:2406.17055, 2024.
  18. 18.Mazeika, M., Yin, X., Tamirisa, R., Lim, J., Lee, B. W., Ren, R., Phan, L., Mu, N., Khoja, A., Zhang, O., and Hendrycks, D. Utility engineering: Analyzing and controlling emergent value systems in ais, 2025. URL https://arxiv.org/abs/2502.08640.
  19. 19.McFadden, D. Conditional logit analysis of qualitative choice behavior. Frontiers in Econometrics, pp. 105–142, 1973.
  20. 20.MetaAI. Introducing llama 4: Advancing multimodal intelligence. https://ai.meta.com/blog/llama-4-multimodal-intelligence/, 2024.
  21. 21.Ouyang, S., Yun, H., and Zheng, X. Ai as decision-maker: Ethics and risk preferences of llms, 2025. URL https://arxiv.org/abs/2406.01168.
  22. 22.Paruchuri, A., Garrison, J., Liao, S., Hernandez, J., Sunshine, J., Althoff, T., Liu, X., and McDuff, D. What are the odds? language models are capable of probabilistic reasoning, 2024. URL https://arxiv.org/abs/2406.12830.
  23. 23.Pfohl, S. R., Cole-Lewis, H., Sayres, R., Neal, D., Asiedu, M., Dieng, A., Tomasev, N., Rashid, Q. M., Azizi, S., Rostamzadeh, N., McCoy, L. G., Celi, L. A., Liu, Y., Schaekermann, M., Walton, A., Parrish, A., Nagpal, C., Singh, P., Dewitt, A., Mansfield, P., Prakash, S., Heller, K., Karthikesalingam, A., Semturs, C., Barral, J., Corrado, G., Matias, Y., Smith-Loud, J., Horn, I., and Singhal, K. A toolbox for surfacing health equity harms and biases in large language models. Nature Medicine, 30(12):3590–3600, September 2024. ISSN 1546-170X. doi: 10.1038/s41591-024-03258-2. URL http://dx.doi.org/10.1038/s41591-024-03258-2.
  24. 24.Samway, K., Kleiman-Weiner, M., Piedrahita, D. G., Mihalcea, R., Schölkopf, B., and Jin, Z. Are language models consequentialist or deontological moral reasoners? In Christodoulopoulos, C., Chakraborty, T., Rose, C., and Peng, V. (eds.), Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pp. 30699–30726, Suzhou, China, November 2025. Association for Computational Linguistics. ISBN 979-8-89176-332-6. doi: 10.18653/v1/2025.emnlp-main.1563. URL https://aclanthology.org/2025.emnlp-main.1563/.
  25. 25.Singh, A., Fry, A., Perelman, A., Tart, A., Ganesh, A., El-Kishky, A., McLaughlin, A., Low, A., Ostrow, A., Ananthram, A., Nathan, A., Luo, A., Helyar, A., Madry, A., Efremov, A., Spyra, A., Baker-Whitcomb, A., Beutel, A., Karpenko, A., Makelov, A., Neitz, A., Wei, A., Barr, A., Kirchmeyer, A., Ivanov, A., Christakis, A., Gillespie, A., Tam, A., Bennett, A., Wan, A., Huang, A., Sandjideh, A. M., Yang, A., Kumar, A., Saraiva, A., Vallone, A., Gheorghe, A., Garcia, A. G., Braunstein, A., Liu, A., Schmidt, A., Mereskin, A., Mishchenko, A., Applebaum, A., Rogerson, A., Rajan, A., Wei, A., Kotha, A., Srivastava, A., Agrawal, A., Vijayvergiya, A., Tyra, A., Nair, A., Nayak, A., Eggers, B., Ji, B., Hoover, B., Chen, B., Chen, B., Barak, B., Minaiev, B., Hao, B., Baker, B., Lightcap, B., McKinzie, B., Wang, B., Quinn, B., Fioca, B., Hsu, B., Yang, B., Yu, B., Zhang, B., Brenner, B., Zetino, C. R., Raymond, C., Lugaresi, C., Paz, C., Hudson, C., Whitney, C., Li, C., Chen, C., Cole, C., Voss, C., Ding, C., Shen, C., Huang, C., Colby, C., Hallacy, C., Koch, C., Lu, C., Kaplan, C., Kim, C., Minott-Henriques, C., Frey, C., Yu, C., Czarnecki, C., Reid, C., Wei, C., Decareaux, C., Scheau, C., Zhang, C., Forbes, C., Tang, D., Goldberg, D., Roberts, D., Palmie, D., Kappler, D., Levine, D., Wright, D., Leo, D., Lin, D., Robinson, D., Grabb, D., Chen, D., Lim, D., Salama, D., Bhattacharjee, D., Tsipras, D., Li, D., Yu, D., Strouse, D., Williams, D., Hunn, D., Bayes, E., Arbus, E., Akyurek, E., Le, E. Y., Widmann, E., Yani, E., Proehl, E., Sert, E., Cheung, E., Schwartz, E., Han, E., Jiang, E., Mitchell, E., Sigler, E., Wallace, E., Ritter, E., Kavanaugh, E., Mays, E., Nikishin, E., Li, F., Such, F. P., de Avila Belbute Peres, F., Raso, F., Bekerman, F., Tsimpourlas, F., Chantzis, F., Song, F., Zhang, F., Raila, G., McGrath, G., Briggs, G., Yang, G., Parascandolo, G., Chabot, G., Kim, G., Zhao, G., Valiant, G., Leclerc, G., Salman, H., Wang, H., Sheng, H., Jiang, H., Wang, H., Jin, H., Sikchi, H., Schmidt, H., Aspegren, H., Chen, H., Qiu, H., Lightman, H., Covert, I., Kivlichan, I., Silber, I., Sohl, I., Hammoud, I., Clavera, I., Lan, I., Akkaya, I., Kostrikov, I., Kofman, I., Etinger, I., Singal, I., Hehir, J., Huh, J., Pan, J., Wilczynski, J., Pachocki, J., Lee, J., Quinn, J., Kiros, J., Kalra, J., Samaroo, J., Wang, J., Wolfe, J., Chen, J., Wang, J., Harb, J., Han, J., Wang, J., Zhao, J., Chen, J., Yang, J., Tworek, J., Chand, J., Landon, J., Liang, J., Lin, J., Liu, J., Wang, J., Tang, J., Yin, J., Jang, J., Morris, J., Flynn, J., Ferstad, J., Heidecke, J., Fishbein, J., Hallman, J., Grant, J., Chien, J., Gordon, J., Park, J., Liss, J., Kraaijeveld, J., Guay, J., Mo, J., Lawson, J., McGrath, J., Vendrow, J., Jiao, J., Lee, J., Steele, J., Wang, J., Mao, J., Chen, K., Hayashi, K., Xiao, K., Salahi, K., Wu, K., Sekhri, K., Sharma, K., Singhal, K., Li, K., Nguyen, K., Gu-Lemberg, K., King, K., Liu, K., Stone, K., Yu, K., Ying, K., Georgiev, K., Lim, K., Tirumala, K., Miller, K., Ahmad, L., Lv, L., Clare, L., Fauconnet, L., Itow, L., Yang, L., Romaniuk, L., Anise, L., Byron, L., Pathak, L., Maksin, L., Lo, L., Ho, L., Jing, L., Wu, L., Xiong, L., Mamitsuka, L., Yang, L., McCallum, L., Held, L., Bourgeois, L., Engstrom, L., Kuhn, L., Feuvrier, L., Zhang, L., Switzer, L., Kondraciuk, L., Kaiser, L., Joglekar, M., Singh, M., Shah, M., Stratta, M., Williams, M., Chen, M., Sun, M., Cayton, M., Li, M., Zhang, M., Aljubeh, M., Nichols, M., Haines, M., Schwarzer, M., Gupta, M., Shah, M., Huang, M., Dong, M., Wang, M., Glaese, M., Carroll, M., Lampe, M., Malek, M., Sharman, M., Zhang, M., Wang, M., Pokrass, M., Florian, M., Pavlov, M., Wang, M., Chen, M., Wang, M., Feng, M., Bavarian, M., Lin, M., Abdool, M., Rohaninejad, M., Soto, N., Staudacher, N., LaFontaine, N., Marwell, N., Liu, N., Preston, N., Turley, N., Ansman, N., Blades, N., Pancha, N., Mikhaylin, N., Felix, N., Handa, N., Rai, N., Keskar, N., Brown, N., Nachum, O., Boiko, O., Murk, O., Watkins, O., Gleeson, O., Mishkin, P., Lesiewicz, P., Baltescu, P., Belov, P., Zhokhov, P., Pronin, P., Guo, P., Thacker, P., Liu, Q., Yuan, Q., Liu, Q., Dias, R., Puckett, R., Arora, R., Mullapudi, R. T., Gaon, R., Miyara, R., Song, R., Aggarwal, R., Marsan, R., Yemiru, R., Xiong, R., Kshirsagar, R., Nuttall, R., Tsiupa, R., Eldan, R., Wang, R., James, R., Ziv, R., Shu, R., Nigmatullin, R., Jain, S., Talaie, S., Altman, S., Arnesen, S., Toizer, S., Toyer, S., Miserendino, S., Agarwal, S., Yoo, S., Heon, S., Ethersmith, S., Grove, S., Taylor, S., Bubeck, S., Banesiu, S., Amdo, S., Zhao, S., Wu, S., Santurkar, S., Zhao, S., Chaudhuri, S. R., Krishnaswamy, S., Shuaiqi, Xia, Cheng, S., Anadkat, S., Fishman, S. P., Tobin, S., Fu, S., Jain, S., Mei, S., Egoian, S., Kim, S., Golden, S., Mah, S., Lin, S., Imm, S., Sharpe, S., Yadlowsky, S., Choudhry, S., Eum, S., Sanjeev, S., Khan, T., Stramer, T., Wang, T., Xin, T., Gogineni, T., Christianson, T., Sanders, T., Patwardhan, T., Degry, T., Shadwell, T., Fu, T., Gao, T., Garipov, T., Sriskandarajah, T., Sherbakov, T., Kaftan, T., Hiratsuka, T., Wang, T., Song, T., Zhao, T., Peterson, T., Kharitonov, V., Chernova, V., Kosaraju, V., Kuo, V., Pong, V., Verma, V., Petrov, V., Jiang, W., Zhang, W., Zhou, W., Xie, W., Zhan, W., McCabe, W., DePue, W., Ellsworth, W., Bain, W., Thompson, W., Chen, X., Qi, X., Xiang, X., Shi, X., Dubois, Y., Yu, Y., Khakbaz, Y., Wu, Y., Qian, Y., Lee, Y. T., Chen, Y., Zhang, Y., Xiong, Y., Tian, Y., Cha, Y., Bai, Y., Yang, Y., Yuan, Y., Li, Y., Zhang, Y., Yang, Y., Jin, Y., Jiang, Y., Wang, Y., Wang, Y., Liu, Y., Stubenvoll, Z., Dou, Z., Wu, Z., and Wang, Z. Openai gpt-5 system card, 2025. URL https://arxiv.org/abs/2601.03267.
  26. 26.Singhal, K., Tu, T., Gottweis, J., Sayres, R., Wulczyn, E., Amin, M., Hou, L., Clark, K., Pfohl, S., Cole-Lewis, H., Neal, D., Rashid, Q., Schaekermann, M., Wang, A., Dash, D., Chen, J., Shah, N., Lachgar, S., Mansfield, P., Prakash, S., Green, B., Dominowska, E., Agüera y Arcas, B., Tomašev, N., Liu, Y., Wong, R., Semturs, C., Mahdavi, S., Barral, J., Webster, D., Corrado, G., Matias, Y., Azizi, S., Karthikesalingam, A., and Natarajan, V. Toward expert-level medical question answering with large language models. Nature Medicine, 31:943–950, January 2025. doi: 10.1038/s41591-024-03423-7. URL https://doi.org/10.1038/s41591-024-03423-7.
  27. 27.Slama, K., Souly, A., Bansal, D., Davidson, H., Summerfield, C., and Luettgau, L. When do llm preferences predict downstream behavior?, 2026. URL https://arxiv.org/abs/2602.18971.
  28. 28.Train, K. E. Discrete Choice Methods with Simulation. Cambridge University Press, 2nd edition, 2009.
  29. 29.Williams, C. Y. K., Bains, J., Tang, T., Patel, K., Lucas, A. N., Chen, F., Miao, B. Y., Butte, A. J., and Kornblith, A. E. Evaluating large language models for drafting emergency department encounter summaries. PLOS Digital Health, 4(6):1–14, 06 2025. doi: 10.1371/journal.pdig.0000899. URL https://doi.org/10.1371/journal.pdig.0000899.
  30. 30.Xiao, F. and Wang, X. X. Evaluating the ability of large language models to predict human social decisions. Scientific Reports, 15(1):32290, 2025.
  31. 31.Yamin, K., Tang, J., Cortes-Gomez, S., Sharma, A., Horvitz, E., and Wilder, B. Do llms act like rational agents? measuring belief coherence in probabilistic decision making, 2026. URL https://arxiv.org/abs/2602.06286.
  32. 32.Zhu, J.-Q., Yan, H., and Griffiths, T. L. Steering risk preferences in large language models by aligning behavioral and neural representations, 2025. URL https://arxiv.org/abs/2505.11615.

Citation

MLA
Yamin, K., et al. “Can Revealed Preferences Clarify LLM Alignment and Steering?”. arXiv, 2026, http://arxiv.org/abs/2605.08556v2.
APA
Yamin, K., Tang, J., Horvitz, E., & Wilder, B. (2026). Can Revealed Preferences Clarify LLM Alignment and Steering?. arXiv. http://arxiv.org/abs/2605.08556v2
Chicago
Yamin, K., J. Tang, E. Horvitz, and B. Wilder. 2026. “Can Revealed Preferences Clarify LLM Alignment and Steering?”. arXiv. http://arxiv.org/abs/2605.08556v2.
Harvard
Yamin, K. et al. (2026) “Can Revealed Preferences Clarify LLM Alignment and Steering?”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2605.08556v2.
Vancouver
1. Yamin K, Tang J, Horvitz E, Wilder B (2026) Can Revealed Preferences Clarify LLM Alignment and Steering?. arXiv

BibTeX

@article{yamin2026can,
  title = {Can Revealed Preferences Clarify LLM Alignment and Steering?},
  author = {Yamin, Khurram and Tang, Jingjing and Horvitz, Eric and Wilder, Bryan},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2605.08556v2},
  eprint = {2605.08556}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/