Can Revealed Preferences Clarify LLM Alignment and Steering?
Khurram YaminJingjing TangEric HorvitzBryan Wilder
Introduces an empirical method using revealed preferences to recover the latent cost functions governing language model decisions, demonstrating that frontier models often fail to faithfully report or adjust their decision tradeoffs in high-stakes medical tasks.
As large language models are increasingly deployed to support high-stakes choices under uncertainty, evaluating their alignment requires understanding not only their factual accuracy but also how they weigh tradeoffs among competing risks. In critical applications such as clinical decision support, failures often arise from implicitly misweighting false positives, false negatives, or the decision to defer to human review. The article develops a formal, decision-theoretic framework based on revealed preferences to rigorously evaluate whether large language models make decisions as if pursuing coherent goals, whether they can accurately verbalize those goals, and whether prompt instructions can reliably steer their decision policies toward user-specified objectives.
To establish this framework, the article introduces a statistical pipeline that elicits a model's numeric beliefs regarding unknown factors alongside its selected decisions for the same scenario, and then fits a discrete-choice multinomial logit model using maximum likelihood estimation to recover the implied cost ratios that best rationalize observed behavior. The study evaluates frontier and open-source models—specifically GPT-5 (Thinking High and Minimal configurations), DeepSeek-R1 (671B), and Llama-4 Scout (17B)—across four clinical diagnosis domains: structural heart disease, diabetes, pediatric fever, and infant crying. The authors evaluate decision consistency, test whether stated cost preferences match implied behavior, and conduct counterfactual simulations to measure the effects of steering interventions.
Key findings show that while models exhibit a substantial degree of internal decision-making consistency, their self-reports and steerability suffer from major limitations. First, models fail to faithfully verbalize their operative objectives: self-reported cost tradeoffs match realized choices poorly (between 24.2% and 60.4% consistency across models), whereas revealed-preference estimates rationalize baseline decisions far better (61.3% to 82.2%). Second, attempting to steer models by explicitly specifying cost functions yields highly erratic outcomes; while implied preferences often shift toward the target, models only land within 80% to 120% of the target in 14.6% to 31.2% of settings, frequently undershooting, overshooting (up to 45.8% for Llama), or moving in the wrong direction (up to 27.1% for DeepSeek-R1). Third, counterfactual analyses confirm that cost-function steering effects are tightly explained by changes in the revealed cost parameters (Pearson correlation of r = 0.91), validating the utility-model framework. Finally, supplying explicit probability values in prompts improves decision consistency (up to 100% in GPT-5 High) but unintentionally alters the model's underlying cost weighting, making models substantially less willing to defer.
These findings have direct operational and safety implications for decision-support deployments. Relying on an AI system’s stated reasoning or self-reported priorities introduces substantial risk because its actual behavior follows an unstated and often skewed decision rule. Furthermore, simple prompt instructions cannot be trusted to reliably enforce organizational risk tolerances, as models either misinterpret explicit penalties or produce off-target preference shifts. However, because the revealed-preference framework reliably predicts post-steering loss reductions, decision-theoretic modeling provides a rigorous, black-box diagnostic tool to measure misalignment and evaluate model behavior prior to real-world integration.
Organizations considering the deployment of language models for high-stakes decision workflows should avoid relying on verbalized self-reports or ad-hoc prompting to control decision policies. Instead, teams should implement empirical revealed-preference benchmarking to quantify operative error tradeoffs and verify alignment against ground-truth benchmarks. Before adopting these systems in practice, further testing across non-medical operational domains and broader multi-class decision spaces is recommended. Because this approach operates externally on observable input-output behavior, it does not uncover internal neural mechanisms, but extensive sensitivity analyses confirm high confidence and robustness against noisy probability elicitation.
- Paper: When Agents Say One Thing and Do Another: Validating Elicited Beliefs from LLMs, Khurram Yamin et al. (2026). This foundational study introduces the formal decision-theoretic and perturbed utility framework for testing whether elicited LLM beliefs faithfully rationalize their actions in the same medical diagnosis settings.
- Paper: Maximum-Likelihood Inverse Reinforcement Learning with Finite-Time Guarantees, Siliang Zeng et al. (2022). It develops the maximum-likelihood discrete choice and inverse reinforcement learning foundations used to recover cost and reward functions from observed decision policies.
- Paper: Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting, Miles Turpin et al. (2023). It provides crucial empirical groundwork demonstrating that models' stated explanations frequently diverge from the actual internal drivers of their decisions.
- Paper: Rewards-in-Context: Multi-objective Alignment of Foundation Models with Dynamic Preference Adjustment, Rui Yang et al. (2024). It establishes techniques for dynamically steering model tradeoffs and aligning models to user-specified cost and reward objectives at inference time.
- Paper: Measuring Inductive Biases of In-Context Learning with Underspecified Demonstrations, Chenglei Si et al. (2023). It analyzes the underlying inductive biases in model decision policies and assesses how prompt steering influences model tradeoffs under underspecification.
- Paper: KTO: Model Alignment as Prospect Theoretic Optimization, Kawin Ethayarajh et al. (2024). It frames LLM alignment through behavioral economic and prospect-theoretic utility functions under uncertainty, which underpins revealed-preference formulations.
- Paper: Discovering Latent Knowledge in Language Models Without Supervision, Collin Burns et al. (2023). It provides key methodology for extracting implicit, latent beliefs from language models without relying on unfaithful verbalized explanations.
- Paper: Reward Models Inherit Value Biases from Pretraining, Brian Christian et al. (2026). It examines how the underlying preference and value biases analyzed via revealed preferences are systematically inherited from pretraining into downstream alignment models.
- Paper: Reinforcement Learning Towards Broadly and Persistently Beneficial Models, Akshay V. Jagadeesh et al. (2026). It extends the evaluation of model alignment and steerability by training models for persistent, robust beneficial preferences across high-stakes domains including healthcare.
