Built independently by an author, for readers. Read the story and support ChapterPal

keyword

LLM decision-making

Large language model decision-making refers to the process by which an artificial intelligence model evaluates input information, weighs alternatives, and selects a specific output or course of action. Driven by probabilistic calculations over learned linguistic and semantic patterns, this mechanism enables models to perform tasks such as selecting among discrete options, answering complex questions, generating multistep reasoning paths, and acting in task-oriented environments. Rather than strictly executing deterministic rules or relying purely on verbatim memorization, model decisions emerge from interactions between contextual representations, prior statistical associations, structural cues within prompts, and inferred task objectives. Understanding this process involves analyzing how internal representations, reasoning heuristics, and prompt dynamics guide model choices, helping researchers distinguish robust generalization and problem-solving abilities from reliance on superficial artifacts or statistical shortcuts.

1 item

Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?

Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?

Nishant Balepur, Abhilasha Ravichander, Rachel Rudinger

OrganizationsAllen Institute for AIUniversity of Maryland

Why you should read this

Reveals that large language models frequently select the correct answer in multiple-choice benchmarks using only the answer options by abductively inferring missing questions and exploiting choice group dynamics rather than relying on memorization alone.

Multiple-choice question answering (MCQA) is often used to evaluate large language models (LLMs). To see if MCQA assesses LLMs as intended, we probe if LLMs can perform MCQA with choices-only prompts, where models must select the correct answer only from the choices. In three MCQA datasets and four LLMs, this prompt bests a majority baseline in 11/12 cases, with up to 0.33 accuracy gain. To help explain this behavior, we conduct an in-depth, black-box analysis on memorization, choice dynamics, and question inference. Our key findings are threefold. First, we find no evidence that the choices-only accuracy stems from memorization alone. Second, priors over individual choices do not fully explain choices-only accuracy, hinting that LLMs use the group dynamics of choices. Third, LLMs have some ability to infer a relevant question from choices, and surprisingly can sometimes even match the original question. Inferring the original question is an impressive reasoning strategy, but it cannot fully explain the high choices-only accuracy of LLMs in MCQA. Thus, while LLMs are not fully incapable of reasoning in MCQA, we still advocate for the use of stronger baselines in MCQA benchmarks, the design of robust MCQA datasets for fair evaluations, and further efforts to explain LLM decision-making.

Added

2026-09-26