Premise Order Matters in Reasoning with Large Language Models
Xinyun ChenRyan A. ChiXuezhi WangDenny Zhou
Demonstrates that large language models suffer performance drops of over 30% on deductive and mathematical reasoning tasks when premises are simply reordered, revealing an autoregressive positional bias and introducing the R-GSM benchmark to measure this vulnerability.
Large language models are increasingly deployed to automate critical reasoning tasks, from financial analysis to decision support. However, real-world information is rarely organized sequentially to match the logical order needed to reach a conclusion. The article investigates a key operational vulnerability: large language models are highly brittle to premise ordering, experiencing steep performance declines when factual statements or rules are rearranged, even when the underlying logic and truth value remain completely unchanged.
The article's main objective is to evaluate how premise ordering influences model reasoning performance across deductive logic and mathematical problem-solving tasks. It systematically measures whether varying the sequence of identical premises degrades the accuracy and proof validity of state-of-the-art language models.
To conduct this evaluation, the researchers tested leading industry models, including GPT-4-turbo, GPT-3.5-turbo, PaLM 2-L, and Gemini 1.0 Pro. The study generated a logical reasoning dataset containing 27,000 problem variants across varying proof lengths and levels of distracting information, using pseudowords to isolate logical reasoning from pretrained internal knowledge. In addition, the researchers developed R-GSM, a mathematical reasoning benchmark consisting of 220 problem pairs adapted from grade-school math word problems, in which the sentence order was reordered without altering the mathematical meaning or final answer.
The findings show that all evaluated models perform best when premises follow the exact sequence of intermediate reasoning steps, termed the forward order. Permuting this order causes severe performance drops exceeding 30% on complex logical problems for top-tier models and over 40% for weaker models. On mathematical reasoning tasks, models failed on 10% to over 35% of problems they had solved correctly in the original order. Introducing irrelevant or distracting premises further magnifies these accuracy drops. Error analyses revealed that out-of-order premises primarily lead to fact hallucinations and ignored temporal relationships, because the models default to processing rules sequentially rather than retrieving facts across the prompt.
These results indicate that current language models struggle with non-linear, back-and-forth reasoning, acting as greedy left-to-right readers rather than true logical reasoners. For organizations relying on automated AI reasoning, this creates significant reliability and compliance risks, as minor phrasing variations in documentation can silently trigger incorrect conclusions. The findings also demonstrate that providing few-shot demonstration examples does not effectively eliminate this order sensitivity.
Organizations deploying language models for complex analytical tasks should not assume model robustness to raw, unformatted inputs. Where feasible, engineering teams should implement preprocessing steps that structure and order relevant context before feeding prompts into models. Moving forward, AI developers must design novel model architectures, training objectives, and broader benchmarks that natively support robust reasoning over unordered information.
The conclusions are supported by extensive empirical testing across multiple commercial models and thousands of test cases. However, the study focuses on short-context problems (under 300 tokens) using synthetic logic and word problems, and it did not conduct a direct baseline study with human subjects. Decision-makers should exercise caution when extrapolating these specific error rates to highly domain-specific, unstructured long-context workflows without dedicated validation.
- Paper: Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought, Abulhair Saparov et al. (2023). Its formal analysis of LLMs’ greedy, sequential reasoning provides the conceptual groundwork for understanding why rearranged premises can derail otherwise valid deductions.
No sufficiently relevant recommendations were found.
