Premise Order Matters in Reasoning with Large Language Models

Xinyun ChenRyan A. ChiXuezhi WangDenny Zhou

article2024ICML72 citations

Demonstrates that large language models suffer performance drops of over 30% on deductive and mathematical reasoning tasks when premises are simply reordered, revealing an autoregressive positional bias and introducing the R-GSM benchmark to measure this vulnerability.

Listen

Large language models are increasingly deployed to automate critical reasoning tasks, from financial analysis to decision support. However, real-world information is rarely organized sequentially to match the logical order needed to reach a conclusion. The article investigates a key operational vulnerability: large language models are highly brittle to premise ordering, experiencing steep performance declines when factual statements or rules are rearranged, even when the underlying logic and truth value remain completely unchanged.

The article's main objective is to evaluate how premise ordering influences model reasoning performance across deductive logic and mathematical problem-solving tasks. It systematically measures whether varying the sequence of identical premises degrades the accuracy and proof validity of state-of-the-art language models.

To conduct this evaluation, the researchers tested leading industry models, including GPT-4-turbo, GPT-3.5-turbo, PaLM 2-L, and Gemini 1.0 Pro. The study generated a logical reasoning dataset containing 27,000 problem variants across varying proof lengths and levels of distracting information, using pseudowords to isolate logical reasoning from pretrained internal knowledge. In addition, the researchers developed R-GSM, a mathematical reasoning benchmark consisting of 220 problem pairs adapted from grade-school math word problems, in which the sentence order was reordered without altering the mathematical meaning or final answer.

The findings show that all evaluated models perform best when premises follow the exact sequence of intermediate reasoning steps, termed the forward order. Permuting this order causes severe performance drops exceeding 30% on complex logical problems for top-tier models and over 40% for weaker models. On mathematical reasoning tasks, models failed on 10% to over 35% of problems they had solved correctly in the original order. Introducing irrelevant or distracting premises further magnifies these accuracy drops. Error analyses revealed that out-of-order premises primarily lead to fact hallucinations and ignored temporal relationships, because the models default to processing rules sequentially rather than retrieving facts across the prompt.

These results indicate that current language models struggle with non-linear, back-and-forth reasoning, acting as greedy left-to-right readers rather than true logical reasoners. For organizations relying on automated AI reasoning, this creates significant reliability and compliance risks, as minor phrasing variations in documentation can silently trigger incorrect conclusions. The findings also demonstrate that providing few-shot demonstration examples does not effectively eliminate this order sensitivity.

Organizations deploying language models for complex analytical tasks should not assume model robustness to raw, unformatted inputs. Where feasible, engineering teams should implement preprocessing steps that structure and order relevant context before feeding prompts into models. Moving forward, AI developers must design novel model architectures, training objectives, and broader benchmarks that natively support robust reasoning over unordered information.

The conclusions are supported by extensive empirical testing across multiple commercial models and thousands of test cases. However, the study focuses on short-context problems (under 300 tokens) using synthetic logic and word problems, and it did not conduct a direct baseline study with human subjects. Decision-makers should exercise caution when extrapolating these specific error rates to highly domain-specific, unstructured long-context workflows without dedicated validation.

arXiv: 2402.08939

No sufficiently relevant recommendations were found.

Cover for Premise Order Matters in Reasoning with Large Language Models

Abstract

Large language models (LLMs) have accomplished remarkable reasoning performance in various domains. However, in the domain of reasoning tasks, we discover a frailty: LLMs are surprisingly brittle to the ordering of the premises, despite the fact that such ordering does not alter the underlying task. In particular, we observe that LLMs achieve the best performance when the premise order aligns with the context required in intermediate reasoning steps. For example, in deductive reasoning tasks, presenting the premises in the same order as the ground-truth proof in the prompt (as opposed to random ordering) drastically increases the model’s accuracy. We first examine the effect of premise ordering on deductive reasoning on a variety of LLMs, and our evaluation shows that even if the model performance is decent on the optimal order, permuting the premise order can cause a performance drop of over 30%. In addition, we release the benchmark R-GSM, based on GSM8K, to examine the ordering effect for mathematical problem-solving, and we again observe a significant drop in accuracy, relative to the original GSM8K benchmark.

Table of Contents

  • 1. Introduction
  • 2. Benchmarks
  • 2.1. Logical Reasoning
  • 2.2. R-GSM for Mathematical Reasoning
  • 3. Experiments
  • 3.1. Experimental Setup
  • 3.2. Logical Reasoning
  • 3.3. R-GSM for Mathematical Reasoning
  • 4. Related Work
  • 5. Conclusion
  • Impact Statement
  • Acknowledgment
  • References
  • A. R-GSM Dataset Statistics
  • B. Logical Reasoning Examples
  • C. R-GSM Examples
  • D. Discussion: Does Logical Reasoning Suffer from the Lost-in-the-middle Issue?
  • E. Full Results for Logical Reasoning
  • F. Full Results on R-GSM

Knowls

  1. Knowl 1 — Synthetic deductive benchmark isolates premise-order effects

    model/method

    The paper evaluates deductive reasoning using synthetic propositional-logic problems adapted from SimpleLogic. Each problem contains true facts, rules with one to three antecedents, and a conclusion known to be true. Predicates are pseudowords, reducing the possibility that a model can rely on familiar world knowledge. A response is correct only if it gives a completely valid step-by-step proof; unsupported facts or rules and incorrect refutations count as errors.

    Problems require 4–12 rules in the proof. For each proof length, the authors generate 200 problems and create variants across five premise orders and three quantities of irrelevant rules (0, 5, or 10), producing 27,000 problem variants. The evaluated models are GPT-4-turbo, GPT-3.5-turbo, PaLM 2-L, and Gemini 1.0 Pro. Decoding is greedy at temperature 0. Prompts are zero-shot; logical-reasoning prompts request a derivation identifying the premise used at each step.

  2. Knowl 2 — Premise order is measured by alignment with a forward proof

    definition

    For a logical-reasoning problem, the forward order presents the rules sequentially in the order they are applied in a ground-truth forward-chaining proof. The backward order reverses that order and corresponds to the proof direction of backward chaining. Other premise orders are categorized by normalized Kendall tau distance τ\tau from the forward order: τ=1\tau=1 is forward, τ=−1\tau=-1 is backward, and values near 00 indicate little rank correlation with the proof order. The experiments use τ∈{1,0.5,0,−0.5,−1}\tau\in\{1,0.5,0,-0.5,-1\}. Changing the order does not change the underlying logical task or its correct conclusion.

  3. Knowl 3 — Logical proof accuracy depends strongly on premise order

    empirical result

    With 12 rules required in the proof and no distracting rules, all four evaluated models achieve their highest accuracy in the forward order. The table compares forward, backward, and shuffled orders; shuffled accuracy aggregates results for τ=0.5,0,−0.5\tau=0.5,0,-0.5.

    Model Forward Backward Shuffled
    GPT-4-turbo 96.5% 84.0% 80.8%
    PaLM 2-L 88.0% 57.5% 66.5%
    Gemini 1.0 Pro 16.5% 0.5% 0.2%
    GPT-3.5-turbo 30.0% 1.0% 1.2%

    The effect becomes more consequential on longer proofs: across the benchmark’s range of 4–12 required rules, alternative orderings generally hurt more as the number of rules increases. For Gemini 1.0 Pro and GPT-3.5-turbo, the reported accuracy drop from forward to alternative orders can exceed 40 percentage points; for GPT-4-turbo and PaLM 2-L, drops reach roughly 20–30 points.

  4. Knowl 4 — Distracting rules further reduce accuracy and amplify order sensitivity

    empirical result

    For 12-rule proofs, adding irrelevant rules reduces accuracy for both GPT-4-turbo and PaLM 2-L and increases the gap between forward and alternative orders. The table reports accuracy for the forward order, backward order, and shuffled orders (aggregated over τ=0.5,0,−0.5\tau=0.5,0,-0.5).

    Model Distracting rules Forward Backward Shuffled
    GPT-4-turbo 0 96.5% 84.0% 80.8%
    GPT-4-turbo 5 80.5% 72.5% 57.2%
    GPT-4-turbo 10 57.5% 46.5% 40.0%
    PaLM 2-L 0 88.0% 57.5% 66.5%
    PaLM 2-L 5 64.0% 32.5% 38.2%
    PaLM 2-L 10 36.5% 15.5% 18.2%

    A position-control experiment with PaLM 2-L and 10 distracting rules found little difference when the same relevant rules were placed at the beginning, middle, or end. For example, with 12 proof rules, forward accuracy ranged from 35.0% to 36.5% across those placements. The longest logical inputs were under 300 tokens, so the authors conclude that a lost-in-the-middle effect is not the primary explanation for the ordering results in these experiments.

  5. Knowl 5 — Models differ in their preferences among non-forward orders

    empirical result

    Although the forward order is the best-performing order overall, the models do not respond alike to other orderings. GPT-4-turbo generally performs better on the backward order than on other non-forward orders. PaLM 2-L tends to perform worst on the backward order, and its accuracy generally declines as τ\tau decreases away from the forward order. Gemini 1.0 Pro and GPT-3.5-turbo show less consistent preferences, but more often favor backward ordering over other non-forward orders. Thus, random shuffling is not uniformly more difficult than reversing the proof order; the relative performance depends on the model.

  6. Knowl 6 — Logical-reasoning errors include hallucinated facts and rules

    empirical result

    The paper groups logical-reasoning mistakes into wrong refutations, hallucinated rules, and hallucinated facts (including facts not yet derived). In the 12-rule, no-distractor setting, the authors report that fact hallucination is typically the most common error type and generally becomes more frequent as premise order moves away from forward order. For example, the fact-hallucination rate from forward to backward order rises from 1.5% to 12.5% for GPT-4-turbo, from 8.5% to 30.0% for PaLM 2-L, from 50.5% to 60.5% for Gemini 1.0 Pro, and from 35.5% to 47.0% for GPT-3.5-turbo.

    The paper attributes this pattern to models following rules in their presented sequence: when the next encountered rule cannot yet be applied, a model may invent a fact to continue the proof. Wrong refutations are generally less frequent for the backward order than for intermediate orders.

  7. Knowl 7 — R-GSM tests sentence-order changes in grade-school math problems

    model/method

    R-GSM is a benchmark built from GSM8K test problems to study premise ordering in mathematical reasoning. The authors select problems with at least five description sentences and exclude problems for which changing sentence order cannot preserve a valid problem and the same answer. For each retained problem, they leave the final sentence in place and manually rewrite the ordering of the other sentences, allowing minor wording edits to retain grammaticality. Each rewrite is checked to preserve the ground-truth answer. The benchmark contains 220 original/reordered pairs; problems have 5–8 sentences and fewer than 200 tokens. In R-GSM evaluation, the model receives the problem description without additional instructions.

  8. Knowl 8 — Reordered R-GSM problems lower accuracy, including on initially solved items

    data/table

    All four models have lower accuracy on the manually reordered R-GSM problems than on the original descriptions. The conditional results measure accuracy on the subset of problems each model solved in the original form, so original accuracy in that subset is 100% by construction.

    Model Original Reordered Original, initially solved subset Reordered, same subset
    GPT-4-turbo 94.1% 85.0% 100% 89.9%
    PaLM 2-L 86.4% 79.5% 100% 87.9%
    Gemini 1.0 Pro 80.5% 69.1% 100% 74.6%
    GPT-3.5-turbo 67.3% 51.8% 100% 64.9%

    Thus, every model fails on some reordered problems that it had solved originally; the conditional decline is 10.1 percentage points for GPT-4-turbo, 12.1 for PaLM 2-L, 25.4 for Gemini 1.0 Pro, and 35.1 for GPT-3.5-turbo. The accuracy gap tends to grow with more reasoning steps and longer descriptions for GPT-4-turbo and Gemini 1.0 Pro; the gap is more similar across complexity levels for PaLM 2-L and GPT-3.5-turbo.

  9. Knowl 9 — R-GSM failures reflect sequential processing of reordered information

    empirical result

    In error cases where a model solved the original R-GSM problem but failed on its reordered version, the authors identify errors involving temporal order, unknown intermediate quantities, and other mistakes. Their analysis links these errors to treating statements or numbers in their presented sequence rather than tracking the dependencies needed by the solution. A reordered description can, for example, mention a quantity before the information needed to compute it, requiring the model to use later sentences first.

    Model Temporal Unknown Others
    GPT-4-turbo 45.0% 15.0% 40.0%
    GPT-3.5-turbo 21.6% 19.6% 58.8%
    PaLM 2-L 34.8% 4.3% 60.9%
    Gemini 1.0 Pro 29.5% 18.2% 52.3%

    Here, “Temporal” denotes errors involving event order, “Unknown” denotes errors involving unknown variables, and “Others” covers the remaining analyzed errors. The authors also find that few-shot chain-of-thought examples do not reliably remove the ordering effect. For GPT-3.5-turbo, five reordered examples improve accuracy to 79.1% on original problems and 66.4% on reordered problems, leaving a 13-percentage-point gap.

  10. Knowl 10 — The paper leaves causes, remedies, and human comparison unresolved

    limitation

    The study establishes an empirical ordering effect but does not provide a theoretical explanation or a general method for eliminating it. The authors suggest autoregressive design, training objectives, and training-data mixture as possible factors, not as established causes, and leave mitigation techniques for future work. They also did not conduct a rigorous human study on the benchmarks, so the reported model performance cannot be directly compared with human performance.

Coverage note — Detailed per-proof-length and per-order accuracy curves and the illustrated individual examples are omitted because they reinforce the benchmark-level results and error patterns captured here rather than adding a separate contribution.

References

  1. 1.Abdou, M., Ravishankar, V., Kulmizev, A., and Søgaard, A. Word order does matter and shuffled language models know it. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 6907–6919, 2022.
  2. 2.Austin, J., Odena, A., Nye, M., Bosma, M., Michalewski, H., Dohan, D., Jiang, E., Cai, C., Terry, M., Le, Q., et al. Program synthesis with large language models. arXiv preprint arXiv:2108.07732, 2021.
  3. 3.Berglund, L., Tong, M., Kaufmann, M., Balesni, M., Stickland, A. C., Korbak, T., and Evans, O. The reversal curse: Llms trained on” a is b” fail to learn” b is a”. arXiv preprint arXiv:2309.12288, 2023.
  4. 4.Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., et al. Sparks of artificial general intelligence: Early experiments with gpt-4. arXiv preprint arXiv:2303.12712, 2023.
  5. 5.Cao, Q., Kojima, T., Matsuo, Y., and Iwasawa, Y. Unnatural error correction: Gpt-4 can almost perfectly handle unnatural scrambled text. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 8898–8913, 2023.
  6. 6.Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. d. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374, 2021.
  7. 7.Cicirello, V. A. Kendall tau sequence distance: Extending kendall tau from ranks to sequences. arXiv preprint arXiv:1905.02752, 2019.
  8. 8.Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021.
  9. 9.Dekeyser, M., Schroyens, W., Schaeken, W., Spitaels, O., and d’Ydewalle, G. Preferred premise order in propositional reasoning: Semantic informativeness and co-reference. Deductive reasoning and strategies, pp. 73–95, 2000.
  10. 10.Ferreira, D. and Freitas, A. Premise selection in natural language mathematical texts. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 7365–7374, 2020.
  11. 11.Gemini. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805, 2023.
  12. 12.Girotto, V., Mazzocco, A., and Tasso, A. The effect of premise order in conditional reasoning: A test of the mental model theory. Cognition, 63(1):1–28, 1997.
  13. 13.Google. Palm 2 technical report. arXiv preprint arXiv:2305.10403, 2023.
  14. 14.Hagendorff, T., Fabi, S., and Kosinski, M. Human-like intuitive behavior and reasoning biases emerged in large language models but disappeared in chatgpt. Nature Computational Science, 3(10):833–838, 2023.
  15. 15.Han, S., Schoelkopf, H., Zhao, Y., Qi, Z., Riddell, M., Benson, L., Sun, L., Zubova, E., Qiao, Y., Burtell, M., et al. Folio: Natural language reasoning with first-order logic. arXiv preprint arXiv:2209.00840, 2022.
  16. 16.Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., and Steinhardt, J. Measuring mathematical problem solving with the math dataset. arXiv preprint arXiv:2103.03874, 2021.
  17. 17.Irving, G., Szegedy, C., Alemi, A. A., Een, N., Chollet, F., and Urban, J. Deepmath-deep sequence models for premise selection. Advances in neural information processing systems, 29, 2016.
  18. 18.Jones, E. and Steinhardt, J. Capturing failures of large language models via human cognitive biases. Advances in Neural Information Processing Systems, 35:11785–11799, 2022.
  19. 19.Lee, N., Sreenivasan, K., Lee, J. D., Lee, K., and Papailiopoulos, D. Teaching arithmetic to small transformers. arXiv preprint arXiv:2307.03381, 2023.
  20. 20.Li, Y., Choi, D., Chung, J., Kushman, N., Schrittwieser, J., Leblond, R., Eccles, T., Keeling, J., Gimeno, F., Dal Lago, A., et al. Competition-level code generation with alphacode. Science, 378(6624):1092–1097, 2022.
  21. 21.Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., and Liang, P. Lost in the middle: How language models use long contexts. Transactions of the Association for Computational Linguistics, 12:157–173, 2024.
  22. 22.McCoy, R. T., Yao, S., Friedman, D., Hardy, M., and Griffiths, T. L. Embers of autoregression: Understanding large language models through the problem they are trained to solve. arXiv preprint arXiv:2309.13638, 2023.
  23. 23.OpenAI. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023.
  24. 24.Saparov, A. and He, H. Language models are greedy reasoners: A systematic formal analysis of chain-of-thought. arXiv preprint arXiv:2210.01240, 2022.
  25. 25.Saparov, A., Pang, R. Y., Padmakumar, V., Joshi, N., Kazemi, S. M., Kim, N., and He, H. Testing the general deductive reasoning capacity of large language models using ood examples. arXiv preprint arXiv:2305.15269, 2023.
  26. 26.Sen, P. K. Estimates of the regression coefficient based on kendall’s tau. Journal of the American statistical association, 63(324):1379–1389, 1968.
  27. 27.Shi, F., Chen, X., Misra, K., Scales, N., Dohan, D., Chi, E. H., Scharli, N., and Zhou, D. Large language models can be easily distracted by irrelevant context. In International Conference on Machine Learning, pp. 31210–31227. PMLR, 2023.
  28. 28.Sinha, K., Parthasarathi, P., Pineau, J., and Williams, A. Unnatural language inference. arXiv preprint arXiv:2101.00010, 2020.
  29. 29.Wan, Y., Wang, W., Yang, Y., Yuan, Y., Huang, J.-t., He, P., Jiao, W., and Lyu, M. R. A & b== b & a: Triggering logical reasoning failures in large language models. arXiv preprint arXiv:2401.00757, 2024.
  30. 30.Wang, M., Tang, Y., Wang, J., and Deng, J. Premise selection for theorem proving by deep graph embedding. Advances in neural information processing systems, 30, 2017.
  31. 31.Wang, P., Li, L., Chen, L., Cai, Z., Zhu, D., Lin, B., Cao, Y., Liu, Q., Liu, T., and Sui, Z. Large language models are not fair evaluators. arXiv preprint arXiv:2305.17926, 2023.
  32. 32.Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35:24824–24837, 2022.
  33. 33.Xu, F., Lin, Q., Han, J., Zhao, T., Liu, J., and Cambria, E. Are large language models really good logical reasoners? a comprehensive evaluation from deductive, inductive and abductive views. arXiv preprint arXiv:2306.09841, 2023.
  34. 34.Yan, S., Shen, C., Liu, J., and Ye, J. Concise and organized perception facilitates large language models for deductive reasoning. arXiv preprint arXiv:2310.03309, 2023.
  35. 35.Zhang, H., Li, L. H., Meng, T., Chang, K.-W., and Broeck, G. V. d. On the paradox of learning to reason from data. arXiv preprint arXiv:2205.11502, 2022.
  36. 36.Zhou, H., Bradley, A., Littwin, E., Razin, N., Saremi, O., Susskind, J., Bengio, S., and Nakkiran, P. What algorithms can transformers learn? a study in length generalization. arXiv preprint arXiv:2310.16028, 2023.
  37. 37.Zhou, Y., Alon, U., Chen, X., Wang, X., Agarwal, R., and Zhou, D. Transformers can achieve length generalization but not robustly. arXiv preprint arXiv:2402.09371, 2024.
  38. 38.Zhu, Z., Xue, Y., Chen, X., Zhou, D., Tang, J., Schuurmans, D., and Dai, H. Large language models can learn rules. arXiv preprint arXiv:2310.07064, 2023.

Citation

MLA
Chen, X., et al. “Premise Order Matters in Reasoning with Large Language Models”. arXiv, 2024, http://arxiv.org/abs/2402.08939v3.
APA
Chen, X., Chi, R. A., Wang, X., & Zhou, D. (2024). Premise Order Matters in Reasoning with Large Language Models. arXiv. http://arxiv.org/abs/2402.08939v3
Chicago
Chen, X., R. A. Chi, X. Wang, and D. Zhou. 2024. “Premise Order Matters in Reasoning with Large Language Models”. arXiv. http://arxiv.org/abs/2402.08939v3.
Harvard
Chen, X. et al. (2024) “Premise Order Matters in Reasoning with Large Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.08939v3.
Vancouver
1. Chen X, Chi RA, Wang X, Zhou D (2024) Premise Order Matters in Reasoning with Large Language Models. arXiv

BibTeX

@article{chen2024premise,
  title = {Premise Order Matters in Reasoning with Large Language Models},
  author = {Chen, Xinyun and Chi, Ryan A. and Wang, Xuezhi and Zhou, Denny},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.08939v3},
  eprint = {2402.08939}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/