ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning

Bill Yuchen LinRonan Le BrasKyle RichardsonAshish SabharwalRadha PoovendranPeter ClarkYejin Choi

article2025ICML225 citations

Introduces the ZebraLogic benchmark to quantify the scaling limits of large language models on logic grid puzzles, revealing that current models suffer a severe collapse in reasoning accuracy as problem complexity grows unless inference-time compute and backtracking mechanisms are scaled.

Listen

Organizations increasingly deploy large language models to automate complex decision-making, resource scheduling, and operational planning. These workflows depend heavily on pure deductive logic rather than general factual recall. However, standard evaluations often conflate domain memorization with underlying reasoning capabilities, obscuring how reliably these models perform as problem difficulty scales.

The article evaluates the formal logical reasoning limits of leading artificial intelligence models across varying levels of problem complexity. Specifically, it assesses whether increasing model parameters, generating multiple candidate solutions, or allocating more test-time computational tokens can overcome performance bottlenecks in complex deductive reasoning.

To conduct this assessment in a controlled environment, the authors introduced ZebraLogic, an evaluation framework comprising 1,000 logic grid puzzles modeled as constraint satisfaction problems. The framework isolates pure logic from factual knowledge and systematically adjusts difficulty across 25 grid sizes. Problem complexity is measured through the mathematical search space size as well as the average number of logical conflicts identified by an automated theorem solver. The authors tested a broad range of open-weight and proprietary models, including OpenAI's o1 suite, DeepSeek-R1, GPT-4o, and Meta's Llama series, using a standardized, one-shot prompt template.

The investigation produced four primary findings. First, models suffer from a severe curse of complexity: accuracy collapses dramatically as search spaces exceed ten million possibilities or logical conflicts rise above 20. Second, simply scaling model parameters fails to solve this issue; the 405-billion-parameter Llama model achieved 81.3% accuracy on small problems but collapsed to 1.5% on large problems and 0.0% on extra-large puzzles. Third, specialized reasoning models that allocate hidden inference-time reasoning tokens significantly outperform standard architectures. OpenAI's o1 achieved an overall accuracy of 81.0% and DeepSeek-R1 reached 78.7%, compared to only 36.2% for Claude 3.5 Sonnet and 31.7% for GPT-4o. Fourth, while sampling multiple answers improves coverage under theoretical ideal selection, practical selection methods such as majority voting or standard reward models yield only modest gains, with majority voting reaching 38.0% on GPT-4o.

These findings indicate that relying on larger general-purpose models will not resolve failures in complex logic and constraint-driven planning. Systems that rely on straightforward forward-chaining deduction fail when problems require non-monotonic reasoning and counterfactual backtracking. Consequently, deploying standard foundation models in mission-critical scheduling or compliance environments introduces severe operational risks.

Strategic decision-makers should avoid relying solely on expanding base model size for complex reasoning tasks. Instead, organizations should prioritize models trained to perform explicit step-by-step reasoning with backtracking mechanisms, such as reinforcement-learning-driven reasoning architectures. Where standard language models must be deployed, teams should combine them with symbolic solvers or automated constraint checkers rather than relying on self-verification prompts or standard reward scoring.

Confidence in these findings is strong regarding grid-based constraint problems, but the evaluation relies on a synthetic puzzle framework that may not fully reflect informal, ill-defined business scenarios. Additionally, because the internal reasoning chains of proprietary models remain concealed, researchers must interpret their intermediate reasoning behavior based on high-level summaries and visible outputs rather than direct inspection.

Lin et al (2025).pdf
Cover for ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning

Table of Contents

  • Abstract
  • 1. Introduction
  • 2. Problem Formulation of Logical Reasoning
  • 2.1. Logic Grid Puzzles
  • 2.2. Problem Formulation
  • 2.3. ZebraLogic Dataset Creation
  • 2.4. Theoretical Problem Complexity
  • 2.5. Measuring Effective Instance Complexity
  • 3. Evaluation
  • 3.1. Main results
  • 3.2. Curse of Complexity in Reasoning with LLMs
  • 3.3. Scaling Behavior of LLMs in Logical Reasoning
  • 4. Scaling Model Size Can Hardly Break the Curse of Complexity in Reasoning
  • 5. Scaling Test-Time Compute with Repeated Sampling: Promises & Challenges
  • 6. Scaling Test-Time Compute with Extensive Chain-of-Thoughts Tokens
  • 6.1. o1 Generates More Hidden Reasoning Tokens
  • 6.2. Self-Refinement is Limited but Promising
  • 7. Related Work
  • 8. Conclusion
  • Impact Statement
  • Acknowledgments
  • References
  • A. Additional Experimental Results and Analysis
  • B. Details of the ZebraLogic Dataset
  • C. Additional Analysis
  • C.1. Human Evaluation of o1's Reasoning
  • C.2. Comparison with LMSYS Arena Rankings.
  • C.3. o1 generates large-scale hidden reasoning tokens.
  • D. Further Discussion on o1's Reasoning
  • D.1. Prompt template to evaluate ZebraLogic

Knowls

  1. Knowl 1 — ZebraLogic benchmark structure and dataset

    definition

    ZebraLogic is a benchmark of logic-grid puzzles in which NN houses are arranged from left to right and each of MM attributes takes one of NN distinct values, with every value used exactly once for that attribute. A puzzle supplies clues relating attribute values or their positions; the solver must recover the unique assignment of values to houses. The benchmark contains 1,000 puzzles: 40 for each of the 25 grid sizes with N,M∈{2,3,4,5,6}N,M\in\{2,3,4,5,6\}. Puzzles have an average of 10.4 clues and a median of 9. The attribute pool contains Name, Color, Nationality, Animal, Drink, Cigar, Food, Flower, PhoneModel, Children, Smoothie, Birthday, Occupation, Height, CarModel, FavoriteSport, MusicGenre, BookGenre, HairColor, Mother, HouseStyle, Education, Hobby, Vacation, and Pet. Name is included in every puzzle; the pool provides at least six possible values for each attribute. Clue families include fixed-position (FOUNDAT), same-house and different-house relations (SAMEHOUSE and NOTAT), direct-left/right, side-by-side, left/right ordering, and one- or two-house-between relations.

  2. Knowl 2 — Constraint-satisfaction formulation of a logic-grid puzzle

    model/method

    Represent a puzzle with houses H={1,…,N}H=\{1,\ldots,N\}, an attribute set AA of size MM, and a value domain VaV_a of size NN for each attribute a∈Aa\in A. Let xa,k∈Vax_{a,k}\in V_a denote the value of attribute aa assigned to house kk. The uniqueness constraint for each attribute is {xa,k:k∈H}=Va\{x_{a,k}:k\in H\}=V_a, so each value appears in exactly one house. Translate every verbal clue into a logical constraint on these variables, including positional constraints such as direct-left relations, which only apply where a house to the right exists. A valid solution satisfies all uniqueness and clue constraints; ZebraLogic puzzles are generated to have a unique valid assignment.

  3. Knowl 3 — Puzzle generation by clue removal while preserving uniqueness

    algorithm

    The generator first samples MM attributes, creates a random complete assignment of their values to NN houses, and generates clues consistent with that assignment using the available clue types and language templates. It then attempts to remove clues one at a time. For a candidate clue pp, it checks with a SAT solver whether the remaining clues still admit exactly one solution, namely the original assignment. If so, pp is removed; otherwise it is retained. The process continues until no further clue can be removed without losing uniqueness, yielding a puzzle with a locally minimal clue set. Clue selection uses weighted sampling that favors simpler clue types over harder ones; the paper does not report the numerical weights. The output is the complete assignment and the retained clues.

  4. Knowl 4 — NP-completeness of ZebraLogic solving

    theoretical result

    The paper establishes that solving ZebraLogic puzzles is NP-complete by reduction from the Quasigroup (Latin-square) Completion Problem. The result already holds when the clue language includes the FOUNDAT and NOTAT clue types. A proposed assignment can be checked against the uniqueness and clue constraints, but finding a solution can become computationally intractable as instances grow. The paper consequently notes that, for a fixed-size language model, the reasoning-token requirement may grow exponentially with puzzle size; this is an implication drawn by the authors, not a separately established bound on language models.

  5. Knowl 5 — Two complementary measures of puzzle complexity

    definition

    For a puzzle with NN houses and MM attributes, the number of assignments satisfying only the per-attribute uniqueness constraints is ∣S∣=(N!)M|S|=(N!)^M. ZebraLogic groups instances by this search-space size: Small when ∣S∣<103|S|<10^3, Medium when 103≤∣S∣<10610^3\leq |S|<10^6, Large when 106≤∣S∣<101010^6\leq |S|<10^{10}, and X-Large when ∣S∣≥1010|S|\geq 10^{10}. As a complementary, solver-dependent measure, the benchmark runs the Z3 SMT solver 32 times per puzzle and uses the average number of conflicts. A conflict is a contradiction in the solver's current assignment that prompts backtracking. Search-space size counts candidate configurations, whereas the conflict count estimates the backtracking burden encountered by Z3; zero-conflict instances are generally amenable to simple forward chaining.

  6. Knowl 6 — Evaluation protocol and performance across models

    empirical result

    Models receive a one-shot example puzzle and are prompted to return reasoning and a solution in JSON. The same prompts, greedy decoding, and parsing procedure are used across models; because o1 does not support only greedy decoding, it is run three times and the best result is retained. Puzzle-level accuracy requires every cell in the grid to be correct; cell-level accuracy counts individual cells. The table gives reported accuracies (%) for representative models, with complexity-group columns in the benchmark's reported order.

    Model Overall Small Medium Large X-Large Cell-level
    o1-full 81.0 97.2 92.1 78.0 42.5 78.7
    DeepSeek-R1 78.7 98.4 95.7 73.5 28.5 80.5
    o1-preview 71.4 98.1 88.2 59.5 17.0 75.1
    o1-mini 59.7 87.5 76.8 39.0 12.0 70.3
    Claude Sonnet 3.5 36.2 84.7 28.9 4.0 1.0 54.3
    Llama-3.1-405B 32.6 81.3 22.5 1.5 0.0 45.8
    GPT-4o 31.7 80.0 19.6 2.5 0.5 50.3

    The reasoning-focused models o1-full and DeepSeek-R1 lead overall, while even the strongest conventional model in this comparison, Claude Sonnet 3.5, solves only 4.0% of Large and 1.0% of X-Large puzzles. The table demonstrates that high accuracy on smaller grids does not carry over to the most complex groups.

  7. Knowl 7 — Model scaling does not remove the complexity-related accuracy collapse

    empirical result

    Across the evaluated Llama model sizes, accuracy falls rapidly as puzzle complexity rises, whether complexity is indexed by search-space size or by average Z3 conflicts. Increasing model size helps on relatively small search spaces, including instances with ∣S∣≤106|S|\leq 10^6, but its benefit diminishes on harder instances. For example, Llama-3.1-405B achieves 22.5% accuracy on Medium puzzles but only 1.5% on Large and 0.0% on X-Large puzzles. More generally, the paper reports that most models struggle when the search space exceeds roughly 10710^7 possibilities or the average Z3 conflict count exceeds 20. The authors call this performance collapse the “curse of complexity”; it persists despite substantial increases in parameter count.

  8. Knowl 8 — Repeated sampling: oracle coverage rises more than practical selection accuracy

    empirical result

    The experiments sampled up to 128 candidate solutions per puzzle from GPT-4o and GPT-4o-mini. BoN-Oracle selects a correct candidate whenever one occurs in the sample pool, so it is the oracle-selection upper bound (pass@N), not a deployable selection rule. Majority voting scores each candidate by the summed frequency of its cell assignments and selects the highest-scoring candidate. BoN-RM selects with Skywork-Reward-Llama-3.1-8B-v0.2. The accuracies below are puzzle-level percentages, ordered Overall, Small, Medium, Large, and X-Large.

    Method Overall Small Medium Large X-Large
    GPT-4o baseline 31.7 80.0 19.6 2.5 0.5
    GPT-4o BoN-Oracle, N=32N=32 60.3 98.4 81.1 28.0 2.5
    GPT-4o BoN-Oracle, N=128N=128 69.1 99.7 92.9 49.0 7.0
    GPT-4o Majority Voting, N=32N=32 38.0 84.1 34.3 7.0 0.5
    GPT-4o Majority Voting, N=128N=128 37.6 84.7 32.1 7.5 0.0
    GPT-4o BoN-RM, N=32N=32 33.9 77.8 28.9 4.5 0.0
    GPT-4o-mini baseline 20.1 58.8 4.6 0.0 0.0
    GPT-4o-mini BoN-Oracle, N=128N=128 51.2 99.7 61.8 10.0 0.0
    GPT-4o-mini Majority Voting, N=128N=128 25.0 69.4 8.9 1.5 0.0

    Oracle coverage can make GPT-4o competitive with reasoning models overall, but practical majority voting yields much smaller gains and does not consistently improve with more samples. Reward-model selection also gives little improvement over the baseline. Even oracle selection at 128 samples leaves GPT-4o at 7.0% on X-Large puzzles, so repeated sampling does not eliminate the complexity-related failure.

  9. Knowl 9 — o1 allocates substantially more hidden reasoning tokens, but scaling plateaus

    empirical result

    On average, o1-mini generates 5,144.6 hidden reasoning tokens per puzzle and o1-preview 5,346.3, compared with 502.9 for GPT-4o-mini and 543.7 for GPT-4o. The o1 models' hidden-token counts generally rise with puzzle complexity, measured by Z3 conflicts. For puzzles with fewer than 20 conflicts, o1-preview produces approximately 400 hidden reasoning tokens per conflict; the relationship plateaus once conflict counts exceed 30, suggesting a limit to the model's inference-time reasoning capacity at the tested scale. o1-full also averages about 5,000 hidden reasoning tokens. The paper reports that o1-preview often uses more hidden tokens on puzzles it gets wrong than on puzzles it solves, consistent with harder instances consuming more reasoning without guaranteeing success. Human inspection further found that o1's visible explanations on large puzzles are often incomplete and can contain unsupported or incorrect steps, so the visible text does not reliably expose the process behind a correct answer.

  10. Knowl 10 — Self-verification prompts yield only small accuracy gains

    empirical result

    The study tested multi-turn self-verification by asking a model to review its initial solution, recheck the clues, and correct errors. For GPT-4o, this prompt without access to the correct answer raises overall accuracy from 31.7% to 33.0%; applying the prompt a second time lowers it to 32.1%. With oracle feedback indicating whether the initial answer was correct, accuracy is 34.8%. For GPT-4o-mini, self-verification without an oracle raises accuracy from 20.1% to 21.1%, and oracle-assisted verification reaches 22.3%. These results show modest gains, substantially smaller than oracle selection from a large candidate pool, and repeated reflection does not guarantee further improvement.

Coverage note — Detailed puzzle-by-puzzle case studies of o1 and secondary appendix analyses are omitted; their general findings about hidden-token scaling and incomplete visible explanations are captured without reproducing individual examples.

References

  1. 1.AI@Meta. The llama 3 herd of models. 2024. URL https://arxiv.org/abs/2407.21783.
  2. 2.Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. Language models are few-shot learners, 2020.
  3. 3.Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., Nori, H., Palangi, H., Ribeiro, M. T., and Zhang, Y. Sparks of artificial general intelligence: Early experiments with gpt-4. arXiv preprint arXiv:2303.12712, 2023.
  4. 4.Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., Schuh, P., Shi, K., Tsvyashchenko, S., Maynez, J., Rao, A., Barnes, P., Tay, Y., Shazeer, N. M., Prabhakaran, V., Reif, E., Du, N., Hutchinson, B., Pope, R., Bradbury, J., Austin, J., Isard, M., Gur-Ari, G., Yin, P., Duke, T., Levskaya, A., Ghemawat, S., Dev, S., Michalewski, H., García, X., Misra, V., Robinson, K., Fedus, L., Zhou, D., Ippolito, D., Luan, D., Lim, H., Zoph, B., Spiridonov, A., Sepassi, R., Dohan, D., Agrawal, S., Omernick, M., Dai, A. M., Pillai, T. S., Pellat, M., Lewkowycz, A., Moreira, E., Child, R., Polozov, O., Lee, K., Zhou, Z., Wang, X., Saeta, B., Díaz, M., Firat, O., Catasta, M., Wei, J., Meier-Hellstern, K. S., Eck, D., Dean, J., Petrov, S., and Fiedel, N. Palm: Scaling language modeling with pathways. ArXiv, abs/2204.02311, 2022. URL https://api.semanticscholar.org/CorpusID:247951931.
  5. 5.Clark, P., Tafjord, O., and Richardson, K. Transformers as soft reasoners over language. In International Joint Conference on Artificial Intelligence, 2020. URL https://api.semanticscholar.org/CorpusID:211126663.
  6. 6.Colbourn, C. J. The complexity of completing partial latin squares. Discret. Appl. Math., 8:25–30, 1984. URL https://api.semanticscholar.org/CorpusID:33813226.
  7. 7.de Moura, L. M. and Bjørner, N. S. Z3: An efficient smt solver. In International Conference on Tools and Algorithms for Construction and Analysis of Systems, 2008. URL https://api.semanticscholar.org/CorpusID:15912959.
  8. 8.Dechter, R. Constraint Processing. Morgan Kaufmann, 2003.
  9. 9.DeepSeek-AI. DeepSeek-R1: Incentivizing reasoning capability in LLMs via reinforcement learning. arXiv, abs/2501.12948, 2025.
  10. 10.Dziri, N., Lu, X., Sclar, M., Li, X. L., Jian, L., Lin, B. Y., West, P., Bhagavatula, C., Le Bras, R., Hwang, J. D., Sanyal, S., Welleck, S., Ren, X., Ettinger, A., Harchaoui, Z., and Choi, Y. Faith and fate: Limits of transformers on compositionality. ArXiv, abs/2305.18654, 2023. URL https://api.semanticscholar.org/CorpusID:258967391.
  11. 11.Gomes, C. P. and Shmoys, D. B. Completing quasigroups or latin squares: A structured graph coloring problem. 2002. URL https://api.semanticscholar.org/CorpusID:10410543.
  12. 12.Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., and Steinhardt, J. Measuring mathematical problem solving with the math dataset. NeurIPS, 2021.
  13. 13.Jain, N., Han, K., Gu, A., Li, W.-D., Yan, F., Zhang, T., Wang, S., Solar-Lezama, A., Sen, K., and Stoica, I. Livecodebench: Holistic and contamination free evaluation of large language models for code. arXiv preprint arXiv:2403.07974, 2024.
  14. 14.Lam, L. H. M., Thatikonda, R. K., and Shareghi, E. A closer look at logical reasoning with llms: The choice of tool matters. 2024. URL https://api.semanticscholar.org/CorpusID:270219176.
  15. 15.Lambert, N., Morrison, J. D., Pyatkin, V., Huang, S., Ivison, H., Brahman, F., Miranda, L. J. V., Liu, A., Dziri, N., Lyu, X., Gu, Y., Malik, S., Graf, V., Hwang, J. D., Yang, J., Le Bras, R., Tafjord, O., Wilhelm, C., Soldaini, L., Smith, N. A., Wang, Y., Dasigi, P., and Hajishirzi, H. Tulu 3: Pushing frontiers in open language model post-training. ArXiv, abs/2411.15124, 2024a. URL https://api.semanticscholar.org/CorpusID:274192505.
  16. 16.Lambert, N., Pyatkin, V., Morrison, J. D., Miranda, L. J. V., Lin, B. Y., Chandu, K. R., Dziri, N., Kumar, S., Zick, T., Choi, Y., Smith, N. A., and Hajishirzi, H. Rewardbench: Evaluating reward models for language modeling. ArXiv, abs/2403.13787, 2024b. URL https://api.semanticscholar.org/CorpusID:268537409.
  17. 17.Liu, C. Y., Zeng, L., Liu, J., Yan, R., He, J., Wang, C., Yan, S., Liu, Y., and Zhou, Y. Skywork-reward: Bag of tricks for reward modeling in llms. arXiv preprint arXiv:2410.18451, 2024.
  18. 18.Liu, H., Liu, J., Cui, L., Teng, Z., Duan, N., Zhou, M., and Zhang, Y. Logiqa 2.0—an improved dataset for logical reasoning in natural language understanding. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 31:2947–2962, 2023. URL https://api.semanticscholar.org/CorpusID:259515154.
  19. 19.Liu, J., Cui, L., Liu, H., Huang, D., Wang, Y., and Zhang, Y. Logiqa: A challenge dataset for machine reading comprehension with logical reasoning. ArXiv, abs/2007.08124, 2020. URL https://api.semanticscholar.org/CorpusID:220483148.
  20. 20.Madusanka, T., Batista-navarro, R., and Pratt-hartmann, I. Identifying the limits of transformers when performing model-checking with natural language. In Vlachos, A. and Augenstein, I. (eds.), Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, pp. 3539–3550, Dubrovnik, Croatia, May 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.eacl-main.257. URL https://aclanthology.org/2023.eacl-main.257.
  21. 21.Madusanka, T., Pratt-Hartmann, I., and Batista-Navarro, R. Natural language satisfiability: Exploring the problem distribution and evaluating transformer-based language models. In Ku, L.-W., Martins, A., and Srikumar, V. (eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 15278–15294, Bangkok, Thailand, August 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.acl-long.815. URL https://aclanthology.org/2024.acl-long.815.
  22. 22.Mitra, A. and Baral, C. Learning to automatically solve logic grid puzzles. In Conference on Empirical Methods in Natural Language Processing, 2015. URL https://api.semanticscholar.org/CorpusID:2684696.
  23. 23.OpenAI. Openai o1 system card. arXiv, abs/2412.16720, 2024.
  24. 24.Pan, L., Ganesh, V., Abernethy, J., Esposo, C., and Lee, W. Can transformers reason logically? a study in sat solving. 2024. URL https://api.semanticscholar.org/CorpusID:273234137.
  25. 25.Parmar, M., Patel, N., Varshney, N., Nakamura, M., Luo, M., Mashetty, S., Mitra, A., and Baral, C. Logicbench: Towards systematic evaluation of logical reasoning ability of large language models. In Annual Meeting of the Association for Computational Linguistics, 2024. URL https://api.semanticscholar.org/CorpusID:269330143.
  26. 26.Prosser, P. Hybrid algorithms for the constraint satisfaction problem. Computational Intelligence, 9, 1993. URL https://api.semanticscholar.org/CorpusID:36951414.
  27. 27.Richardson, K. and Sabharwal, A. Pushing the limits of rule reasoning in transformers through natural language satisfiability. In AAAI Conference on Artificial Intelligence, 2022. URL https://api.semanticscholar.org/CorpusID:245219217.
  28. 28.Ryu, H., Kim, G., Lee, H. S., and Yang, E. Divide and translate: Compositional first-order logic translation and verification for complex logical reasoning. 2024. URL https://api.semanticscholar.org/CorpusID:273233577.
  29. 29.Schlegel, V., Pavlov, K. V., and Pratt-Hartmann, I. Can transformers reason in fragments of natural language? In Conference on Empirical Methods in Natural Language Processing, 2022. URL https://api.semanticscholar.org/CorpusID:253446947.
  30. 30.Sempolinski, P. Automatic solutions of logic puzzles. 2009. URL https://api.semanticscholar.org/CorpusID:125304065.
  31. 31.Tyagi, N., Parmar, M., Kulkarni, M., Rrv, A., Patel, N., Nakamura, M., Mitra, A., and Baral, C. Step-by-step reasoning to solve grid puzzles: Where do llms falter? In Conference on Empirical Methods in Natural Language Processing, 2024. URL https://api.semanticscholar.org/CorpusID:271329041.
  32. 32.Wei, J., Wang, X., Schuurmans, D., Bosma, M., Chi, E. H., Xia, F., Le, Q., and Zhou, D. Chain of thought prompting elicits reasoning in large language models. ArXiv, abs/2201.11903, 2022. URL https://api.semanticscholar.org/CorpusID:246411621.
  33. 33.Xie, C., Huang, Y., Zhang, C., Yu, D., Chen, X., Lin, B. Y., Li, B., Ghazi, B., and Kumar, R. On memorization of large language models in logical reasoning. 2024. URL https://api.semanticscholar.org/CorpusID:273695832.
  34. 34.Yan, J., Wang, C., Huang, J., and Zhang, W. Do large language models understand logic or just mimick context? ArXiv, abs/2402.12091, 2024. URL https://api.semanticscholar.org/CorpusID:267751049.

Citation

MLA
Lin, B. Y., et al. “ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning”. arXiv, 2025, https://doi.org/10.48550/arxiv.2502.01100.
APA
Lin, B. Y., Bras, R. L., Richardson, K., Sabharwal, A., Poovendran, R., Clark, P., & Choi, Y. (2025). ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning. arXiv. https://doi.org/10.48550/arxiv.2502.01100
Chicago
Lin, B. Y., R. L. Bras, K. Richardson, et al. 2025. “ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.2502.01100.
Harvard
Lin, B.Y. et al. (2025) “ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning”. arXiv. Available at: https://doi.org/10.48550/arxiv.2502.01100.
Vancouver
1. Lin BY, Bras RL, Richardson K, Sabharwal A, Poovendran R, Clark P, Choi Y (2025) ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning. https://doi.org/10.48550/arxiv.2502.01100

BibTeX

@misc{https://doi.org/10.48550/arxiv.2502.01100,
  doi = {10.48550/ARXIV.2502.01100},
  url = {https://arxiv.org/abs/2502.01100},
  author = {Lin, Bill Yuchen and Bras, Ronan Le and Richardson, Kyle and Sabharwal, Ashish and Poovendran, Radha and Clark, Peter and Choi, Yejin},
  keywords = {Artificial Intelligence (cs.AI), Computation and Language (cs.CL), Machine Learning (cs.LG), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning},
  publisher = {arXiv},
  year = {2025},
  copyright = {Creative Commons Attribution 4.0 International}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/