Puzzle Solving using Reasoning of Large Language Models: A Survey
Panagiotis GiadikiaroglouMaria LymperaiouGiorgos FilandrianosGiorgos Stamou
Categorizes puzzle-solving benchmarks into rule-based and rule-less domains to systematically evaluate large language models across prompting, neuro-symbolic, and fine-tuning strategies while identifying current gaps in automated logical reasoning.
As organizations deploy large language models to automate complex decision-making, understanding their actual logical reasoning and strategic capabilities is critical. Standard benchmarks often fail to isolate whether these models genuinely reason through unfamiliar situations or simply retrieve memorized training data. Puzzles serve as a rigorous testing ground to evaluate core cognitive functions, including deduction, lateral thinking, spatial orientation, and planning under uncertainty. The article systematically reviews how language models perform on text-based puzzles and evaluates the prompting, translation, and fine-tuning strategies developed to improve their reasoning.
To conduct this evaluation, the article introduces a functional framework that divides text-based challenges into rule-based puzzles—which feature formal constraints and closed environments, either deterministic (such as Sudoku or Rubik's Cube) or stochastic (such as Minesweeper or card games)—and rule-less puzzles, which require lateral thinking and contextual interpretation (such as riddles, programming snippets, and detective-style commonsense problems). The review synthesizes empirical findings across recent benchmarks, datasets, and methodologies published over a four-year period, comparing prompting frameworks, neuro-symbolic solvers, and specialized model fine-tuning against traditional computational algorithms.
The analysis reveals several core findings. First, advanced structured prompting methods, such as Tree-of-Thoughts and Everything-of-Thoughts, substantially outperform basic sequential reasoning prompts on deterministic puzzles, raising success rates by 26% to 70% and reaching up to 93.2% accuracy on spatial benchmarks like the 8-puzzle. Second, neuro-symbolic translation delivers the highest accuracy on formal logic tasks; translating natural-language rules into Answer Set Programming allowed GPT-4 to achieve 92% accuracy, compared to just 7% under standard few-shot prompting. Third, a substantial performance gap remains between language models and human reasoning on rule-less puzzles, particularly those involving metaphors, counterfactual deduction, and code analysis. Fourth, models struggle heavily in stochastic environments; in games like Minesweeper and poker, they fail to complete full boards or sustain multi-step risk calculations despite grasping isolated rules. Finally, task formatting heavily biases performance, with multiple-choice structures artificially inflating accuracy compared to open-ended, free-text formulations.
These results demonstrate that organizations cannot rely on base language models for mission-critical tasks requiring multi-step deduction, complete verification, or risk management under incomplete information. While traditional deterministic algorithms guarantee complete solutions with transparent execution, standalone language models remain unpredictable and prone to logical breakdown. Advanced prompting approaches mitigate some failures but introduce operational trade-offs, significantly increasing computational latency and invocation costs. Fine-tuning improves domain-specific execution but exhibits narrow transferability across distinct reasoning tasks.
Decision-makers implementing reasoning pipelines should adopt hybrid neuro-symbolic architectures that utilize language models to interpret natural language specifications and formal external solvers to verify logic and execute decisions. Research and development teams should prioritize constructing robust benchmarks for stochastic puzzles and code-translation tasks, which remain underrepresented in current evaluations. Further pilot testing is necessary before deploying purely generative models in complex reasoning workflows.
These findings are based on empirical literature spanning four years, and confidence in the broad performance trends is high. However, readers should note that the field is evolving rapidly, and newer model architectures or alternative prompting designs may alter specific accuracy baselines across individual puzzle benchmarks.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). Its foundational chain-of-thought results establish the prompting baseline that the survey compares with more structured puzzle-solving strategies.
- Paper: Tree of Thoughts: Deliberate Problem Solving with Large Language Models, Shunyu Yao et al. (2023). Tree of Thoughts introduces the branching search approach behind the survey’s discussion of structured exploration and backtracking on puzzles.
- Paper: PAL: Program-aided Language Models, Luyu Gao et al. (2023). PAL shows how translating reasoning into executable programs delegates exact calculation to an external interpreter, clarifying the survey’s neuro-symbolic comparisons.
- Paper: LINC: A Neurosymbolic Approach for Logical Reasoning by Combining Language Models with First-Order Logic Provers, Theo Olausson et al. (2023). LINC demonstrates the formal-logic-and-prover pipeline that helps explain the survey’s reported gains from symbolic verification over prompting alone.
- Paper: Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought, Abulhair Saparov et al. (2023). Its formal analysis of chain-of-thought exposes the gap between locally valid deductions and complete proof planning that informs the survey’s account of reasoning failures.
- Paper: Algorithm of Thoughts: Enhancing Exploration of Ideas in Large Language Models, Bilgehan Sel et al. (2024). Algorithm of Thoughts develops in-context search for puzzle-like tasks, providing a concrete precursor to the survey’s comparison of advanced reasoning prompts.
- Paper: ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning, Bill Yuchen Lin et al. (2025). ZebraLogic extends the survey’s puzzle-based evaluation agenda by measuring how deductive performance collapses as constraint-grid complexity scales.
- Paper: Steering Large Language Models between Code Execution and Textual Reasoning, Yongchao Chen et al. (2025). This study carries the survey’s hybrid-reasoning implications forward by testing when language models should hand precise tasks to code execution rather than reason in text.
- Paper: GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models, Iman Mirzadeh et al. (2025). GSM-Symbolic extends the survey’s concern about benchmark validity with controlled variants that reveal instability and possible memorization in mathematical reasoning.
- Paper: Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Chain-of-Thought, Violet Xiang et al. (2025). Meta Chain-of-Thought continues the survey’s account of search-based reasoning by training models to explore, backtrack, and verify rather than rely on a single reasoning path.
