Puzzle Solving using Reasoning of Large Language Models: A Survey

Panagiotis GiadikiaroglouMaria LymperaiouGiorgos FilandrianosGiorgos Stamou

article2024EMNLP64 citations

Categorizes puzzle-solving benchmarks into rule-based and rule-less domains to systematically evaluate large language models across prompting, neuro-symbolic, and fine-tuning strategies while identifying current gaps in automated logical reasoning.

Listen

As organizations deploy large language models to automate complex decision-making, understanding their actual logical reasoning and strategic capabilities is critical. Standard benchmarks often fail to isolate whether these models genuinely reason through unfamiliar situations or simply retrieve memorized training data. Puzzles serve as a rigorous testing ground to evaluate core cognitive functions, including deduction, lateral thinking, spatial orientation, and planning under uncertainty. The article systematically reviews how language models perform on text-based puzzles and evaluates the prompting, translation, and fine-tuning strategies developed to improve their reasoning.

To conduct this evaluation, the article introduces a functional framework that divides text-based challenges into rule-based puzzles—which feature formal constraints and closed environments, either deterministic (such as Sudoku or Rubik's Cube) or stochastic (such as Minesweeper or card games)—and rule-less puzzles, which require lateral thinking and contextual interpretation (such as riddles, programming snippets, and detective-style commonsense problems). The review synthesizes empirical findings across recent benchmarks, datasets, and methodologies published over a four-year period, comparing prompting frameworks, neuro-symbolic solvers, and specialized model fine-tuning against traditional computational algorithms.

The analysis reveals several core findings. First, advanced structured prompting methods, such as Tree-of-Thoughts and Everything-of-Thoughts, substantially outperform basic sequential reasoning prompts on deterministic puzzles, raising success rates by 26% to 70% and reaching up to 93.2% accuracy on spatial benchmarks like the 8-puzzle. Second, neuro-symbolic translation delivers the highest accuracy on formal logic tasks; translating natural-language rules into Answer Set Programming allowed GPT-4 to achieve 92% accuracy, compared to just 7% under standard few-shot prompting. Third, a substantial performance gap remains between language models and human reasoning on rule-less puzzles, particularly those involving metaphors, counterfactual deduction, and code analysis. Fourth, models struggle heavily in stochastic environments; in games like Minesweeper and poker, they fail to complete full boards or sustain multi-step risk calculations despite grasping isolated rules. Finally, task formatting heavily biases performance, with multiple-choice structures artificially inflating accuracy compared to open-ended, free-text formulations.

These results demonstrate that organizations cannot rely on base language models for mission-critical tasks requiring multi-step deduction, complete verification, or risk management under incomplete information. While traditional deterministic algorithms guarantee complete solutions with transparent execution, standalone language models remain unpredictable and prone to logical breakdown. Advanced prompting approaches mitigate some failures but introduce operational trade-offs, significantly increasing computational latency and invocation costs. Fine-tuning improves domain-specific execution but exhibits narrow transferability across distinct reasoning tasks.

Decision-makers implementing reasoning pipelines should adopt hybrid neuro-symbolic architectures that utilize language models to interpret natural language specifications and formal external solvers to verify logic and execute decisions. Research and development teams should prioritize constructing robust benchmarks for stochastic puzzles and code-translation tasks, which remain underrepresented in current evaluations. Further pilot testing is necessary before deploying purely generative models in complex reasoning workflows.

These findings are based on empirical literature spanning four years, and confidence in the broad performance trends is high. However, readers should note that the field is evolving rapidly, and newer model architectures or alternative prompting designs may alter specific accuracy baselines across individual puzzle benchmarks.

Cover for Puzzle Solving using Reasoning of Large Language Models: A Survey

Abstract

Exploring the capabilities of Large Language Models (LLMs) in puzzle solving unveils critical insights into their potential and challenges in AI, marking a significant step towards understanding their applicability in complex reasoning tasks. This survey leverages a unique taxonomy—dividing puzzles into rule-based and rule-less categories—to critically assess LLMs through various methodologies, including prompting techniques, neuro-symbolic approaches, and fine-tuning. Through a critical review of relevant datasets and benchmarks, we assess LLMs’ performance, identifying significant challenges in complex puzzle scenarios. Our findings highlight the disparity between LLM capabilities and human-like reasoning, particularly in those requiring advanced logical inference. The survey underscores the necessity for novel strategies and richer datasets to advance LLMs’ puzzle-solving proficiency and contribute to AI’s logical reasoning and creative problem-solving advancements.

Table of Contents

  • 1 Introduction
  • 2 Categorization of Puzzle Problems
  • 2.1 Rule-based Puzzles
  • 2.2 Rule-less Puzzles
  • 3 Methods and Strategies
  • 3.1 Prompting Methods
  • 3.2 Puzzle Translation
  • 3.3 Fine-Tuning
  • 4 Datasets, Benchmarks and Tasks
  • 4.1 Rule-based Puzzles
  • 4.1.1 Deterministic Puzzles
  • 4.1.2 Stochastic Puzzles
  • 4.2 Rule-less Puzzles
  • 4.2.1 Riddles
  • 4.2.2 Programming Puzzles
  • 4.2.3 Commonsense Reasoning Puzzles
  • 5 Discussion and Future Directions
  • 6 Conclusion
  • Limitations
  • Acknowledgments
  • References
  • A Appendix
  • A.1 Prompting Topologies
  • A.2 Conventional Methods
  • A.3 Tables

Knowls

  1. Knowl 1 — Taxonomy of text-based puzzles for evaluating LLM reasoning

    definition

    The survey defines a puzzle as a problem that tests cognitive abilities such as logical reasoning, spatial cognition, or creative thinking by asking a solver to detect patterns, deduce from available information, and combine insights to reach a solution. Its taxonomy separates puzzles according to whether solving depends on explicit formal rules or on flexible inference and world knowledge:

    • Rule-based puzzles specify victory conditions, legal moves, or state-transition rules. Deterministic puzzles have a uniquely determined next state for a given current state and action; examples include Sudoku, mazes, Rubik’s Cube, and chess. Stochastic puzzles involve randomness or hidden information, so actions can have uncertain outcomes; examples include Minesweeper, card games, and social-deduction games.
    • Rule-less puzzles rely more on contextual interpretation, general knowledge, and inference than on a formal move system. The survey groups these into riddles (wordplay and metaphor), programming puzzles (analyze code or find an input that makes a program meet a condition), and commonsense-reasoning puzzles (infer unstated causes, assumptions, or events in a situation).

    The distinction is about the demands of solving, not the question format or whether reasoning is deductive, inductive, or abductive.

  2. Knowl 2 — Prompting strategies and their reported puzzle-solving effects

    empirical result

    Across the puzzle studies reviewed, prompting approaches that elicit intermediate reasoning generally outperform simple input-output prompting, but no single strategy wins across all puzzle types and models. Few-shot prompting supplies worked examples in context; Chain-of-Thought (CoT) elicits stepwise reasoning. More structured approaches explore multiple candidate thoughts: Self-Consistency samples reasoning paths and selects a consistent answer; Tree-of-Thoughts (ToT) searches a tree of candidate steps; Tree-of-Uncertain-Thoughts (TouT) branches over uncertain paths; and Everything-of-Thoughts (XoT) combines LLM thought generation with Monte Carlo Tree Search.

    Reported gains are task-specific and should not be compared as if measured under one shared experiment. Self-Refine achieved a 13% higher success rate than CoT on Game of 24. ToT gains over CoT ranged from 26% to 70%, depending on puzzle and tree depth, while requiring more LLM calls. TouT exceeded ToT by 9% on Game of 24 and 3% on mini-crosswords. GoT scored 2%–6% below ToT in the cited puzzle comparisons. XoT improved reported results from 53% to 69% relative to ToT and used fewer LLM calls than the other tested thought-topology methods. In the MARB evaluations, combining Inference-Exclusion-Prompting with CoT scored 82% on puzzles versus 81% for zero-shot CoT, but scored 79% on riddles versus 82% for zero-shot CoT. These mixed results illustrate that prompting benefits depend on the benchmark.

  3. Knowl 3 — Benchmark coverage across the puzzle taxonomy

    data/table

    The survey’s benchmark inventory is unevenly distributed across its puzzle categories. The collected datasets and tasks include:

    • Rule-based, deterministic: BoardgameQA, Sudoku, Rubik’s Cube, mazes, crosswords, the 8-puzzle, Game of 24, and chess.
    • Rule-based, stochastic or incomplete-information: Minesweeper, BoardgameQA scenarios with missing information, card games, and social-deduction games such as Werewolf and Avalon.
    • Rule-less, riddles: BrainTeaser, RiddleSense, BiRdQA, CC-Riddle, PUZZLEQA, and MARB.
    • Rule-less, programming: P3 and a dataset of programming-course code snippets.
    • Rule-less, commonsense reasoning: LatEval, True Detective, DetectBench, and MARB.

    Several dataset sizes and formats show the range of evaluation settings: RiddleSense contains 5.7K riddles; BrainTeaser has 1,119 lateral-thinking puzzles; CC-Riddle contains 27K Chinese character riddles; PUZZLEQA has 558 word puzzles; DetectBench has 1,200 questions; and the programming-course dataset has 530 code snippets. Formats include multiple-choice and free-text QA, code-input generation, long-form detective stories, and interactive yes/no questioning. BoardgameQA is used for both rule-following and incomplete-information settings.

  4. Knowl 4 — Performance on deterministic rule-based puzzles varies sharply by task

    empirical result

    The reviewed deterministic-puzzle evaluations show that LLM performance can be strong on selected tasks but remains inconsistent across puzzle types. In a 2×2×2 Rubik’s Cube evaluation, XoT with self-revision achieved a 77.6% success rate. On the 8-puzzle, XoT with revision reached 93.2% accuracy across 419 puzzles. By contrast, fine-tuned GPT-2 solved the Rubik’s Cube in only 1 of 7 attempts in an earlier study, despite having been tested on more than 2,400 Rubik’s Cube samples and 10K mazes during that work’s evaluation and training setup. For crosswords, GPT-4 with five-shot prompting and ToT solved 4 of 20 puzzles and achieved 60% word-level success. BoardgameQA evaluations found that fine-tuning BERT-large and T5-XXL with proofs was more effective than few-shot PaLM with CoT; extra or contradictory information reduced accuracy. These results come from different studies and conditions, so they indicate variation across tasks rather than a common ranking of systems.

  5. Knowl 5 — Rule-less puzzle evaluations expose limits in inference and lateral thinking

    empirical result

    The rule-less benchmarks reviewed in the survey reveal persistent difficulty in deriving answers from context, especially when a task demands lateral inference or synthesis across clues. On True Detective, vanilla and CoT approaches performed near random despite the stories containing the information needed to solve them. Golden-CoT, which supplies the reasoning behind the correct answer for the model to interpret, improved performance: GPT-3.5 reached a 63% solve rate, while GPT-4 matched human solver results in the reported comparison, even though those human solvers did not receive the answer reasoning. In DetectBench, the best reported model accuracy did not exceed 61.6%; hints and detective-style prompting helped, and larger models generally performed better. On riddle benchmarks, larger models often improved accuracy, but systems still struggled with metaphor, counterfactual situations, and lateral thinking. In LatEval’s interactive story setting, GPT-4 had the strongest answer consistency, but question relevance remained a challenge. Overall, the reviewed evidence shows a gap between LLM performance and human-level understanding on many rule-less tasks.

  6. Knowl 6 — Translating natural-language puzzles into symbolic programs

    model/method

    A neuro-symbolic approach to rule-based puzzles uses an LLM to translate a natural-language puzzle into a formal representation, then delegates solving to a symbolic solver. In the reviewed Answer Set Programming (ASP) approach, the model generates predicates and rules encoding logic puzzles, including chess puzzles, Jobs puzzles, and Sudoku; an ASP solver then operates on that encoding. On a logic-puzzles dataset, GPT-4 used in this translation-and-solving pipeline achieved 92% accuracy, compared with 7% in few-shot prompting and 21% in zero-shot prompting with the same model. The result concerns the combined pipeline and does not isolate puzzle-solving ability from the model’s ability to construct a usable formal encoding. The survey identifies natural-language-to-code translation as an unstudied direction in the puzzle benchmarks it reviewed, and reports no comparable translation studies for rule-less puzzles.

  7. Knowl 7 — Fine-tuning benefits depend on puzzle domain and training data

    empirical result

    The survey finds that fine-tuning can help models acquire puzzle-specific skills, but its effects vary across domains and depend on the training data. For riddles, training models on both RiddleSense and the general commonsense dataset CommonsenseQA improved performance over training on RiddleSense alone in the cited work; transfer learning from CommonsenseQA added a reported 4% accuracy improvement for ALBERT-XXL over simple fine-tuning. For rule-based puzzles, fine-tuning GPT-2 on Sudoku, Rubik’s Cube, and mazes produced suboptimal results in one study, while crosswords showed mixed comparisons with non-neural baselines. Fine-tuning with proofs and CoT performed especially well in BoardgameQA evaluations. The survey also notes the limitation of task specificity: training for one puzzle domain does not establish that a model will transfer to a substantially different one.

  8. Knowl 8 — The survey identifies benchmark and method gaps

    limitation

    The survey identifies a marked shortage of benchmarks for rule-based stochastic puzzles and rule-less programming puzzles, compared with the greater availability of deterministic-puzzle and riddle datasets. Because of the shortage of direct stochastic-puzzle benchmarks, the review includes card and social-deduction games as related tasks involving hidden information or uncertain outcomes. The surveyed puzzle studies also make little use of neuro-symbolic translation from natural language into code, and prompting methods such as ToT and GoT can require more model calls than CoT or XoT, raising efficiency and scalability concerns. The paper therefore points to richer datasets, more work on uncertainty, and further exploration of translation methods as research opportunities; it does not present these opportunities as experimentally validated solutions.

  9. Knowl 9 — Conventional solvers provide a reliability contrast for rule-based puzzles

    empirical result

    The survey’s comparison with conventional puzzle-solving methods emphasizes that explicit algorithms can be more reliable on structured, rule-based tasks than the LLM approaches reviewed. For Sudoku, backtracking and constraint programming are reported to solve puzzles across difficulty levels, with constraint programming often finding solutions within milliseconds; the cited LLM evaluation did not exceed 80% on 5×5 Sudoku puzzles. Conventional search and group-theoretic methods also solve Rubik’s Cube, while standard search algorithms can solve mazes and constraint satisfaction methods can encode Minesweeper. The survey notes a trade-off: conventional methods are generally more deterministic and interpretable, while some methods can be computationally time-consuming. This comparison is intended to contextualize LLM reasoning, not to establish that one solver family is superior for every puzzle.

  10. Knowl 10 — Review scope and stated limitations

    limitation

    The survey reviews LLM puzzle-solving work primarily from the four years preceding publication, drawing mainly on ACL, EMNLP, NAACL, NeurIPS, ICLR, and arXiv sources. The authors caution that the field changes rapidly, so some recent developments may be absent; page limits also prevent exhaustive technical descriptions of every method. Their conclusions are based on empirical analysis and the authors’ interpretation of the reviewed studies. The survey’s scope excludes puzzles that cannot be represented textually, puzzles requiring multimodal understanding, and mathematical puzzles, and it focuses on solving rather than systematically reviewing puzzle generation.

Coverage note — Puzzle generation is only briefly surveyed and is not developed into a standalone knowl because the paper focuses on puzzle solving; prompting-topology definitions are incorporated into the prompting-results knowl rather than separated.

References

  1. 1.Forest Agostinelli, Stephen Marcus McAleer, Alexander Shmakov, and Pierre Baldi. 2019. Solving the rubik’s cube with deep reinforcement learning and search. Nature Machine Intelligence, 1:356 – 363.
  2. 2.Rohan Anil, Andrew M. Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Tachard Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Z. Chen, Eric Chu, J. Clark, Laurent El Shafey, Yanping Huang, Kathleen S. Meier-Hellstern, Gaurav Mishra, Erica Moreira, Mark Omernick, Kevin Robinson, Sebastian Ruder, Yi Tay, Kefan Xiao, Yuanzhong Xu, Yujing Zhang, Gustavo Hernandez Abrego, Junwhan Ahn, Jacob Austin, Paul Barham, Jan A. Botha, James Bradbury, Siddhartha Brahma, Kevin Michael Brooks, Michele Catasta, Yongzhou Cheng, Colin Cherry, Christopher A. Choquette-Choo, Aakanksha Chowdhery, C Crépy, Shachi Dave, Mostafa Dehghani, Sunipa Dev, Jacob Devlin, M. C. D’iaz, Nan Du, Ethan Dyer, Vladimir Feinberg, Fan Feng, Vlad Fienber, Markus Freitag, Xavier García, Sebastian Gehrmann, Lucas González, Guy Gur-Ari, Steven Hand, Hadi Hashemi, Le Hou, Joshua Howland, An Ren Hu, Jeffrey Hui, Jeremy Hurwitz, Michael Isard, Abe Ittycheriah, Matthew Jagielski, Wen Hao Jia, Kathleen Kenealy, Maxim Krikun, Sneha Kudugunta, Chang Lan, Katherine Lee, Benjamin Lee, Eric Li, Mu-Li Li, Wei Li, Yaguang Li, Jun Yu Li, Hyeontaek Lim, Han Lin, Zhong-Zhong Liu, Frederick Liu, Marcello Maggioni, Aroma Mahendru, Joshua Maynez, Vedant Misra, Maysam Moussalem, Zachary Nado, John Nham, Eric Ni, Andrew Nystrom, Alicia Parrish, Marie Pellat, Martin Polacek, Oleksandr Polozov, Reiner Pope, Siyuan Qiao, Emily Reif, Bryan Richter, Parker Riley, Alexandra Ros, Aurko Roy, Brennan Saeta, Rajkumar Samuel, Renee Marie Shelby, Ambrose Slone, Daniel Smilkov, David R. So, Daniela Sohn, Simon Tokumine, Dasha Valter, Vijay Vasudevan, Kiran Vodrahalli, Xuezhi Wang, Pidong Wang, Zirui Wang, Tao Wang, John Wieting, Yuhuai Wu, Ke Xu, Yunhan Xu, Lin Wu Xue, Pengcheng Yin, Jiahui Yu, Qiaoling Zhang, Steven Zheng, Ce Zheng, Wei Zhou, Denny Zhou, Slav Petrov, and Yonghui Wu. 2023. Palm 2 technical report. ArXiv, abs/2305.10403.
  3. 3.Yejin Bang, Samuel Cahyawijaya, Nayeon Lee, Wenliang Dai, Dan Su, Bryan Wilie, Holy Lovenia, Ziwei Ji, Tiezheng Yu, Willy Chung, Quyet V. Do, Yan Xu, and Pascale Fung. 2023. A multitask, multilingual, multimodal evaluation of chatgpt on reasoning, hallucination, and interactivity. ArXiv, abs/2302.04023.
  4. 4.Qiming Bao, Gaël Gendron, Alex Yuxuan Peng, Wanjun Zhong, Ne¸set Özkan Tan, Yang Chen, Michael Witbrock, and Jiamou Liu. 2023. A systematic evaluation of large language models on out-of-distribution logical reasoning tasks. ArXiv, abs/2310.09430.
  5. 5.Houda Nait El Barj and Theophile Sautory. 2024. Reinforcement learning from llm feedback to counteract goal misgeneralization.
  6. 6.Maciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger, Lukas Gianinazzi, Joanna Gajda, Tomasz Lehmann, Michal Podstawski, Hubert Niewiadomski, Piotr Nyczyk, and Torsten Hoefler. 2023. Graph of thoughts: Solving elaborate problems with large language models.
  7. 7.Maciej Besta, Florim Memedi, Zhenyu Zhang, Robert Gerstenberger, Guangyuan Piao, Nils Blach, Piotr Nyczyk, Marcin Copik, Grzegorz Kwasniewski, Jürgen Müller, Lukas Gianinazzi, Ales Kubicek, Hubert Niewiadomski, Aidan O’Mahony, Onur Mutlu, and Torsten Hoefler. 2024. Demystifying chains, trees, and graphs of thoughts.
  8. 8.Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi. 2019. Piqa: Reasoning about physical commonsense in natural language. ArXiv, abs/1911.11641.
  9. 9.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, T. J. Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeff Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. ArXiv, abs/2005.14165.
  10. 10.Murray Campbell, A.Joseph Hoane, and Feng hsiung Hsu. 2002. Deep blue. Artificial Intelligence, 134(1):57–83.
  11. 11.Banghao Chen, Zhaofeng Zhang, Nicolas Langren’e, and Shengxin Zhu. 2023. Unleashing the potential of prompt engineering in large language models: a comprehensive review. ArXiv, abs/2310.14735.
  12. 12.Juntao Chen. 2022. Different algorithms to solve a rubik’s cube. Journal of Physics: Conference Series, 2386(1):012018.
  13. 13.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde, Jared Kaplan, Harrison Edwards, Yura Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, David W. Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert-Voss, William H. Guss, Alex Nichol, Igor Babuschkin, Suchir Balaji, Shantanu Jain, Andrew Carr, Jan Leike, Joshua Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew M. Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, and Wojciech Zaremba. 2021. Evaluating large language models trained on code. ArXiv, abs/2107.03374.
  14. 14.Wenhu Chen, Xueguang Ma, Xinyi Wang, and William W. Cohen. 2022. Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks. ArXiv, abs/2211.12588.
  15. 15.Eric C. Chi and Kenneth Lange. 2013. Techniques for solving sudoku puzzles.
  16. 16.Yew Ken Chia, Vernon Toh Yan Han, Deepanway Ghosal, Lidong Bing, and Soujanya Poria. 2024. Puzzlevqa: Diagnosing multimodal reasoning challenges of language models with abstract visual patterns.
  17. 17.Zheng Chu, Jingchang Chen, Qianglong Chen, Weijiang Yu, Tao He, Haotian Wang, Weihua Peng, Ming Liu, Bing Qin, and Ting Liu. 2023. A survey of chain of thought reasoning: Advances, frontiers and future. ArXiv, abs/2309.15402.
  18. 18.Hyung Won Chung, Le Hou, S. Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Dasha Valter, Sharan Narang, Gaurav Mishra, Adams Wei Yu, Vincent Zhao, Yanping Huang, Andrew M. Dai, Hongkun Yu, Slav Petrov, Ed Huai hsin Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, and Jason Wei. 2022. Scaling instruction-finetuned language models. ArXiv, abs/2210.11416.
  19. 19.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Unsupervised cross-lingual representation learning at scale. In Annual Meeting of the Association for Computational Linguistics.
  20. 20.Antonia Creswell, Murray Shanahan, and Irina Higgins. 2022. Selection-inference: Exploiting large language models for interpretable logical reasoning. ArXiv, abs/2205.09712.
  21. 21.Maksym Del and Mark Fishel. 2022. True detective: A deep abductive reasoning benchmark undoable for gpt-3 and challenging for gpt-4. In STARSEM.
  22. 22.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding.
  23. 23.Ruomeng Ding, Chaoyun Zhang, Lu Wang, Yong Xu, Ming-Jie Ma, Wei Zhang, Si Qin, S. Rajmohan, Qingwei Lin, and Dongmei Zhang. 2023. Everything of thoughts: Defying the law of penrose triangle for thought generation. ArXiv, abs/2311.04254.
  24. 24.Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Zhiyong Wu, Baobao Chang, Xu Sun, Jingjing Xu, Lei Li, and Zhifang Sui. 2023. A survey on in-context learning.
  25. 25.Avia Efrat, Uri Shaham, Dan Kilman, and Omer Levy. 2021. Cryptonite: A cryptic crossword benchmark for extreme ambiguity in language. ArXiv, abs/2103.01242.
  26. 26.Jiazhan Feng, Ruochen Xu, Junheng Hao, Hiteshi Sharma, Yelong Shen, Dongyan Zhao, and Weizhu Chen. 2023a. Language models can be logical solvers. ArXiv, abs/2311.06158.
  27. 27.Xidong Feng, Yicheng Luo, Ziyan Wang, Hongrui Tang, Mengyue Yang, Kun Shao, David Henry Mguni, Yali Du, and Jun Wang. 2023b. Chessgpt: Bridging policy learning and language modeling. ArXiv, abs/2306.09200.
  28. 28.Peter A. Flach and Antonis C. Kakas. 2000. Abductive and inductive reasoning: background and issues.
  29. 29.Yao Fu, Hao-Chun Peng, Ashish Sabharwal, Peter Clark, and Tushar Khot. 2022. Complexity-based prompting for multi-step reasoning. ArXiv, abs/2210.00720.
  30. 30.Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig. 2022. Pal: Program-aided language models. ArXiv, abs/2211.10435.
  31. 31.Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth, and Jonathan Berant. 2021. Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies. Transactions of the Association for Computational Linguistics, 9:346–361.
  32. 32.Deepanway Ghosal, Vernon Toh Yan Han, Chia Yew Ken, and Soujanya Poria. 2024. Are language models puzzle prodigies? algorithmic puzzles unveil serious challenges in multimodal reasoning.
  33. 33.Zhouhong Gu, Zihan Li, Lin Zhang, Zhuozhi Xiong, Sihang Jiang, Xiaoxuan Zhu, Shusen Wang, Zili Wang, Jianchen Wang, Haoning Ye, Wenhao Huang, Yikai Zhang, Hongwei Feng, and Yanghua Xiao. 2023. Go beyond the obvious: Probing the gap of informal reasoning ability between humanity and llms by detective reasoning puzzle benchmark.
  34. 34.Jiaxian Guo, Bo Yang, Paul Yoo, Bill Yuchen Lin, Yusuke Iwasawa, and Yutaka Matsuo. 2023. Suspicion-agent: Playing imperfect information games with theory of mind aware gpt-4. ArXiv, abs/2309.17277.
  35. 35.Akshat Gupta. 2023. Are chatgpt and gpt-4 good poker players? - a pre-flop analysis. ArXiv, abs/2308.12466.
  36. 36.Chenghao Huang, Yanbo Cao, Yinlong Wen, Tao Zhou, and Yanru Zhang. 2024. Pokergpt: An end-to-end lightweight solver for multi-player texas hold’em via large language model. ArXiv, abs/2401.06781.
  37. 37.Jie Huang and Kevin Chen-Chuan Chang. 2022. Towards reasoning in large language models: A survey. ArXiv, abs/2212.10403.
  38. 38.Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xinying Song, and Denny Zhou. 2023a. Large language models cannot self-correct reasoning yet. ArXiv, abs/2310.01798.
  39. 39.Shulin Huang, Shirong Ma, Yinghui Li, Mengzuo Huang, Wuhe Zou, Weidong Zhang, and Haitao Zheng. 2023b. Lateval: An interactive llms evaluation benchmark with incomplete information from lateral thinking puzzles. ArXiv, abs/2308.10855.
  40. 40.Adam Ishay, Zhun Yang, and Joohyung Lee. 2023. Leveraging large language models to generate answer set programs. ArXiv, abs/2307.07699.
  41. 41.Yifan Jiang, Filip Ilievski, and Kaixin Ma. 2023. Brainteaser: Lateral thinking puzzles for large language models. In Conference on Empirical Methods in Natural Language Processing.
  42. 42.Mehran Kazemi, Quan Yuan, Deepti Bhatia, Najoung Kim, Xin Xu, Vaiva Imbrasaite, and Deepak Ramachandran. 2023. Boardgameqa: A dataset for natural language reasoning with contradictory information. ArXiv, abs/2306.07934.
  43. 43.Daniel Khashabi, Sewon Min, Tushar Khot, Ashish Sabharwal, Oyvind Tafjord, Peter Clark, and Hannaneh Hajishirzi. 2020. Unifiedqa: Crossing format boundaries with a single qa system. In Findings.
  44. 44.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. ArXiv, abs/2205.11916.
  45. 45.Richard E. Korf. 1997. Finding optimal solutions to rubik’s cube using pattern databases. In AAAI/IAAI.
  46. 46.Saurabh Kulshreshtha, Olga Kovaleva, Namrata Shivagunde, and Anna Rumshisky. 2022. Down and across: Introducing crossword-solving as a new nlp benchmark. ArXiv, abs/2205.10442.
  47. 47.Yihuai Lan, Zhiqiang Hu, Lei Wang, Yang Wang, DeYong Ye, Peilin Zhao, Ee-Peng Lim, Hui Xiong, and Hao Wang. 2023. Llm-based agent society investigation: Collaboration and confrontation in avalon gameplay. ArXiv, abs/2310.14985.
  48. 48.Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2019. Albert: A lite bert for self-supervised learning of language representations. ArXiv, abs/1909.11942.
  49. 49.Bin Lei, Pei-Hung Lin, Chunhua Liao, and Caiwen Ding. 2023. Boosting logical reasoning in large language models through a new framework: The graph of thought. ArXiv, abs/2308.08614.
  50. 50.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdel rahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2019. Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Annual Meeting of the Association for Computational Linguistics.
  51. 51.Yinghao Li, Haorui Wang, and Chao Zhang. 2023. Assessing logical puzzle solving in large language models: Insights from a minesweeper case study. ArXiv, abs/2311.07387.
  52. 52.Bill Yuchen Lin, Ziyi Wu, Yichi Yang, Dong-Ho Lee, and Xiang Ren. 2021. Riddlesense: Reasoning about riddle questions featuring linguistic creativity and commonsense knowledge. In Findings.
  53. 53.Hanmeng Liu, Ruoxi Ning, Zhiyang Teng, Jian Liu, Qiji Zhou, and Yuexin Zhang. 2023a. Evaluating the logical reasoning ability of chatgpt and gpt-4. ArXiv, abs/2304.03439.
  54. 54.Hanmeng Liu, Zhiyang Teng, Ruoxi Ning, Jian Liu, Qiji Zhou, and Yuexin Zhang. 2023b. Glore: Evaluating logical reasoning of large language models. ArXiv, abs/2310.09107.
  55. 55.Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2021. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. ACM Computing Surveys, 55:1 – 35.
  56. 56.Wentao Liu, Hanglei Hu, Jie Zhou, Yuyang Ding, Junsong Li, Jiayi Zeng, Mengliang He, Qin Chen, Bo Jiang, Aimin Zhou, and Liang He. 2023c. Mathematical language models: A survey.
  57. 57.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. ArXiv, abs/1907.11692.
  58. 58.Jieyi Long. 2023. Large language model guided tree-of-thought. ArXiv, abs/2305.08291.
  59. 59.Man Luo, Shrinidhi Kumbhar, Ming shen, Mihir Parmar, Neeraj Varshney, Pratyay Banerjee, Somak Aditya, and Chitta Baral. 2023. Towards logiglue: A brief survey and a benchmark for analyzing logical reasoning capabilities of language models. ArXiv, abs/2310.00836.
  60. 60.Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Sean Welleck, Bodhisattwa Prasad Majumder, Shashank Gupta, Amir Yazdanbakhsh, and Peter Clark. 2023. Self-refine: Iterative refinement with self-feedback. ArXiv, abs/2303.17651.
  61. 61.Smaragda Markaki and Costas Panagiotakis. 2022. Jigsaw puzzle solving techniques and applications: a survey. The Visual Computer, 39:4405 – 4421.
  62. 62.Stephen McAleer, Forest Agostinelli, Alexander Shmakov, and Pierre Baldi. 2018. Solving the rubik’s cube without human knowledge.
  63. 63.Arindam Mitra and Chitta Baral. 2015. Learning to automatically solve logic grid puzzles. In Conference on Empirical Methods in Natural Language Processing.
  64. 64.Shentong Mo and Miao Xin. 2023. Tree of uncertain thoughts reasoning for large language models. ArXiv, abs/2309.07694.
  65. 65.David A. Noever and Ryerson Burdick. 2021. Puzzle solving without search or human knowledge: An unnatural language approach. ArXiv, abs/2109.02797.
  66. 66.Theo X. Olausson, Alex Gu, Benjamin Lipkin, Cedegao Zhang, Armando Solar-Lezama, Josh Tenenbaum, and Roger Levy. 2023. Linc: A neurosymbolic approach for logical reasoning by combining language models with first-order logic provers. In Conference on Empirical Methods in Natural Language Processing.
  67. 67.OpenAI, :, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mo Bavarian, Jeff Belgum, Irwan Bello, Jake Berdine, Gabriel Bernadett-Shapiro, Christopher Berner, Lenny Bogdonoff, Oleg Boiko, Madeleine Boyd, Anna-Luisa Brakman, Greg Brockman, Tim Brooks, Miles Brundage, Kevin Button, Trevor Cai, Rosie Campbell, Andrew Cann, Brittany Carey, Chelsea Carlson, Rory Carmichael, Brooke Chan, Che Chang, Fotis Chantzis, Derek Chen, Sully Chen, Ruby Chen, Jason Chen, Mark Chen, Ben Chess, Chester Cho, Casey Chu, Hyung Won Chung, Dave Cummings, Jeremiah Currier, Yunxing Dai, Cory Decareaux, Thomas Degry, Noah Deutsch, Damien Deville, Arka Dhar, David Dohan, Steve Dowling, Sheila Dunning, Adrien Ecoffet, Atty Eleti, Tyna Eloundou, David Farhi, Liam Fedus, Niko Felix, Simón Posada Fishman, Juston Forte, Isabella Fulford, Leo Gao, Elie Georges, Christian Gibson, Vik Goel, Tarun Gogineni, Gabriel Goh, Rapha Gontijo-Lopes, Jonathan Gordon, Morgan Grafstein, Scott Gray, Ryan Greene, Joshua Gross, Shixiang Shane Gu, Yufei Guo, Chris Hallacy, Jesse Han, Jeff Harris, Yuchen He, Mike Heaton, Johannes Heidecke, Chris Hesse, Alan Hickey, Wade Hickey, Peter Hoeschele, Brandon Houghton, Kenny Hsu, Shengli Hu, Xin Hu, Joost Huizinga, Shantanu Jain, Shawn Jain, Joanne Jang, Angela Jiang, Roger Jiang, Haozhun Jin, Denny Jin, Shino Jomoto, Billie Jonn, Heewoo Jun, Tomer Kaftan, Łukasz Kaiser, Ali Kamali, Ingmar Kanitscheider, Nitish Shirish Keskar, Tabarak Khan, Logan Kilpatrick, Jong Wook Kim, Christina Kim, Yongjik Kim, Hendrik Kirchner, Jamie Kiros, Matt Knight, Daniel Kokotajlo, Łukasz Kondraciuk, Andrew Kondrich, Aris Konstantinidis, Kyle Kosic, Gretchen Krueger, Vishal Kuo, Michael Lampe, Ikai Lan, Teddy Lee, Jan Leike, Jade Leung, Daniel Levy, Chak Ming Li, Rachel Lim, Molly Lin, Stephanie Lin, Mateusz Litwin, Theresa Lopez, Ryan Lowe, Patricia Lue, Anna Makanju, Kim Malfacini, Sam Manning, Todor Markov, Yaniv Markovski, Bianca Martin, Katie Mayer, Andrew Mayne, Bob McGrew, Scott Mayer McKinney, Christine McLeavey, Paul McMillan, Jake McNeil, David Medina, Aalok Mehta, Jacob Menick, Luke Metz, Andrey Mishchenko, Pamela Mishkin, Vinnie Monaco, Evan Morikawa, Daniel Moss, Tong Mu, Mira Murati, Oleg Murk, David Mély, Ashvin Nair, Reiichiro Nakano, Rajeev Nayak, Arvind Neelakantan, Richard Ngo, Hyeonwoo Noh, Long Ouyang, Cullen O’Keefe, Jakub Pachocki, Alex Paino, Joe Palermo, Ashley Pantuliano, Giambattista Parascandolo, Joel Parish, Emy Parparita, Alex Passos, Mikhail Pavlov, Andrew Peng, Adam Perelman, Filipe de Avila Belbute Peres, Michael Petrov, Henrique Ponde de Oliveira Pinto, Michael, Pokorny, Michelle Pokrass, Vitchyr Pong, Tolly Powell, Alethea Power, Boris Power, Elizabeth Proehl, Raul Puri, Alec Radford, Jack Rae, Aditya Ramesh, Cameron Raymond, Francis Real, Kendra Rimbach, Carl Ross, Bob Rotsted, Henri Roussez, Nick Ryder, Mario Saltarelli, Ted Sanders, Shibani Santurkar, Girish Sastry, Heather Schmidt, David Schnurr, John Schulman, Daniel Selsam, Kyla Sheppard, Toki Sherbakov, Jessica Shieh, Sarah Shoker, Pranav Shyam, Szymon Sidor, Eric Sigler, Maddie Simens, Jordan Sitkin, Katarina Slama, Ian Sohl, Benjamin Sokolowsky, Yang Song, Natalie Staudacher, Felipe Petroski Such, Natalie Summers, Ilya Sutskever, Jie Tang, Nikolas Tezak, Madeleine Thompson, Phil Tillet, Amin Tootoonchian, Elizabeth Tseng, Preston Tuggle, Nick Turley, Jerry Tworek, Juan Felipe Cerón Uribe, Andrea Vallone, Arun Vijayvergiya, Chelsea Voss, Carroll Wainwright, Justin Jay Wang, Alvin Wang, Ben Wang, Jonathan Ward, Jason Wei, CJ Weinmann, Akila Welihinda, Peter Welinder, Jiayi Weng, Lilian Weng, Matt Wiethoff, Dave Willner, Clemens Winter, Samuel Wolrich, Hannah Wong, Lauren Workman, Sherwin Wu, Jeff Wu, Michael Wu, Kai Xiao, Tao Xu, Sarah Yoo, Kevin Yu, Qiming Yuan, Wojciech Zaremba, Rowan Zellers, Chong Zhang, Marvin Zhang, Shengjia Zhao, Tianhao Zheng, Juntang Zhuang, William Zhuk, and Barret Zoph. 2023. Gpt-4 technical report.
  68. 68.Liangming Pan, Alon Albalak, Xinyi Wang, and William Yang Wang. 2023a. Logic-lm: Empowering large language models with symbolic solvers for faithful logical reasoning. ArXiv, abs/2305.12295.
  69. 69.Liangming Pan, Michael Stephen Saxon, Wenda Xu, Deepak Nathani, Xinyi Wang, and William Yang Wang. 2023b. Automatically correcting large language models: Surveying the landscape of diverse self-correction strategies. ArXiv, abs/2308.03188.
  70. 70.Julien Pourcel, Cédric Colas, Pierre-Yves Oudeyer, and Laetitia Teodorescu. 2023. Aces: Generating diverse programming puzzles with autotelic language models and semantic descriptors. ArXiv, abs/2310.10692.
  71. 71.Shuofei Qiao, Yixin Ou, Ningyu Zhang, Xiang Chen, Yunzhi Yao, Shumin Deng, Chuanqi Tan, Fei Huang, and Huajun Chen. 2022. Reasoning with language model prompting: A survey. ArXiv, abs/2212.09597.
  72. 72.Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners.
  73. 73.Colin Raffel, Noam M. Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2019. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21:140:1–140:67.
  74. 74.Josh Rozner, Christopher Potts, and Kyle Mahowald. 2021. Decrypting cryptic crosswords: Semantically complex wordplay puzzles as a target for nlp. ArXiv, abs/2104.08620.
  75. 75.Abulhair Saparov, Richard Yuanzhe Pang, Vishakh Padmakumar, Nitish Joshi, Seyed Mehran Kazemi, Najoung Kim, and He He. 2023. Testing the general deductive reasoning capacity of large language models using ood examples. ArXiv, abs/2305.15269.
  76. 76.Jaromir Savelka, Arav Agarwal, Christopher Bogart, and Majd Sakr. 2023. Large language models (gpt) struggle to answer multiple-choice questions about code.
  77. 77.Tal Schuster, A. Kalyan, Oleksandr Polozov, and Adam Tauman Kalai. 2021. Programming puzzles. ArXiv, abs/2106.05784.
  78. 78.David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, Timothy Lillicrap, Karen Simonyan, and Demis Hassabis. 2017. Mastering chess and shogi by self-play with a general reinforcement learning algorithm.
  79. 79.Helmut Simonis. 2005. Sudoku as a constraint problem.
  80. 80.Chris Studholme. 2001. Minesweeper as a constraint satisfaction problem.
  81. 81.Roxana Szomiu and Adrian Groza. 2021. A puzzle-based dataset for natural language inference. ArXiv, abs/2112.05742.
  82. 82.Kyo Takano. 2023. Self-supervision is all you need for solving rubik’s cube.
  83. 83.Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019. Commonsenseqa: A question answering challenge targeting commonsense knowledge. ArXiv, abs/1811.00937.
  84. 84.Yongqi Tong, Yifan Wang, Dawei Li, Sizhe Wang, Zi Lin, Simeng Han, and Jingbo Shang. 2023. Eliminating reasoning via inferring with planning: A new framework to guide llms’ non-linear thinking. ArXiv, abs/2310.12342.
  85. 85.Gladys Tyen, Hassan Mansoor, Peter Chen, Tony Mak, and Victor Carbune. 2023. Llms cannot find reasoning errors, but can correct them! ArXiv, abs/2311.08516.
  86. 86.Lei Wang, Wanyu Xu, Yihuai Lan, Zhiqiang Hu, Yunshi Lan, Roy Ka-Wei Lee, and Ee-Peng Lim. 2023a. Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models. In Annual Meeting of the Association for Computational Linguistics.
  87. 87.Shenzhi Wang, Chang Liu, Zilong Zheng, Siyuan Qi, Shuo Chen, Qisen Yang, Andrew Zhao, Chaofei Wang, Shiji Song, and Gao Huang. 2023b. Avalon’s game of thoughts: Battle against deception through recursive contemplation. ArXiv, abs/2310.01320.
  88. 88.Weiqi Wang, Tianqing Fang, Wenxuan Ding, Baixuan Xu, Xin Liu, Yangqiu Song, and Antoine Bosselut. 2023c. Car: Conceptualization-augmented reasoner for zero-shot commonsense question answering. In Conference on Empirical Methods in Natural Language Processing.
  89. 89.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Huai hsin Chi, and Denny Zhou. 2022. Self-consistency improves chain of thought reasoning in language models. ArXiv, abs/2203.11171.
  90. 90.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Huai hsin Chi, F. Xia, Quoc Le, and Denny Zhou. 2022. Chain of thought prompting elicits reasoning in large language models. ArXiv, abs/2201.11903.
  91. 91.Yixuan Weng, Minjun Zhu, Fei Xia, Bin Li, Shizhu He, Kang Liu, and Jun Zhao. 2022. Large language models are better reasoners with self-verification. In Conference on Empirical Methods in Natural Language Processing.
  92. 92.Fan Xu, Yunxiang Zhang, and Xiao-Yi Wan. 2022. Cc-riddle: A question answering dataset of chinese character riddles. ArXiv, abs/2206.13778.
  93. 93.Fangzhi Xu, Qika Lin, Jiawei Han, Tianzhe Zhao, Jun Liu, and Erik Cambria. 2023a. Are large language models really good logical reasoners? a comprehensive evaluation and beyond.
  94. 94.Yuzhuang Xu, Shuo Wang, Peng Li, Fuwen Luo, Xiaolong Wang, Weidong Liu, and Yang Liu. 2023b. Exploring large language models for communication games: An empirical study on werewolf. ArXiv, abs/2309.04658.
  95. 95.Sen Yang, Xin Li, Leyang Cui, Li Bing, and Wai Lam. 2023a. Neuro-symbolic integration brings causal and reliable reasoning proofs. ArXiv, abs/2311.09802.
  96. 96.Zonglin Yang, Xinya Du, Rui Mao, Jinjie Ni, and E. Cambria. 2023b. Logical reasoning over natural language as knowledge representation: A survey. ArXiv, abs/2303.12023.
  97. 97.Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik Narasimhan. 2023. Tree of thoughts: Deliberate problem solving with large language models. ArXiv, abs/2305.10601.
  98. 98.Fei Yu, Hongbo Zhang, Prayag Tiwari, and Benyou Wang. 2023a. Natural language reasoning, a survey.
  99. 99.Zihan Yu, Liang He, Zhen Wu, Xinyu Dai, and Jiajun Chen. 2023b. Towards better chain-of-thought prompting strategies: A survey. ArXiv, abs/2310.04959.
  100. 100.Kamyar Zeinalipour, Tommaso laquinta, Asya Zanollo, Giovanni Angelini, Leonardo Rigutini, Marco Maggini, and Marco Gori. 2023a. Italian crossword generator: Enhancing education through interactive word puzzles.
  101. 101.Kamyar Zeinalipour, Mohamed Saad, Marco Maggini, and Marco Gori. 2023b. Arabicros: Ai-powered arabic crossword puzzle generation for educational applications. In Proceedings of ArabicNLP 2023. Association for Computational Linguistics.
  102. 102.Yunxiang Zhang and Xiaojun Wan. 2021. Birdqa: A bilingual dataset for question answering on tricky riddles. ArXiv, abs/2109.11087.
  103. 103.Zhuosheng Zhang, Aston Zhang, Mu Li, and Alexander J. Smola. 2022. Automatic chain of thought prompting in large language models. ArXiv, abs/2210.03493.
  104. 104.Jingmiao Zhao and Carolyn Jane Anderson. 2023. Solving and generating npr sunday puzzles with large language models.
  105. 105.Yongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster, Silviu Pitis, Harris Chan, and Jimmy Ba. 2022. Large language models are human-level prompt engineers. ArXiv, abs/2211.01910.
  106. 106.Andrea Zugarini, Kamyar Zeinalipour, Surya Sai Kadali, Marco Maggini, Marco Gori, and Leonardo Rigutini. 2024. Clue-instruct: Text-based clue generation for educational crossword puzzles.

Citation

MLA
Giadikiaroglou, P., et al. “Puzzle Solving Using Reasoning of Large Language Models: A Survey”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024, pp. 11574–91, https://doi.org/10.18653/v1/2024.emnlp-main.646.
APA
Giadikiaroglou, P., Lymperaiou, M., Filandrianos, G., & Stamou, G. (2024). Puzzle Solving using Reasoning of Large Language Models: A Survey. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 11574–11591. https://doi.org/10.18653/v1/2024.emnlp-main.646
Chicago
Giadikiaroglou, P., M. Lymperaiou, G. Filandrianos, and G. Stamou. 2024. “Puzzle Solving Using Reasoning of Large Language Models: A Survey”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 11574–91. https://doi.org/10.18653/v1/2024.emnlp-main.646.
Harvard
Giadikiaroglou, P. et al. (2024) “Puzzle Solving using Reasoning of Large Language Models: A Survey”, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 11574–11591. Available at: https://doi.org/10.18653/v1/2024.emnlp-main.646.
Vancouver
1. Giadikiaroglou P, Lymperaiou M, Filandrianos G, Stamou G (2024) Puzzle Solving using Reasoning of Large Language Models: A Survey. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 11574–11591

BibTeX

@inproceedings{giadikiaroglou-etal-2024-puzzle,
    title = "Puzzle Solving using Reasoning of Large Language Models: A Survey",
    author = "Giadikiaroglou, Panagiotis  and
      Lymperaiou, Maria  and
      Filandrianos, Giorgos  and
      Stamou, Giorgos",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.646/",
    doi = "10.18653/v1/2024.emnlp-main.646",
    pages = "11574--11591"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/