Algorithm of Thoughts: Enhancing Exploration of Ideas in Large Language Models

Bilgehan SelAhmad Al-TawahaVanshaj KhattarRuoxi JiaMing Jin

article2024ICML121 citations

Proposes an in-context prompting strategy that internalizes algorithmic search directly within large language model generations, matching or exceeding complex multi-query tree-search methods at a fraction of the computational and token cost.

Listen

Large language models often struggle with complex problem-solving tasks that require strategic planning and deep exploration of ideas. While multi-query tree-search methods like Tree of Thoughts improve reasoning, they rely on external software scripts that repeatedly halt, evaluate, and resume model generation. This external loop requires hundreds of model queries per task, dramatically increasing financial costs, processing latency, system memory load, and energy consumption.

The article aims to resolve this operational bottleneck by introducing and evaluating the Algorithm of Thoughts framework. This prompting strategy guides models through structured, algorithmic search pathways entirely in-context, enabling deep idea exploration within a single or minimal query interaction.

To evaluate this framework, the authors tested the approach across challenging reasoning benchmarks, including the Game of 24 mathematical puzzle, 5-by-5 Mini Crosswords, creative writing tasks, and standard dynamic programming problems. The methodology incorporates depth-first and breadth-first search examples directly into the model prompt. This design allows the language model to propose intermediate steps, evaluate progress, and backtrack to alternative solutions within a continuous generation sweep, eliminating the need for external tree-maintenance code.

The results show that the Algorithm of Thoughts outperforms both traditional single-prompt baselines and resource-heavy multi-query methods. In the Game of 24 benchmark, the method achieved a 71% success rate (increasing to 78% with manual solution formatting), outperforming Tree of Thoughts at 69% and Chain-of-Thought at under 10%. Crucially, the method reduced required model queries by more than a factor of 100—from approximately 109 queries to a single query—while consuming significantly fewer total tokens. In the Mini Crosswords benchmark, the approach achieved a 52% word success rate across only two queries, surpassing the 46.5% success rate of Tree of Thoughts across more than 200 queries while reducing total token consumption by 25 times. Furthermore, the model visited fewer search nodes than a conventional programmatic depth-first search, showing that language models effectively combine algorithmic structure with intuitive pattern recognition. When fine-tuned on algorithmic examples, model performance increased by 60 percentage points, compared to just an 8 percentage point gain for standard Chain-of-Thought fine-tuning.

These findings demonstrate that organizations do not need complex, costly external orchestration frameworks to achieve high-level reasoning in artificial intelligence applications. Adopting in-context algorithmic reasoning lowers operational expenses, reduces response latency for real-time systems, and lessens data center energy consumption without sacrificing accuracy.

Decision-makers should consider integrating algorithmic prompting structures into enterprise AI workflows that involve multi-step planning, optimization, or structured problem-solving. Prior to broad deployment, teams should conduct internal pilot tests to balance prompt length against operational requirements, as prompt length directly influences generation speed.

Confidence in these findings is high for advanced frontier models such as GPT-4, Claude 3, and Gemini 1.5 Pro. However, decision-makers should note that the approach relies on the advanced recursive capabilities of top-tier models and may deliver less pronounced improvements on smaller or less capable model architectures. Additionally, while the method is far more efficient than multi-query search tools, it uses more tokens per interaction than direct, single-step prompting.

Cover for Algorithm of Thoughts: Enhancing Exploration of Ideas in Large Language Models

Abstract

Current literature, aiming to surpass the “Chain-of-Thought” approach, often resorts to external modi operandi involving halting, modifying, and then resuming the generation process to boost Large Language Models’ (LLMs) reasoning capacities. Due to their myopic perspective, they escalate the number of query requests, leading to increased costs, memory, and computational overheads. Addressing this, we propose the Algorithm of Thoughts—a novel strategy that propels LLMs through algorithmic reasoning pathways. By employing algorithmic examples fully in-context, this overarching view of the whole process exploits the innate recurrence dynamics of LLMs, expanding their idea exploration with merely one or a few queries. Our technique outperforms earlier single-query methods and even more recent multi-query strategies that employ an extensive tree search algorithms while using significantly fewer tokens. Intriguingly, our results suggest that instructing an LLM using an algorithm can lead to performance surpassing that of the algorithm itself, hinting at LLM’s inherent ability to weave its intuition into optimized searches. We probe into the underpinnings of our method’s efficacy and its nuances in application. The code and related content can be found in: algorithm-of-thoughts.github.io.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Algorithm of Thoughts
  • 4. Experiments
  • 4.1. Game of 24
  • 4.2. Mini Crosswords
  • 4.3. Finetuning
  • 5. Discussion
  • 6. Conclusion
  • 7. Limitations
  • Acknowledgments
  • Impact Statement
  • References
  • B. Creative Writing
  • B.1. Task Setup
  • B.2. Baselines
  • B.3. AoT Setup
  • B.4. Results
  • C. CoT vs. Single Iteration AoT in the Game of 24
  • D. Detailed Analysis on the Effect of the Length of the Prompts
  • E. Proof of Corollary 3.1
  • F. Prompts
  • F.1. Game of 24
  • AoT (DFS)
  • F.1.1. AOT (LONG)
  • F.1.2. AOT (RANDOM)
  • F.1.3. AOT (BFS)
  • F.2. AoT (Short)
  • F.3. 5 × 5 Mini Crosswords Prompts
  • F.3.1. AOT
  • F.3.2. PROPOSE WORDS
  • F.4. Creative Writing
  • F.4.1. AOT
  • F.4.2. SCORE PROMPT

Knowls

  1. Knowl 1 — Algorithm of Thoughts treats search as part of the in-context example

    model/method

    Algorithm of Thoughts (AoT) changes the usual in-context format from a problem paired with an answer or a linear chain of steps to a problem paired with a search process and its solution. The example demonstrates candidate exploration, assessment, and possible backtracking, rather than requiring the language model to follow one fixed linear reasoning path. At inference, the model generates this exploration as a continuous response, typically using one or a small number of interactions instead of separate model queries for each candidate or search step.

  2. Knowl 2 — AoT uses in-context search examples to guide exploration and backtracking

    model/method

    AoT is designed for problems that can be decomposed into states with alternative next actions. Its examples show intermediate states, candidate actions, and the outcomes of considering those actions; they also demonstrate moving on from an unpromising route to another candidate. The paper predominantly uses depth-first-search-style examples, with breadth-first-search-style examples also explored. Candidate assessment and search progression occur within the same generated response, so the method does not require an external controller to issue a new query for every node. The examples are intended to induce the model's own prioritization of promising candidates, not to make it execute a prescribed pseudocode program.

  3. Knowl 3 — AoT improves Game of 24 success with one query and fewer tokens than Tree of Thoughts

    data/table

    The Game of 24 test used games ranked 901–1000 by relative difficulty; a solution had to use the four given numbers exactly once and only addition, subtraction, multiplication, and division to produce 24. Standard prompting and chain-of-thought (CoT) were evaluated in a five-shot setting, with 100 samples averaged; CoT self-consistency (CoT-SC) used 100 votes. Tree of Thoughts (ToT) used breadth 5 and had a maximum of 100 node visits. AoT used five in-context examples and one query per test game. The reported comparison is:

    MethodSuccessAverage queriesPrompt tokensCompletion tokens
    I/O7.3%116418
    CoT4.0%142146.2
    CoT-SC9.0%10042,1004,620
    I/O + Refine27%10458360
    ToT (b = 5)69%109.113,9005,500
    AoT71%15,450998.4

    AoT had the highest reported success rate, 2 percentage points above ToT, while using one query rather than an average of 109.1 and fewer prompt and completion tokens than ToT. The I/O + Refine result was taken from prior work and allowed up to 10 refinement iterations. When the model's discovered solution was manually resolved into a final answer, AoT reached 78% success; the paper attributes a combined 7% of cases to expression missteps and failure to finalize a found solution, indicating that answer articulation can limit measured performance independently of search.

  4. Knowl 4 — AoT solves mini crosswords through sequential word placement and compatibility checks

    empirical result

    For 5×5 mini crosswords, AoT first uses one query to propose candidate words for rows and columns and identify a promising starting word. In a second query, the model generates candidate words for the remaining clues, checks whether their letters agree with already placed intersecting words, and either places a compatible word or shifts to another candidate. The generation ends when the grid is complete or no suitable continuation is found. The five in-context examples comprised three completed puzzles and two that filled most of the grid. Testing used 20 puzzles (games 1, 6, …, 91, and 96); examples came from games 136, 141, 146, 151, and 156. Word success is the percentage of puzzle words completed correctly.

    MethodWord successAverage queriesPrompt tokensCompletion tokens
    I/O14%1790.330.5
    CoT-SC15.6%11,4001,600
    ToT46.5%> 20096,70021.8k
    AoT52%23,800975.6

    AoT achieved the highest reported word success and used over 100 times fewer queries and about 25 times fewer total tokens than ToT. The ToT success rate was taken from its published result.

  5. Knowl 5 — AoT's expressivity result links exponential intermediate generation to exponential-time problems

    theoretical result

    The paper gives the following informal corollary. Let nn be the number of input tokens and let a≥1a \geq 1. If a transformer prompted with AoT can generate ana^n intermediate tokens to solve a problem, then

    TIME(an)⊆AOT(n),\mathrm{TIME}(a^n) \subseteq \mathrm{AOT}(n),

    where TIME(an)\mathrm{TIME}(a^n) denotes problems solvable by a Turing machine in time O(an)O(a^n), and AOT(n)\mathrm{AOT}(n) denotes the number of AoT decoding steps for an input of nn tokens. This is a conditional expressivity claim about AoT under the stated generation capability, not an empirical demonstration that the evaluated models solve all such problems.

  6. Knowl 6 — AoT shows competitive results across question answering, dynamic programming, writing, and models

    empirical result

    The paper evaluates AoT beyond the Game of 24. For question answering, AoT used a zero-shot prompt that proposed three strategies and expanded them before selecting one; results are for the first 100 questions of GSM8K and StrategyQA. For Coin Change and Edit Distance, the authors compare against I/O and CoT but omit ToT because its DFS/BFS search is explosive and it does not use tabulation. For creative writing, each input supplied four arbitrary sentences, and the task was to write four coherent paragraphs ending in those sentences; GPT-4 scored each response five times on a 1–10 coherence scale. AoT's creative-writing prompt generated five plans, selected one, drafted a passage, and refined it. Separate Game of 24 results test Claude 3 and Gemini 1.5 Pro.

    EvaluationI/OCoTCoT-SCToTAoT
    GSM8K, accuracy51%86%—90%89%
    StrategyQA, accuracy73%82%—83%84%
    Coin Change, accuracy72%76%—Not reported96%
    Edit Distance, accuracy61%64%—Not reported90%
    Creative writing, mean coherence score6.196.93—7.567.58
    Creative writing, average queries11—201

    On Game of 24, the reported success rates were GPT-4: I/O 7%, CoT-SC 9%, AoT 71%; Claude 3: I/O 6%, CoT-SC 9%, AoT 68%; and Gemini 1.5 Pro: I/O 6%, CoT-SC not run, AoT 55%. These results show competitive AoT performance on the tested tasks; the creative-writing improvement over ToT was reported as not statistically significant.

  7. Knowl 7 — AoT retains an advantage after fine-tuning on Game of 24 examples

    empirical result

    The authors fine-tuned GPT-3.5-Turbo with 900 examples formatted using either CoT or AoT, then evaluated Game of 24 success rates with the corresponding prompting approach. The reported rates were:

    Prompting approachWithout fine-tuningWith fine-tuning
    CoT3%12%
    AoT3%63%

    In this setup, fine-tuning with AoT examples raised success to 63%, compared with 12% after CoT fine-tuning. The paper interprets the contrast as evidence that fine-tuning alone did not provide the same benefit as explicitly prompting the model to explore possible options.

  8. Knowl 8 — AoT's search behavior reflects its examples and can visit fewer nodes than DFS

    empirical result

    In Game of 24 comparisons, AoT systematically visited fewer search nodes than the depth-first search algorithm represented in its examples. The authors attribute this to the language model using its own heuristic to prioritize candidates rather than following DFS's uniform next-subtree choice; the paper does not report a single numerical node-count summary. Prompt examples also affected search length: AoT (Short), whose examples showed one or two steps to a solution, produced shorter searches, while AoT (Long), with three to five additional subtree explorations, produced longer ones. The authors report that a short, single-initial-operation AoT variant achieved 48% success versus 4% for CoT, showing that even this limited search-style example differed substantially from a linear chain-of-thought prompt.

  9. Knowl 9 — Repeated in-context examples with identical answers can reduce arithmetic accuracy

    empirical result

    To probe possible context effects, the authors prompted text-davinci-003 with arithmetic questions such as 11−2=11 - 2 = after adding multiple correct in-context equations that all had the same output (for example, equations yielding 10). As the number of these examples increased, the probability of generating the correct answer token declined steeply in the reported experiment. The authors suggest that correct reasoning examples can still bias the model toward echoing a repeated output. This finding motivates AoT examples that represent unsuccessful search attempts and subsequent recovery, rather than merely adding examples with varied answers or repeated successful outcomes.

  10. Knowl 10 — AoT reduces query counts but still has token and model-capability limitations

    limitation

    AoT's extensive generated search uses more resources than standard prompting and CoT, particularly in tokens, even though it can use far fewer queries than external tree-search approaches. The paper identifies token-efficient examples as an open need. Its main experiments are centered on GPT-4; a separate Game of 24 comparison includes Claude 3 and Gemini 1.5 Pro, but the evidence across other models is limited to that task. The authors caution that AoT's gains may depend on a sufficiently capable model and may not transfer comparably to weaker models.

Coverage note — Fine-grained mini-crossword error categories and the appendix's full prompt transcripts are omitted because they are task-specific diagnostics and implementation examples rather than additional standalone core findings.

References

  1. 1.Al-Tawaha, A., Kaushik, H., Sel, B., Jia, R., and Jin, M. Decision-focused learning for inverse noncooperative games: Generalization bounds and convergence analysis. IFAC-PapersOnLine, 56(2):9336–9341, 2023.
  2. 2.Al-Tawaha, A. S., Aljanaideh, K., and Alshorman, A. A singular value thresholding algorithm for order estimation. In 2021 American Control Conference (ACC), pp. 4478–4483. IEEE, 2021.
  3. 3.Aminabadi, R. Y., Rajbhandari, S., Awan, A. A., Li, C., Li, D., Zheng, E., Ruwase, O., Smith, S., Zhang, M., Rasley, J., et al. Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale. In SC22: International Conference for High Performance Computing, Networking, Storage and Analysis, pp. 1–15. IEEE, 2022.
  4. 4.Austin, J., Odena, A., Nye, M., Bosma, M., Michalewski, H., Dohan, D., Jiang, E., Cai, C., Terry, M., Le, Q., et al. Program synthesis with large language models. arXiv preprint arXiv:2108.07732, 2021.
  5. 5.Baddeley, A. Working memory: looking back and looking forward. Nature reviews neuroscience, 4(10):829–839, 2003.
  6. 6.Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., et al. Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073, 2022.
  7. 7.Banerjee, S., Bringsjord, S., Giancola, M., and Govindarajulu, N. S. Qualitative mechanical problem-solving by artificial agents:: Further progress, under psychometric ai. In The International FLAIRS Conference Proceedings, volume 35, 2022.
  8. 8.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020.
  9. 9.Chen, L., Zaharia, M., and Zou, J. Frugalgpt: How to use large language models while reducing cost and improving performance. arXiv preprint arXiv:2305.05176, 2023.
  10. 10.Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. d. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374, 2021.
  11. 11.Chiang, D., Cholak, P., and Pillay, A. Tighter bounds on the expressivity of transformer encoders. arXiv preprint arXiv:2301.10743, 2023.
  12. 12.Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311, 2022.
  13. 13.Dhar, P. The carbon impact of artificial intelligence. Nat. Mach. Intell., 2(8):423–425, 2020.
  14. 14.Drozdov, A., Scharli, N., Aky ¨ urek, E., Scales, N., Song, X., Chen, X., Bousquet, O., and Zhou, D. Compositional Semantic Parsing with Large Language Models. September 2022. URL https://openreview.net/forum?id=gJW8hSGBys8.
  15. 15.Feng, G., Gu, Y., Zhang, B., Ye, H., He, D., and Wang, L. Towards revealing the mystery behind chain of thought: a theoretical perspective. arXiv preprint arXiv:2305.15408, 2023.
  16. 16.Gu, S., Sel, B., Ding, Y., Wang, L., Lin, Q., Jin, M., and Knoll, A. Balance reward and safety optimization for safe reinforcement learning: A perspective of gradient manipulation. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pp. 21099–21106, 2024a.
  17. 17.Gu, S., Sel, B., Ding, Y., Wang, L., Lin, Q., Knoll, A., and Jin, M. Safe and balanced: A framework for constrained multi-objective reinforcement learning. arXiv preprint arXiv:2405.16390, 2024b.
  18. 18.Helie, S. and Pizlo, Z. When is psychology research useful in artificial intelligence? a case for reducing computational complexity in problem solving. Topics in Cognitive Science, 14(4):687–701, 2022.
  19. 19.Holyoak, K. J. and Morrison, R. G. The Cambridge handbook of thinking and reasoning. Cambridge University Press, 2005.
  20. 20.Huang, J. and Chang, K. C.-C. Towards reasoning in large language models: A survey. arXiv preprint arXiv:2212.10403, 2022.
  21. 21.Jin, M., Khattar, V., Kaushik, H., Sel, B., and Jia, R. On solution functions of optimization: Universal approximation and covering number bounds. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 37, pp. 8123–8131, 2023a.
  22. 22.Jin, M., Sel, B., Hardeep, F., and Yin, W. A human-on-the-loop optimization autoformalism approach for sustainability. arXiv preprint arXiv:2308.10380, 2023b.
  23. 23.Kadavath, S., Conerly, T., Askell, A., Henighan, T., Drain, D., Perez, E., Schiefer, N., Hatfield-Dodds, Z., DasSarma, N., Tran-Johnson, E., et al. Language models (mostly) know what they know. arXiv preprint arXiv:2207.05221, 2022.
  24. 24.Kahneman, D. Thinking, fast and slow. macmillan, 2011.
  25. 25.Khattar, V. and Jin, M. Winning the citylearn challenge: adaptive optimization with evolutionary search under trajectory-based guidance. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 37, pp. 14286–14294, 2023.
  26. 26.Khattar, V., Ding, Y., Sel, B., Lavaei, J., and Jin, M. A cmdp-within-online framework for meta-safe reinforcement learning. In The Eleventh International Conference on Learning Representations, 2022.
  27. 27.Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., and Iwasawa, Y. Large language models are zero-shot reasoners. Advances in neural information processing systems, 35:22199–22213, 2022.
  28. 28.Lanham, T., Chen, A., Radhakrishnan, A., Steiner, B., Denison, C., Hernandez, D., Li, D., Durmus, E., Hubinger, E., Kernion, J., et al. Measuring faithfulness in chain-of-thought reasoning. arXiv preprint arXiv:2307.13702, 2023.
  29. 29.Liang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., Zhang, Y., Narayanan, D., Wu, Y., Kumar, A., et al. Holistic evaluation of language models. arXiv preprint arXiv:2211.09110, 2022.
  30. 30.Libby, M. E., Weiss, J. S., Bancroft, S., and Ahearn, W. H. A comparison of most-to-least and least-to-most prompting on the acquisition of solitary play skills. Behavior analysis in practice, 1:37–43, 2008.
  31. 31.Lin, T.-W., Khattar, V., Huang, Y., Hong, J., Jia, R., Liu, C.-C., Sangiovanni-Vincentelli, A., and Jin, M. Causal-prompt: Enhancing llms with weakly supervised causal reasoning for robust per-formance in non-language tasks.
  32. 32.Liu, Y., Han, T., Ma, S., Zhang, J., Yang, Y., Tian, J., He, H., Li, A., He, M., Liu, Z., et al. Summary of chatgpt/gpt-4 research and perspective towards the future of large language models. arXiv preprint arXiv:2304.01852, 2023.
  33. 33.Long, J. Large language model guided tree-of-thought. arXiv preprint arXiv:2305.08291, 2023.
  34. 34.Lyu, Q., Havaldar, S., Stein, A., Zhang, L., Rao, D., Wong, E., Apidianaki, M., and Callison-Burch, C. Faithful chain-of-thought reasoning. arXiv preprint arXiv:2301.13379, 2023.
  35. 35.Merrill, W. and Sabharwal, A. The expresssive power of transformers with chain of thought. arXiv preprint arXiv:2310.07923, 2023.
  36. 36.Mialon, G., Dess`ı, R., Lomeli, M., Nalmpantis, C., Pasunuru, R., Raileanu, R., Roziere, B., Schick, T., Dwivedi-Yu, J., Celikyilmaz, A., et al. Augmented language models: a survey. arXiv preprint arXiv:2302.07842, 2023.
  37. 37.Monsell, S. Task switching. Trends in cognitive sciences, 7(3):134–140, 2003.
  38. 38.Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744, 2022.
  39. 39.Robinson, J. and Wingate, D. Leveraging Large Language Models for Multiple Choice Question Answering. September 2022. URL https://openreview.net/forum?id=yKbprarjc5B.
  40. 40.Schuurmans, D. Memory augmented large language models are computationally universal. arXiv preprint arXiv:2301.04589, 2023.
  41. 41.Sel, A., Sel, B., and Kasnakoglu, C. Glsdc based parameter estimation algorithm for a pmsm model. Energies, 14(3):611, 2021.
  42. 42.Sel, A., Sel, B., Coskun, U., and Kasnakoglu, C. Sos-based nonlinear observer design for simultaneous state and disturbance estimation designed for a pmsm model. Sustainability, 14(17):10650, 2022.
  43. 43.Sel, B., Tawaha, A., Ding, Y., Jia, R., Ji, B., Lavaei, J., and Jin, M. Learning-to-learn to guide random search: Derivative-free meta blackbox optimization on manifold. In Learning for Dynamics and Control Conference, pp. 38–50. PMLR, 2023.
  44. 44.Sel, B., Shanmugasundaram, P., Kachuee, M., Zhou, K., Jia, R., and Jin, M. Skin-in-the-game: Decision making via multi-stakeholder alignment in llms. arXiv preprint arXiv:2405.12933, 2024.
  45. 45.Shao, Z., Gong, Y., Shen, Y., Huang, M., Duan, N., and Chen, W. Synthetic Prompting: Generating Chain-of-Thought Demonstrations for Large Language Models. June 2023. URL https://openreview.net/forum?id=RYD1UMgTdk.
  46. 46.Sloman, S. A. The empirical case for two systems of reasoning. Psychological bulletin, 119(1):3, 1996.
  47. 47.Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., et al. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. arXiv preprint arXiv:2206.04615, 2022.
  48. 48.Suzgun, M., Scales, N., Scharli, N., Gehrmann, S., Tay, Y., Chung, H. W., Chowdhery, A., Le, Q. V., Chi, E. H., Zhou, D., and Wei, J. Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them, October 2022. URL http://arxiv.org/abs/2210.09261. arXiv:2210.09261 [cs].
  49. 49.Thoppilan, R., De Freitas, D., Hall, J., Shazeer, N., Kulshreshtha, A., Cheng, H.-T., Jin, A., Bos, T., Baker, L., Du, Y., et al. Lamda: Language models for dialog applications. arXiv preprint arXiv:2201.08239, 2022.
  50. 50.Turpin, M., Michael, J., Perez, E., and Bowman, S. R. Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting. arXiv preprint arXiv:2305.04388, 2023.
  51. 51.Wang, X., Wei, J., Schuurmans, D., Le, Q. V., Chi, E. H., Narang, S., Chowdhery, A., and Zhou, D. Self-Consistency Improves Chain of Thought Reasoning in Language Models. September 2022. URL https://openreview.net/forum?id=1PL1NIMMrw.
  52. 52.Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., Chi, E. H., Hashimoto, T., Vinyals, O., Liang, P., Dean, J., and Fedus, W. Emergent Abilities of Large Language Models, October 2022a. URL http://arxiv.org/abs/2206.07682. arXiv:2206.07682 [cs].
  53. 53.Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824–24837, 2022b.
  54. 54.Wu, C.-J., Raghavendra, R., Gupta, U., Acun, B., Ardalani, N., Maeng, K., Chang, G., Aga, F., Huang, J., Bai, C., et al. Sustainable ai: Environmental implications, challenges and opportunities. Proceedings of Machine Learning and Systems, 4:795–813, 2022.
  55. 55.Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T. L., Cao, Y., and Narasimhan, K. Tree of Thoughts: Deliberate Problem Solving with Large Language Models, May 2023. URL http://arxiv.org/abs/2305.10601. arXiv:2305.10601 [cs].
  56. 56.Zelikman, E., Wu, Y., Mu, J., and Goodman, N. Star: Bootstrapping reasoning with reasoning. Advances in Neural Information Processing Systems, 35:15476–15488, 2022.
  57. 57.Zhang, Z., Zhang, A., Li, M., and Smola, A. Automatic Chain of Thought Prompting in Large Language Models. September 2022. URL https://openreview.net/forum?id=5NTt8GFjUHkr.
  58. 58.Zhou, D., Scharli, N., Hou, L., Wei, J., Scales, N., Wang, X., Schuurmans, D., Cui, C., Bousquet, O., Le, Q. V., and Chi, E. H. Least-to-Most Prompting Enables Complex Reasoning in Large Language Models. September 2022a. URL https://openreview.net/forum?id=WZH7099tgfM.
  59. 59.Zhou, H., Nova, A., Larochelle, H., Courville, A., Neyshabur, B., and Sedghi, H. Teaching algorithmic reasoning via in-context learning. arXiv preprint arXiv:2211.09066, 2022b.

Citation

MLA
Sel, B., et al. “Algorithm of Thoughts: Enhancing Exploration of Ideas in Large Language Models”. arXiv, 2023, http://arxiv.org/abs/2308.10379v3.
APA
Sel, B., Al-Tawaha, A., Khattar, V., Jia, R., & Jin, M. (2023). Algorithm of Thoughts: Enhancing Exploration of Ideas in Large Language Models. arXiv. http://arxiv.org/abs/2308.10379v3
Chicago
Sel, B., A. Al-Tawaha, V. Khattar, R. Jia, and M. Jin. 2023. “Algorithm of Thoughts: Enhancing Exploration of Ideas in Large Language Models”. arXiv. http://arxiv.org/abs/2308.10379v3.
Harvard
Sel, B. et al. (2023) “Algorithm of Thoughts: Enhancing Exploration of Ideas in Large Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2308.10379v3.
Vancouver
1. Sel B, Al-Tawaha A, Khattar V, Jia R, Jin M (2023) Algorithm of Thoughts: Enhancing Exploration of Ideas in Large Language Models. arXiv

BibTeX

@article{sel2023algorithm,
  title = {Algorithm of Thoughts: Enhancing Exploration of Ideas in Large Language Models},
  author = {Sel, Bilgehan and Al-Tawaha, Ahmad and Khattar, Vanshaj and Jia, Ruoxi and Jin, Ming},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2308.10379v3},
  eprint = {2308.10379}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/