Algorithm of Thoughts: Enhancing Exploration of Ideas in Large Language Models
Bilgehan SelAhmad Al-TawahaVanshaj KhattarRuoxi JiaMing Jin
Proposes an in-context prompting strategy that internalizes algorithmic search directly within large language model generations, matching or exceeding complex multi-query tree-search methods at a fraction of the computational and token cost.
Large language models often struggle with complex problem-solving tasks that require strategic planning and deep exploration of ideas. While multi-query tree-search methods like Tree of Thoughts improve reasoning, they rely on external software scripts that repeatedly halt, evaluate, and resume model generation. This external loop requires hundreds of model queries per task, dramatically increasing financial costs, processing latency, system memory load, and energy consumption.
The article aims to resolve this operational bottleneck by introducing and evaluating the Algorithm of Thoughts framework. This prompting strategy guides models through structured, algorithmic search pathways entirely in-context, enabling deep idea exploration within a single or minimal query interaction.
To evaluate this framework, the authors tested the approach across challenging reasoning benchmarks, including the Game of 24 mathematical puzzle, 5-by-5 Mini Crosswords, creative writing tasks, and standard dynamic programming problems. The methodology incorporates depth-first and breadth-first search examples directly into the model prompt. This design allows the language model to propose intermediate steps, evaluate progress, and backtrack to alternative solutions within a continuous generation sweep, eliminating the need for external tree-maintenance code.
The results show that the Algorithm of Thoughts outperforms both traditional single-prompt baselines and resource-heavy multi-query methods. In the Game of 24 benchmark, the method achieved a 71% success rate (increasing to 78% with manual solution formatting), outperforming Tree of Thoughts at 69% and Chain-of-Thought at under 10%. Crucially, the method reduced required model queries by more than a factor of 100—from approximately 109 queries to a single query—while consuming significantly fewer total tokens. In the Mini Crosswords benchmark, the approach achieved a 52% word success rate across only two queries, surpassing the 46.5% success rate of Tree of Thoughts across more than 200 queries while reducing total token consumption by 25 times. Furthermore, the model visited fewer search nodes than a conventional programmatic depth-first search, showing that language models effectively combine algorithmic structure with intuitive pattern recognition. When fine-tuned on algorithmic examples, model performance increased by 60 percentage points, compared to just an 8 percentage point gain for standard Chain-of-Thought fine-tuning.
These findings demonstrate that organizations do not need complex, costly external orchestration frameworks to achieve high-level reasoning in artificial intelligence applications. Adopting in-context algorithmic reasoning lowers operational expenses, reduces response latency for real-time systems, and lessens data center energy consumption without sacrificing accuracy.
Decision-makers should consider integrating algorithmic prompting structures into enterprise AI workflows that involve multi-step planning, optimization, or structured problem-solving. Prior to broad deployment, teams should conduct internal pilot tests to balance prompt length against operational requirements, as prompt length directly influences generation speed.
Confidence in these findings is high for advanced frontier models such as GPT-4, Claude 3, and Gemini 1.5 Pro. However, decision-makers should note that the approach relies on the advanced recursive capabilities of top-tier models and may deliver less pronounced improvements on smaller or less capable model architectures. Additionally, while the method is far more efficient than multi-query search tools, it uses more tokens per interaction than direct, single-step prompting.
- Paper: Tree of Thoughts: Deliberate Problem Solving with Large Language Models, Shunyu Yao et al. (2023). Tree of Thoughts establishes the branching search approach that Algorithm of Thoughts contrasts with its in-context, lower-overhead exploration.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). Chain-of-Thought Prompting introduces the step-by-step reasoning baseline that Algorithm of Thoughts develops into algorithm-guided exploration.
- Paper: Self-Consistency Improves Chain of Thought Reasoning in Language Models, Xuezhi Wang et al. (2023). Self-Consistency provides a key multi-sample reasoning baseline against which Algorithm of Thoughts positions its more query-efficient search.
- Paper: Reasoning with Language Model is Planning with World Model, Shibo Hao et al. (2023). Reasoning via Planning shows how language models can use Monte Carlo Tree Search to explore reasoning paths, clarifying the search-based methods Algorithm of Thoughts seeks to improve.
- Paper: Graph of Thoughts: Solving Elaborate Problems with Large Language Models, Maciej Besta et al. (2023). Graph of Thoughts extends tree-based reasoning with flexible operations, helping situate Algorithm of Thoughts among structured exploration methods.
- Paper: ReAct: Synergizing Reasoning and Acting in Language Models, Shunyu Yao et al. (2023). ReAct exemplifies reasoning methods that interleave model generation with external actions, the broader family of multi-step approaches Algorithm of Thoughts aims to avoid.
- Paper: Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning, Zhenni Bi et al. (2025). Forest-of-Thought carries test-time exploration forward by coordinating multiple reasoning trees, extending the search strategies Algorithm of Thoughts seeks to make efficient.
- Paper: Sketch-of-Thought: Efficient LLM Reasoning with Adaptive Cognitive-Inspired Sketching, Simon A. Aytes et al. (2025). Sketch-of-Thought continues the efficiency agenda by compressing reasoning traces to reduce token use while maintaining problem-solving performance.
- Paper: Training Large Language Models to Reason in a Continuous Latent Space, Shibo Hao et al. (2024). Coconut pursues a further alternative to explicit search traces by carrying reasoning in continuous latent states and eliciting emergent search behavior.
- Paper: Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Chain-of-Thought, Violet Xiang et al. (2025). Meta Chain-of-Thought develops search-generated training and reinforcement learning to teach models richer, internalized reasoning beyond in-context algorithmic examples.
