Language Agent Tree Search Unifies Reasoning, Acting, and Planning in Language Models

Andy ZhouKai YanMichal Shlapentokh-RothmanHaohan WangYu-Xiong Wang

article2024ICML420 citations

Unifies reasoning, acting, and planning in language model agents by integrating Monte Carlo Tree Search with external environment feedback and self-reflection, achieving state-of-the-art results on HumanEval and competitive decision-making on interactive benchmarks without task-specific fine-tuning.

Listen

Autonomous language model agents show strong potential for automating complex digital workflows, yet existing systems often falter in multifaceted environments. Standard methods typically generate actions in a rigid, step-by-step manner without planning ahead, while isolated planning techniques fail to incorporate real-time observations from external tools. The article evaluates Language Agent Tree Search (LATS), a framework designed to unify internal reasoning, external action, and deliberate planning into a single decision-making process without requiring additional model training.

The approach adapts Monte Carlo Tree Search to language models, allowing an agent to explore multiple potential action paths, receive feedback from an external environment, and evaluate progress using internal scoring heuristics and self-reflection. When a simulated path fails, the model generates semantic reflections to inform subsequent attempts. The authors tested this framework across diverse benchmarks, including computer programming on HumanEval and MBPP, multi-hop question answering on HotPotQA, e-commerce web navigation on WebShop, and mathematical problem-solving on Game of 24.

The findings show that deliberate tree search combined with external feedback significantly improves performance across all evaluated domains. In programming, the framework established a state-of-the-art 92.7% pass rate on HumanEval when paired with GPT-4 and achieved 83.8% with GPT-3.5, substantially exceeding existing agent prompting methods. In interactive question answering, combining internal reasoning and external tool use reached a 71% success rate, more than doubling baseline performance. In web navigation, the method attained an average score of 75.9, outperforming specialized reinforcement learning and fine-tuned models without updating model weights. Furthermore, the search strategy expanded fewer total nodes and used fewer tokens upon success than alternative tree-search baselines.

These results indicate that structured planning and environmental feedback make autonomous language agents far more capable and reliable for high-stakes workflows without costly model fine-tuning. Organizations should consider planning-based frameworks for complex reasoning and tool-use applications where execution accuracy is critical. However, decision-makers should note that the approach incurs higher inference costs than simple single-pass prompting and assumes the ability to reset or simulate intermediate task states. Future implementation efforts should focus on optimizing inference efficiency and validating the method within irreversible, live production environments.

Cover for Language Agent Tree Search Unifies Reasoning, Acting, and Planning in Language Models

Abstract

While language models (LMs) have shown potential across a range of decision-making tasks, their reliance on simple acting processes limits their broad deployment as autonomous agents. In this paper, we introduce Language Agent Tree Search (LATS) – the first general framework that synergizes the capabilities of LMs in reasoning, acting, and planning. By leveraging the in-context learning ability of LMs, we integrate Monte Carlo Tree Search into LATS to enable LMs as agents, along with LM-powered value functions and self-reflections for proficient exploration and enhanced decision-making. A key feature of our approach is the incorporation of an environment for external feedback, which offers a more deliberate and adaptive problem-solving mechanism that surpasses the constraints of existing techniques. Our experimental evaluation across diverse domains, including programming, interactive question-answering (QA), web navigation, and math, validates the effectiveness and generality of LATS in decision-making while maintaining competitive or improved reasoning performance. Notably, LATS achieves state-of-the-art pass@1 accuracy (92.7%) for programming on HumanEval with GPT-4 and demonstrates gradient-free performance (average score of 75.9) comparable to gradient-based fine-tuning for web navigation on WebShop with GPT-3.5. Code can be found at https://github.com/lapisrocks/LanguageAgentTreeSearch.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Preliminaries
  • 3.1. Problem Setting and Prompting
  • 3.2. Monte Carlo Tree Search (MCTS)
  • 4. Unifying Reasoning, Acting, and Planning
  • 4.1. LM Agent
  • 4.2. LATS
  • 5. Experiments
  • 5.1. HotPotQA
  • 5.2. Programming
  • 5.3. WebShop
  • 5.4. Ablation Study and Additional Analysis
  • 6. Conclusion
  • Impact Statement
  • Acknowledgements
  • References
  • Appendix of LATS
  • A. LATS Pseudocode
  • B. More Discussion on Limitations
  • C. Additional Ablations
  • D. Environment Details
  • D.1. HotPotQA
  • D.2. Programming
  • D.3. WebShop
  • D.4. Game of 24
  • E. HotPotQA Prompts
  • E.1. Base Acting Prompt
  • E.2. Base Reasoning Prompt
  • E.3. Value Function Prompt
  • E.4. Reflection Prompt
  • F. Programming Prompts
  • F.1. HumanEval function implementation example
  • F.2. Base Acting/Reasoning Prompt
  • F.3. Reflection Prompt
  • F.4. Test Case Generation Prompt
  • G. WebShop Prompts
  • G.1. Acting Prompt
  • G.2. Value Function Prompt
  • G.3. Reflection Prompt

Knowls

  1. Knowl 1 — Language Agent Tree Search (LATS) Algorithm

    algorithm

    Language Agent Tree Search (LATS) unifies language model (LM) reasoning, acting, and planning by embedding a pretrained language model pθp_\theta into a Monte Carlo Tree Search (MCTS) loop over state nodes s=[x,a1…t,o1…t]s = [x, a_{1\dots t}, o_{1\dots t}], where xx is the initial input query, a1…ta_{1\dots t} is the sequence of actions/reasoning thoughts, and o1…to_{1\dots t} is the sequence of environment observations.

    LATS operates over KK iterations (episodes) using six consecutive operations:

    1. Selection: Traverses from the root s0s_0 by picking child nodes that maximize the Upper Confidence bounds applied to Trees (UCT) formula until an unexpanded or non-terminal leaf node is reached: UCT(s)=V(s)+wln⁡N(p)N(s)\text{UCT}(s) = V(s) + w \sqrt{\frac{\ln N(p)}{N(s)}} where N(s)N(s) is the visit count of child node ss, N(p)N(p) is the visit count of its parent node pp, V(s)V(s) is the value estimate of the subtree at ss, and ww is an exploration hyperparameter (typically w=1.0w=1.0).
    2. Expansion: Samples nn alternative candidate actions/thoughts at(i)∼pθ(st)a^{(i)}_t \sim p_\theta(s_t) for i=1,…,ni=1,\dots,n from the policy model, executes them in the external environment to obtain observations ot(i)o^{(i)}_t, and creates child nodes st+1(i)=[st,at(i),ot(i)]s^{(i)}_{t+1} = [s_t, a^{(i)}_t, o^{(i)}_t].
    3. Evaluation: Computes a scalar heuristic value V(st(i))V(s^{(i)}_t) for each child node by combining an LM-generated trajectory evaluation score with a self-consistency frequency score.
    4. Simulation (Rollout): Selectively expands the most promising node depth-wise until reaching a terminal state or horizon limit LL, yielding an objective trajectory outcome/reward rr.
    5. Backpropagation: Propagates the final scalar reward rr up the visited trajectory path s0,s1,…,sTs_0, s_1, \dots, s_T, updating node statistics: N(st)←N(st)+1N(s_t) \leftarrow N(s_t) + 1 V(st)←Vold(st)(N(st)−1)+rN(st)V(s_t) \leftarrow \frac{V_{\text{old}}(s_t)(N(s_t) - 1) + r}{N(s_t)}
    6. Reflection: When a trajectory simulation fails (r<1r < 1), an LM reflection generator prefp_{\text{ref}} analyzes the failed trajectory context ctc_t and reward to produce verbal self-reflections that diagnose mistakes and suggest corrections. These reflections are stored in external memory and injected into the prompt context for subsequent search iterations.
    Input: Initial state s0s_0, action generator pθp_\theta, value function pVp_V, reflection generator prefp_{\text{ref}}, generated actions nn, depth limit LL, roll-out budget KK, exploration weight ww, value weighting λ\lambda
    Output: Best action trajectory or solved solution state
    Initialize action space AA, observation space OO
    Initialize value table V:S→RV: S \to \mathbb{R} and visit counts N:S→NN: S \to \mathbb{N} to 1
    for k←0k \leftarrow 0 to K−1K - 1 do
        st←s0s_t \leftarrow s_0
        for t←0t \leftarrow 0 to L−1L - 1 do
            if sts_t is not terminal then
                for i←1i \leftarrow 1 to nn do
                    Sample at(i)∼pθ(st)a^{(i)}_t \sim p_\theta(s_t)
                    Execute at(i)a^{(i)}_t to get ot(i)o^{(i)}_t from environment
                    st+1(i)←[st,at(i),ot(i)]s^{(i)}_{t+1} \leftarrow [s_t, a^{(i)}_t, o^{(i)}_t]
                    V(st+1(i))←λ⋅pV(st+1(i))+(1−λ)⋅SC(st+1(i))V(s^{(i)}_{t+1}) \leftarrow \lambda \cdot p_V(s^{(i)}_{t+1}) + (1 - \lambda) \cdot \text{SC}(s^{(i)}_{t+1})
                    Add st+1(i)s^{(i)}_{t+1} to children of sts_t
                end for
            end if
            if sts_t is terminal then
                Get ground-truth reward rr from environment
                if rr is not successful then
                    reflection←pref(st,r)\text{reflection} \leftarrow p_{\text{ref}}(s_t, r)
                    Append reflection\text{reflection} to memory context cc
                end if
                break
            end if
            at←arg⁡max⁡a∈children(st)[V(s)+wln⁡N(st)/N(s)]a_t \leftarrow \arg\max_{a \in \text{children}(s_t)} [V(s) + w \sqrt{\ln N(s_t) / N(s)}]
            st+1←s_{t+1} \leftarrow child corresponding to ata_t
            N(st+1)←N(st+1)+1N(s_{t+1}) \leftarrow N(s_{t+1}) + 1
            if ata_t is a finish/terminal action then
                break
            end if
        end for
        T←T \leftarrow number of steps in current trajectory
        for t←T−1t \leftarrow T - 1 down to 00 do
            V(st)←V(st)(N(st)−1)+rN(st)V(s_t) \leftarrow \frac{V(s_t)(N(s_t) - 1) + r}{N(s_t)}
        end for
        if task is successfully solved then
            return trajectory
        end if
    end for
    return node with highest value VV
  2. Knowl 2 — LATS State Value Function Combining LM Evaluation and Self-Consistency

    equation

    In Language Agent Tree Search (LATS), the heuristic value V(s)V(s) assigned to an intermediate state node s=[x,a1…i,o1…i]s = [x, a_{1\dots i}, o_{1\dots i}] during tree expansion is computed via a weighted linear combination of an LM-evaluated correctness score and a self-consistency score:

    V(s)=λ⋅LM(s)+(1−λ)⋅SC(s)V(s) = \lambda \cdot \text{LM}(s) + (1 - \lambda) \cdot \text{SC}(s)

    where:

    • LM(s)∈[0,1]\text{LM}(s) \in [0, 1] is a scalar score produced by prompting the pretrained language model pθp_\theta to reason about the trajectory and output an integer correctness score between 11 and 1010 (normalized to [0.1,1.0][0.1, 1.0]). Crucially, LM(s)\text{LM}(s) is evaluated after observing environmental feedback oio_i for the most recent action aia_i.
    • SC(s)∈(0,1]\text{SC}(s) \in (0, 1] represents the self-consistency heuristic, defined as the empirical frequency of action aia_i when sampled among the nn candidate proposals at parent state si−1s_{i-1}.
    • λ∈[0,1]\lambda \in [0, 1] is a weighting hyperparameter. In experimental evaluations, λ=0.5\lambda = 0.5 is used for reasoning and question-answering tasks (HotPotQA, Game of 24), while λ=0.8\lambda = 0.8 is used for interactive tasks (Programming on HumanEval/MBPP, WebShop).
  3. Knowl 3 — Verbal Self-Reflection and Semantic Memory Mechanism in LATS

    model/method

    LATS incorporates verbal self-reflection to provide a non-differentiable "semantic gradient" that guides tree search across episodes without gradient-based parameter updates.

    When a search trajectory terminates in an unsuccessful state (r<1r < 1 or execution failure), the language model pθp_\theta is queried as a reflection generator prefp_{\text{ref}}. It receives the full trajectory history (thoughts, actions, and observations) along with the final environmental reward rr, and is prompted to:

    1. Diagnose the specific reasoning or acting failure modes in the sequence.
    2. Formulate an actionable, high-level natural language correction plan to avoid repeating the mistake.

    The resulting verbal reflections and failed trajectory summaries are appended to an external episodic memory buffer. In all subsequent MCTS rollouts and state evaluations for the task, these stored reflections are dynamically included in the in-context prompt of both the policy agent and the LM value function, steering the exploration away from previously identified error paths.

  4. Knowl 4 — HumanEval and MBPP Code Generation Performance

    data/table

    On Python program synthesis benchmarks HumanEval (164 problems) and MBPP (397 problem subset), LATS is evaluated against standard and search-based prompting baselines using Pass@1 accuracy. For LATS, each action is a complete generated Python program; synthetic unit tests generated by the LM serve as the environment observation, with the fraction of passed tests used as intermediate feedback (k=8k=8 iterations, n=5n=5 candidate solutions per expansion):

    Prompt Method Base Model HumanEval Pass@1 (%) MBPP Pass@1 (%)
    CoT GPT-3.5 46.9 54.9
    ReAct GPT-3.5 56.9 67.0
    ToT GPT-3.5 54.4 65.8
    RAP GPT-3.5 63.1 71.4
    Reflexion GPT-3.5 68.1 70.0
    LATS (ReAct) GPT-3.5 83.8 81.1
    Base LM GPT-4 80.1 –
    Reflexion GPT-4 91.0 –
    LATS (ReAct) GPT-4 92.7 –

    LATS with GPT-3.5 outperforms the next best method (Reflexion) on HumanEval by +15.7%+15.7\%, and on MBPP by +11.1%+11.1\%. With GPT-4, LATS reaches 92.7%92.7\% Pass@1 on HumanEval.

  5. Knowl 5 — Multi-Hop Question Answering Accuracy on HotPotQA

    data/table

    On a 100-question subset of the multi-hop reasoning benchmark HotPotQA, LATS is evaluated using GPT-3.5 against reasoning-based and acting-based baselines under an oracle setting where the environment returns exact match (EM) correctness feedback upon final answer submission (k=50k=50 trajectories, n=5n=5 expanded candidates):

    Paradigm Prompt Method HotPotQA (EM) ↑\uparrow
    Internal Reasoning Base LM 0.32
    Internal Reasoning CoT 0.34
    Internal Reasoning CoT-SC 0.38
    Internal Reasoning ToT 0.55
    Internal Reasoning RAP (n=5n=5) 0.60
    Internal Reasoning RAP (n=10n=10) 0.60
    Internal Reasoning LATS (CoT) 0.62
    Interactive Acting ReAct 0.32
    Interactive Acting ReAct (best of kk) 0.38
    Interactive Acting Reflexion 0.51
    Interactive Acting ToT (ReAct) 0.39
    Interactive Acting RAP (ReAct) 0.54
    Interactive Acting LATS (ReAct, n=3n=3) 0.58
    Interactive Acting LATS (ReAct, n=5n=5) 0.63
    Interactive Acting LATS (ReAct, n=10n=10) 0.65
    Hybrid Reasoning+Acting LATS (CoT + ReAct) 0.71

    Direct adaptations of tree search to acting (ToT ReAct: 0.39; RAP ReAct: 0.54) perform worse than their pure reasoning counterparts, whereas LATS leverages environment feedback effectively. The hybrid variant LATS (CoT + ReAct), which attempts internal reasoning first and falls back to interactive API search upon failure, reaches the top score of 0.710.71 EM.

  6. Knowl 6 — Interactive Web Navigation Performance on WebShop

    data/table

    In the WebShop interactive e-commerce shopping environment (evaluated over 50 human instructions with a depth limit of 15), agents must search and navigate product pages to fulfill multi-attribute purchase goals. Performance is measured by average Task Score (100×average reward100 \times \text{average reward}, capturing percentage of matched attributes) and Success Rate (SR, fraction of episodes with reward r=1.0r=1.0). LATS (n=5,k=30n=5, k=30) with GPT-3.5 is evaluated against prompting baselines (k=30k=30) and gradient-based methods:

    Method Task Score ↑\uparrow Success Rate (%) ↑\uparrow
    ReAct 53.8 28.0
    ReAct (best of kk) 59.1 32.0
    Reflexion 64.2 35.0
    LATS (ReAct) 75.9 38.0
    Imitation Learning (IL) 59.9 29.1
    IL + RL 62.4 28.7
    Fine-tuning 67.5 45.0
    Expert Human 82.1 59.6

    Without parameter fine-tuning, LATS achieves a Task Score of 75.975.9, exceeding imitation learning (59.959.9), reinforcement learning (62.462.4), and gradient-based fine-tuning (67.567.5), demonstrating effective exploration in large action spaces.

  7. Knowl 7 — Mathematical Reasoning Performance on Game of 24

    empirical result

    On the Game of 24 mathematical reasoning benchmark (evaluated over 50 games with GPT-3.5, sampling n=5n=5 candidate thoughts per step and k=30k=30 trajectories with depth limit 55), the objective is to form an arithmetic expression equating to 24 using four provided numbers exactly once.

    Success rates across prompting methods:

    • CoT: 0.080.08 (8%)
    • Reflexion: 0.120.12 (12%)
    • ToT: 0.200.20 (20%)
    • RAP: 0.400.40 (40%)
    • LATS (CoT, λ=0.5\lambda=0.5): 0.44\mathbf{0.44} (44%)

    Setting λ=1.0\lambda = 1.0 (removing the self-consistency component from the LATS value function) drops LATS performance from 0.440.44 to 0.400.40, demonstrating that incorporating action proposal self-consistency into MCTS value estimation improves reasoning accuracy.

  8. Knowl 8 — Ablation of Search Algorithm, Heuristic, and Exploration Parameters in LATS

    empirical result

    Ablation experiments on HotPotQA using GPT-3.5 (n=5n=5 expanded child nodes, k=50k=50 trajectories) isolate the contribution of individual algorithmic components in LATS:

    1. Component Removal vs Full LATS (EM = 0.63):

      • No Reflection: Removing verbal self-reflection drops EM to 0.580.58 (−0.05-0.05).
      • DFS Search: Replacing MCTS selection and backpropagation with Depth-First Search with tree pruning drops EM to 0.420.42 (−0.21-0.21).
      • No LM Heuristic: Removing the LM state value function pV(s)p_V(s) and relying only on sparse environment rewards drops EM to 0.370.37 (−0.26-0.26).
    2. Exploration Weight (ww):

      • w=0.5w = 0.5: EM decreases to 0.550.55 due to insufficient exploration.
      • w=1.0w = 1.0: EM is 0.630.63.
      • w=2.0w = 2.0: EM remains 0.630.63, but demonstrates faster initial search convergence.
    3. Search Depth (dd):

      • Reducing maximum search depth from d=7d = 7 to d=4d = 4 results in EM =0.58= 0.58, showing that while most multi-hop questions are solvable in four steps, deeper search trees prevent premature truncation on complex queries.
  9. Knowl 9 — Sample Complexity and Token Consumption of Tree-Search Prompting Methods

    theoretical result

    For an MCTS or tree search agent executing kk trajectories with branching factor nn, the asymptotic sample complexity (number of LM queries and token generation upper bound) is O(kn)\mathcal{O}(kn), compared to O(k)\mathcal{O}(k) for linear sampling methods (such as ReAct best-of-kk or CoT-SC).

    Empirical evaluation of total token consumption and expanded nodes upon successful problem-solving on HotPotQA (n=5,k=50n=5, k=50) shows:

    • ToT (ReAct): 0.490.49 EM, 84.0584.05 nodes expanded on average, 210,215210,215 total tokens consumed.
    • RAP (ReAct): 0.540.54 EM, 70.6070.60 nodes expanded on average, 176,500176,500 total tokens consumed.
    • LATS (ReAct): 0.63\mathbf{0.63} EM, 66.65\mathbf{66.65} nodes expanded on average, 173,290\mathbf{173,290} total tokens consumed.

    Across trajectory budgets k∈{10,30,50}k \in \{10, 30, 50\}, LATS consistently expands fewer nodes upon success than RAP (requiring 3.553.55 fewer nodes on average) and ToT (12.1212.12 fewer nodes on average), demonstrating higher search efficiency despite sharing the same O(kn)\mathcal{O}(kn) sample complexity.

  10. Knowl 10 — Environment Reversion Assumption and Computational Limitations in LATS

    limitation

    Language Agent Tree Search operates under two primary limitations:

    1. Environment State Reversibility: LATS relies on a model-free MCTS framework that requires resetting or rolling back the environment state to arbitrary prior decision nodes sts_t during tree search. While easily satisfied in software domains (code evaluation via test replay, API search queries via cached histories, web navigation via URL/state reset, text games), LATS is not directly applicable to non-reversible or real-time physical environments lacking state reset capabilities.
    2. Inference Compute and Latency: Due to expanding nn parallel actions at each search depth and evaluating them with an LM value function across kk rollouts, LATS incurs higher token consumption and wall-clock latency compared to single-pass prompting (ReAct, Reflexion). Branching parameter nn allows tuning the trade-off between task success and computational cost (e.g., n=1n=1 reduces complexity to single-trial efficiency).

Coverage note — None omitted. All substantial contributions—including the LATS algorithm, value function equations, memory/reflection mechanism, empirical results across all four benchmarks (HumanEval/MBPP, HotPotQA, WebShop, Game of 24), ablations, cost/complexity analyses, and stated limitations—are fully represented.

References

  1. 1.Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Chuyuan Fu, Keerthana Gopalakrishnan, Karol Hausman, Alex Herzog, Daniel Ho, Jasmine Hsu, Julian Ibarz, Brian Ichter, Alex Irpan, Eric Jang, Rosario Jauregui Ruano, Kyle Jeffrey, Sally Jesmonth, Nikhil J Joshi, Ryan Julian, Dmitry Kalashnikov, Yuheng Kuang, Kuang-Huei Lee, Sergey Levine, Yao Lu, Linda Luu, Carolina Parada, Peter Pastor, Jornell Quiambao, Kanishka Rao, Jarek Rettinghouse, Diego Reyes, Pierre Sermanet, Nicolas Sievers, Clayton Tan, Alexander Toshev, Vincent Vanhoucke, Fei Xia, Ted Xiao, Peng Xu, Sichun Xu, Mengyuan Yan, and Andy Zeng. Do as I can, not as I say: Grounding language in robotic affordances. In CoRL, 2022.
  2. 2.Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, and Charles Sutton. Program synthesis with large language models. In NeurIPS, 2022.
  3. 3.Bowen Baker, Ilge Akkaya, Peter Zhokhov, Joost Huizinga, Jie Tang, Adrien Ecoffet, Brandon Houghton, Raul Sampedro, and Jeff Clune. Video pretraining (VPT): Learning to act by watching unlabeled online videos. In NeurIPS, 2022.
  4. 4.Maciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger, Lukas Gianinazzi, Joanna Gajda, Tomasz Lehmann, Michal Podstawski, Hubert Niewiadomski, Piotr Nyczyk, and Torsten Hoefler. Graph of thoughts: Solving elaborate problems with large language models. arXiv:2308.09687, 2023.
  5. 5.Samuel R Bowman, Gabor Angeli, Christopher Potts, and Christopher D Manning. A large annotated corpus for learning natural language inference. In EMNLP, 2015.
  6. 6.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. Language models are few-shot learners. In NeurIPS, 2020.
  7. 7.Murray Campbell, A Joseph Hoane Jr, and Feng-hsiung Hsu. Deep blue. Artificial intelligence, 2002.
  8. 8.Bei Chen, Fengji Zhang, Anh Nguyen, Daoguang Zan, Zeqi Lin, Jian-Guang Lou, and Weizhu Chen. CodeT: Code generation with generated tests. In ICLR, 2023a.
  9. 9.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde, Jared Kaplan, Harrison Edwards, Yura Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, David W. Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert-Voss, William H. Guss, Alex Nichol, Igor Babuschkin, Suchir Balaji, Shantanu Jain, Andrew Carr, Jan Leike, Joshua Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew M. Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, and Wojciech Zaremba. Evaluating large language models trained on code. arXiv:2107.03374, 2021.
  10. 10.Wenhu Chen, Xueguang Ma, Xinyi Wang, and William W. Cohen. Program of thoughts prompting: disentangling computation from reasoning for numerical reasoning tasks. TMLR, 2023b. ISSN 2835-8856.
  11. 11.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. PaLM: Scaling language modeling with pathways. JMLR, 24 (240):1–113, 2023.
  12. 12.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training verifiers to solve math word problems. arXiv:2110.14168, 2021.
  13. 13.Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen, Samuel Stevens, Boshi Wang, Huan Sun, and Yu Su. Mind2Web: Towards a generalist agent for the web. In NeurIPS Datasets and Benchmarks Track, 2023.
  14. 14.Danny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, Wenlong Huang, Yevgen Chebotar, Pierre Sermanet, Daniel Duckworth, Sergey Levine, Vincent Vanhoucke, Karol Hausman, Marc Toussaint, Klaus Greff, Andy Zeng, Igor Mordatch, and Pete Florence. PaLM-E: An embodied multimodal language model. In ICML, 2023.
  15. 15.Yilun Du, Mengjiao Yang, Bo Dai, Hanjun Dai, Ofir Nachum, Joshua B. Tenenbaum, Dale Schuurmans, and Pieter Abbeel. Learning universal policies via text-guided video generation. In NeurIPS, 2023.
  16. 16.Jonathan St BT Evans. Intuition and reasoning: A dual-process perspective. Psychological Inquiry, pages 313 – 326, 2010.
  17. 17.Linxi Fan, Guanzhi Wang, Yunfan Jiang, Ajay Mandlekar, Yuncong Yang, Haoyi Zhu, Andrew Tang, De-An Huang, Yuke Zhu, and Anima Anandkumar. MineDojo: Building open-ended embodied agents with internet-scale knowledge. In NeurIPS Datasets and Benchmarks Track, 2022.
  18. 18.Hiroki Furuta, Ofir Nachum, Kuang-Huei Lee, Yutaka Matsuo, Shixiang Shane Gu, and Izzeddin Gur. Multimodal web navigation with instruction-finetuned foundation models. In ICLR, 2024.
  19. 19.Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig. PAL: Program-aided language models. In ICML, 2023.
  20. 20.Jiaxian Guo, Sidi Lu, Han Cai, Weinan Zhang, Yong Yu, and Jun Wang. Long text generation via adversarial training with leaked information. In AAAI, 2018.
  21. 21.William H. Guss, Brandon Houghton, Nicholay Topin, Phillip Wang, Cayden Codel, Manuela Veloso, and Ruslan Salakhutdinov. MineRL: A large-scale dataset of Minecraft demonstrations. In IJCAI, 2019.
  22. 22.Danijar Hafner, Timothy Lillicrap, Ian Fischer, Ruben Villegas, David Ha, Honglak Lee, and James Davidson. Learning latent dynamics for planning from pixels. In ICML, 2019.
  23. 23.Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy Lillicrap. Mastering diverse domains through world models. arXiv:2301.04104, 2023.
  24. 24.Shibo Hao, Yi Gu, Haodi Ma, Joshua Jiahua Hong, Zhen Wang, Daisy Zhe Wang, and Zhiting Hu. Reasoning with language model is planning with world model. In EMNLP, 2023.
  25. 25.Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xinying Song, and Denny Zhou. Large language models cannot self-correct reasoning yet. In ICLR, 2024.
  26. 26.Wenlong Huang, F. Xia, Ted Xiao, Harris Chan, Jacky Liang, Peter R. Florence, Andy Zeng, Jonathan Tompson, Igor Mordatch, Yevgen Chebotar, Pierre Sermanet, Noah Brown, Tomas Jackson, Linda Luu, Sergey Levine, Karol Hausman, and Brian Ichter. Inner monologue: Embodied reasoning through planning with language models. In CoRL, 2022.
  27. 27.Levente Kocsis and Csaba Szepesvári. Bandit based monte-carlo planning. In ECML, 2006.
  28. 28.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. Large language models are zero-shot reasoners. In NeurIPS, 2022.
  29. 29.Steven M. LaValle. Rapidly-exploring random trees : A new tool for path planning. The Annual Research Report, 1998.
  30. 30.Evan Zheran Liu, Kelvin Guu, Panupong Pasupat, Tianlin Shi, and Percy Liang. Reinforcement learning on web interfaces using workflow-guided exploration. In ICLR, 2018.
  31. 31.Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, Shudan Zhang, Xiang Deng, Aohan Zeng, Zhengxiao Du, Chenhui Zhang, Sheng Shen, Tianjun Zhang, Yu Su, Huan Sun, Minlie Huang, Yuxiao Dong, and Jie Tang. AgentBench: Evaluating LLMs as agents. In ICLR, 2024.
  32. 32.Zhihan Liu, Hao Hu, Shenao Zhang, Hongyi Guo, Shuqi Ke, Boyi Liu, and Zhaoran Wang. Reason for future, act for now: A principled framework for autonomous LLM agents with provable sample efficiency. arXiv:2309.17382, 2023.
  33. 33.Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Shashank Gupta, Bodhisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, and Peter Clark. Self-refine: Iterative refinement with self-feedback. In NeurIPS, 2023.
  34. 34.Ramesh Nallapati, Bowen Zhou, Cicero dos Santos, Caglar Gulcehre, and Bing Xiang. Abstractive text summarization using sequence-to-sequence RNNs and beyond. In Special Interest Group on Natural Language Learning, 2016.
  35. 35.OpenAI. GPT-4 technical report. arXiv:2303.08774, 2023.
  36. 36.Yujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu, Lan Yan, Yaxi Lu, Yankai Lin, Xin Cong, Xiangru Tang, Bill Qian, Sihan Zhao, Runchu Tian, Ruobing Xie, Jie Zhou, Mark Gerstein, Dahai Li, Zhiyuan Liu, and Maosong Sun. ToolLLM: Facilitating large language models to master 16000+ real-world APIs. In ICLR, 2024.
  37. 37.Abulhair Saparov and He He. Language models are greedy reasoners: A systematic formal analysis of chain-of-thought. In ICLR, 2023.
  38. 38.Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. Toolformer: Language models can teach themselves to use tools. In NeurIPS, 2023.
  39. 39.Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, and Yueting Zhuang. HuggingGPT: Solving AI tasks with ChatGPT and its friends in Hugging Face. In NeurIPS, 2023.
  40. 40.Noah Shinn, Federico Cassano, Beck Labash, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. Reflexion: Language agents with verbal reinforcement learning. In NeurIPS, 2023.
  41. 41.Mohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk, Adam Trischler, and Matthew Hausknecht. ALFWorld: Aligning text and embodied environments for interactive learning. In ICLR, 2020.
  42. 42.David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, L. Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Vedavyas Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy P. Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis. Mastering the game of Go with deep neural networks and tree search. Nature, 529:484–489, 2016.
  43. 43.David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, L. Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Vedavyas Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy P. Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis. Mastering chess and Shogi by self-play with a general reinforcement learning algorithm. arXiv:1712.01815, 2017.
  44. 44.Steven A. Sloman. The empirical case for two systems of reasoning. Psychological Bulletin, 119:3–22, 1996.
  45. 45.Haotian Sun, Yuchen Zhuang, Lingkai Kong, Bo Dai, and Chao Zhang. AdaPlanner: Adaptive planning from feedback with language models. In NeurIPS, 2023.
  46. 46.Dídac Surís, Sachit Menon, and Carl Vondrick. ViperGPT: Visual inference via Python execution for reasoning. In ICCV, 2023.
  47. 47.Maciej Swiechowski, Konrad Godlewski, Bartosz Sawicki, and Jacek Mańdziuk. Monte Carlo tree search: A review of recent modifications and applications. Artificial Intelligence Review, 56:2497–2562, 2021.
  48. 48.Hugo Touvron, Louis Martin, Kevin R. Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Daniel M. Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony S. Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel M. Kloumann, A. V. Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, R. Subramanian, Xia Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zhengxu Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. Llama 2: Open foundation and fine-tuned chat models. arXiv:2307.09288, 2023.
  49. 49.Tom Vodopivec, Spyridon Samothrakis, and Branko Ster. On Monte Carlo tree search and reinforcement learning. Journal of Artificial Intelligence Research, 60:881–936, 2017.
  50. 50.Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar. Voyager: An open-ended embodied agent with large language models. arXiv:2305.16291, 2023.
  51. 51.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, and Denny Zhou. Self-consistency improves chain of thought reasoning in language models. In ICLR, 2022.
  52. 52.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. Chain of thought prompting elicits reasoning in large language models. In NeurIPS, 2022.
  53. 53.Michael Wooldridge and Nicholas R Jennings. Intelligent agents: Theory and practice. The Knowledge Engineering Review, 10:115 – 152, 1995.
  54. 54.Philipp Wu, Alejandro Escontrela, Danijar Hafner, Pieter Abbeel, and Ken Goldberg. Daydreamer: World models for physical robot learning. In CoRL, 2023.
  55. 55.Yuxi Xie, Kenji Kawaguchi, Yiran Zhao, Xu Zhao, Min-Yen Kan, Junxian He, and Qizhe Xie. Decomposition enhances reasoning via self-evaluation guided decoding. arXiv:2305.00633, 2023.
  56. 56.Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W Cohen, Ruslan Salakhutdinov, and Christopher D Manning. HotpotQA: A dataset for diverse, explainable multi-hop question answering. In EMNLP, 2018.
  57. 57.Shunyu Yao, Howard Chen, John Yang, and Karthik R Narasimhan. WebShop: Towards scalable real-world web interaction with grounded language agents. In NeurIPS, 2022.
  58. 58.Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik Narasimhan. Tree of thoughts: deliberate problem solving with large language models. In NeurIPS, 2023a.
  59. 59.Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. ReAct: Synergizing reasoning and acting in language models. In ICLR, 2023b.
  60. 60.Weirui Ye, Shaohuai Liu, Thanard Kurutach, Pieter Abbeel, and Yang Gao. Mastering Atari games with limited data. In NeurIPS, 2021.
  61. 61.Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Olivier Bousquet, Quoc Le, and Ed Chi. Least-to-most prompting enables complex reasoning in large language models. In ICLR, 2022.
  62. 62.Yuchen Zhuang, Xiang Chen, Tong Yu, Saayan Mitra, Victor Bursztyn, Ryan A. Rossi, Somdeb Sarkhel, and Chao Zhang. ToolChain*: Efficient action space navigation in large language models with A* search. In ICLR, 2023.

Citation

MLA
Zhou, A., et al. “Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models”. arXiv, 2023, http://arxiv.org/abs/2310.04406v3.
APA
Zhou, A., Yan, K., Shlapentokh-Rothman, M., Wang, H., & Wang, Y.-X. (2023). Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models. arXiv. http://arxiv.org/abs/2310.04406v3
Chicago
Zhou, A., K. Yan, M. Shlapentokh-Rothman, H. Wang, and Y.-X. Wang. 2023. “Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models”. arXiv. http://arxiv.org/abs/2310.04406v3.
Harvard
Zhou, A. et al. (2023) “Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2310.04406v3.
Vancouver
1. Zhou A, Yan K, Shlapentokh-Rothman M, Wang H, Wang Y-X (2023) Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models. arXiv

BibTeX

@article{zhou2023language,
  title = {Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models},
  author = {Zhou, Andy and Yan, Kai and Shlapentokh-Rothman, Michal and Wang, Haohan and Wang, Yu-Xiong},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2310.04406v3},
  eprint = {2310.04406}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/