Prompt-Based Monte-Carlo Tree Search for Goal-oriented Dialogue Policy Planning
Xiao YuMaximillian ChenZhou Yu
Proposes GDP-ZERO, a training-free framework that integrates large language model prompting with Open-Loop Monte Carlo Tree Search to plan strategic dialogue actions, outperforming standard ChatGPT in persuasive conversation tasks without requiring annotated training data.
Building automated conversational agents that can effectively steer conversations toward specific objectives—such as persuasion, negotiation, or emotional support—remains a major practical challenge. Traditional approaches rely on training neural networks on human conversation data to guide dialogue strategies. However, these methods often falter because strategic dialogue is inherently subjective, high-quality human demonstrations are difficult to obtain, and annotated training data frequently contains noise. Furthermore, errors in simulated responses tend to compound over multiple conversation turns, degrading the system's ability to plan ahead effectively.
The article demonstrates a novel, zero-training dialogue policy planning framework called GDP-ZERO. The primary objective is to evaluate whether prompting large language models to perform open-loop Monte Carlo Tree Search at decision time can successfully plan strategic conversational moves and achieve goal-oriented outcomes without requiring any model training or specialized annotated datasets.
To evaluate this approach, the authors formulated dialogue policy planning as a stochastic decision process. Rather than relying on static or trained policy classifiers, GDP-ZERO prompts a large language model—specifically ChatGPT—to simultaneously act as a policy generator, user simulator, value function, and dialogue generator during tree search. The open-loop design avoids fixed state representations by dynamically regenerating candidate user and system turns, mitigating simulation error compounding. Credibility was assessed using the PersuasionForGood dataset across both static evaluations and real-time interactive evaluations involving human crowdworkers on Amazon Mechanical Turk, comparing GDP-ZERO against standard ChatGPT prompting and a state-of-the-art rule-based planner called RAP.
The findings demonstrate clear performance advantages. In static comparisons, responses generated via GDP-ZERO planning were preferred over standard ChatGPT responses up to 59.32% of the time, with larger search simulation counts yielding stronger preferences. In live interactive human evaluations, GDP-ZERO achieved the highest scores across all persuasion metrics: it achieved a donation probability of 0.79 (compared to 0.73 for ChatGPT and 0.72 for RAP), produced significantly stronger arguments (4.28 out of 5), and was rated significantly more convincing (4.38) and more natural (4.38). Additionally, human evaluators rated GDP-ZERO as significantly less manipulative (2.29 out of 5) than both baseline systems. The analysis revealed that GDP-ZERO achieves this by adopting a more patient, conservative strategy, avoiding premature requests for donations and balancing logical appeals, emotional appeals, and credibility evidence.
These results show that high-performing conversational planning can be achieved entirely without domain-specific model training, dramatically reducing data collection and engineering costs. This makes advanced conversational agents viable in data-scarce domains where collecting expert dialogue demonstrations is prohibitive. However, decision-makers must weigh this against practical trade-offs. GDP-ZERO introduces notable computational latency; in interactive tests, performing 10 look-ahead simulations required approximately 35 seconds per conversational turn, which can impact real-time user experience.
For organizations considering deployment, the source supports piloting zero-training planning systems in strategic, high-stakes communication settings where turn latency is acceptable or where back-end systems can parallelize simulation queries. Future development should focus on optimizing search runtime, exploring tree-parallelization methods, and testing the framework across other complex dialogue domains beyond charitable persuasion. While results are statistically significant within the evaluated benchmark, stakeholders should exercise caution regarding response latency and potential misalignment risks before deploying simulation-driven planners in consumer-facing environments.
- Paper: Tree of Thoughts: Deliberate Problem Solving with Large Language Models, Shunyu Yao et al. (2023). Introduces the Tree of Thoughts framework for deliberate planning and evaluation using language models, establishing the conceptual foundation for tree search-based LLM planning that GDP-ZERO adapts for dialogue.
- Paper: Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model, Julian Schrittwieser et al. (2020). Presents MuZero's integration of learned transition, value, and policy models with Monte Carlo Tree Search, providing the algorithmic inspiration for prompting LLMs to serve those core planning functions.
- Paper: ReAct: Synergizing Reasoning and Acting in Language Models, Shunyu Yao et al. (2023). Establishes how language models can interleave reasoning traces and environment actions, underlying prompt-based multi-step decision-making pipelines.
- Paper: Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents, Wenlong Huang et al. (2022). Demonstrates that pre-trained language models can generate structured, multi-step action plans without task-specific training, motivating zero-shot planning in dialogue.
- Paper: Deep Reinforcement Learning for Dialogue Generation, Jiwei Li et al. (2016). Shows the importance of long-term conversational planning and lookahead simulation over short-sighted single-turn generation in dialogue systems.
- Paper: Language Agent Tree Search Unifies Reasoning, Acting, and Planning in Language Models, Andy Zhou et al. (2024). Extends Monte Carlo Tree Search planning with language models across broader autonomous agent decision-making environments by incorporating external feedback and self-reflection.
- Paper: WebPilot: A Versatile and Autonomous Multi-Agent System for Web Task Execution with Strategic Exploration, Yao Zhang et al. (2025). Applies reflection-guided tree search and strategic exploration to autonomous agents executing complex tasks in dynamic web environments.
- Paper: Graph of Thoughts: Solving Elaborate Problems with Large Language Models, Maciej Besta et al. (2023). Generalizes tree-structured LLM search into arbitrary graph topologies to support thought aggregation and dynamic feedback loops during complex problem-solving.
- Paper: Code World Models for General Game Playing, Wolfgang Lehrach et al. (2026). Advances MCTS-based LLM planning by synthesizing formal, executable world models in code for general game playing in both perfect- and imperfect-information settings.
- Paper: ProAgent: Building Proactive Cooperative Agents with Large Language Models, Ceyao Zhang et al. (2024). Explores proactive intention reasoning and multi-turn coordination with other agents using large language models without task-specific fine-tuning.
