Prompt-Based Monte-Carlo Tree Search for Goal-oriented Dialogue Policy Planning

Xiao YuMaximillian ChenZhou Yu

article2023EMNLP84 citations

Proposes GDP-ZERO, a training-free framework that integrates large language model prompting with Open-Loop Monte Carlo Tree Search to plan strategic dialogue actions, outperforming standard ChatGPT in persuasive conversation tasks without requiring annotated training data.

Listen

Building automated conversational agents that can effectively steer conversations toward specific objectives—such as persuasion, negotiation, or emotional support—remains a major practical challenge. Traditional approaches rely on training neural networks on human conversation data to guide dialogue strategies. However, these methods often falter because strategic dialogue is inherently subjective, high-quality human demonstrations are difficult to obtain, and annotated training data frequently contains noise. Furthermore, errors in simulated responses tend to compound over multiple conversation turns, degrading the system's ability to plan ahead effectively.

The article demonstrates a novel, zero-training dialogue policy planning framework called GDP-ZERO. The primary objective is to evaluate whether prompting large language models to perform open-loop Monte Carlo Tree Search at decision time can successfully plan strategic conversational moves and achieve goal-oriented outcomes without requiring any model training or specialized annotated datasets.

To evaluate this approach, the authors formulated dialogue policy planning as a stochastic decision process. Rather than relying on static or trained policy classifiers, GDP-ZERO prompts a large language model—specifically ChatGPT—to simultaneously act as a policy generator, user simulator, value function, and dialogue generator during tree search. The open-loop design avoids fixed state representations by dynamically regenerating candidate user and system turns, mitigating simulation error compounding. Credibility was assessed using the PersuasionForGood dataset across both static evaluations and real-time interactive evaluations involving human crowdworkers on Amazon Mechanical Turk, comparing GDP-ZERO against standard ChatGPT prompting and a state-of-the-art rule-based planner called RAP.

The findings demonstrate clear performance advantages. In static comparisons, responses generated via GDP-ZERO planning were preferred over standard ChatGPT responses up to 59.32% of the time, with larger search simulation counts yielding stronger preferences. In live interactive human evaluations, GDP-ZERO achieved the highest scores across all persuasion metrics: it achieved a donation probability of 0.79 (compared to 0.73 for ChatGPT and 0.72 for RAP), produced significantly stronger arguments (4.28 out of 5), and was rated significantly more convincing (4.38) and more natural (4.38). Additionally, human evaluators rated GDP-ZERO as significantly less manipulative (2.29 out of 5) than both baseline systems. The analysis revealed that GDP-ZERO achieves this by adopting a more patient, conservative strategy, avoiding premature requests for donations and balancing logical appeals, emotional appeals, and credibility evidence.

These results show that high-performing conversational planning can be achieved entirely without domain-specific model training, dramatically reducing data collection and engineering costs. This makes advanced conversational agents viable in data-scarce domains where collecting expert dialogue demonstrations is prohibitive. However, decision-makers must weigh this against practical trade-offs. GDP-ZERO introduces notable computational latency; in interactive tests, performing 10 look-ahead simulations required approximately 35 seconds per conversational turn, which can impact real-time user experience.

For organizations considering deployment, the source supports piloting zero-training planning systems in strategic, high-stakes communication settings where turn latency is acceptable or where back-end systems can parallelize simulation queries. Future development should focus on optimizing search runtime, exploring tree-parallelization methods, and testing the framework across other complex dialogue domains beyond charitable persuasion. While results are statistically significant within the evaluated benchmark, stakeholders should exercise caution regarding response latency and potential misalignment risks before deploying simulation-driven planners in consumer-facing environments.

arXiv: 2305.13660
Cover for Prompt-Based Monte-Carlo Tree Search for Goal-oriented Dialogue Policy Planning

Abstract

Planning for goal-oriented dialogue often requires simulating future dialogue interactions and estimating task progress. Many approaches thus consider training neural networks to perform look-ahead search algorithms such as A* search and Monte Carlo Tree Search (MCTS). However, this training often requires abundant annotated data, which creates challenges when faced with noisy annotations or low-resource settings. We introduce GDP-ZERO, an approach using Open-Loop MCTS to perform goal-oriented dialogue policy planning without any model training. GDP-ZERO prompts a large language model to act as a policy prior, value function, user simulator, and system model during the tree search. We evaluate GDP-ZERO on the goal-oriented task PersuasionForGood, and find that its responses are preferred over ChatGPT up to 59.32% of the time, and are rated more persuasive than ChatGPT during interactive evaluations1.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Problem Definition
  • 3.2 Dialogue Planning as a Stochastic MDP
  • 3.3 GDP-ZERO
  • 4 Experiments
  • 4.1 Static Evaluation
  • 4.2 Interactive Human Evaluation
  • 4.3 Ablation Studies
  • 5 Conclusion
  • 6 Limitations
  • 7 Ethical Considerations
  • References
  • A Additional details on GDP-ZERO
  • B Prompting Details on P4G
  • C Ablation Studies
  • D Analysis of GDP-ZERO Dialogues
  • E GDP-ZERO Setup on P4G
  • F Additional details on static evaluation
  • G Additional details on interactive study
  • H Additional details on survey results
  • I Example Interactive Conversations

Knowls

  1. Knowl 1 — GDP-ZERO: zero-training dialogue policy planning

    model/method

    GDP-ZERO, or Goal-oriented Dialogue Planning with Zero training, performs dialogue policy planning at decision time by combining Open-Loop Monte-Carlo Tree Search with a prompted generative large language model (LLM). Given the current dialogue history and a fixed dialogue-action space, GDP-ZERO runs multiple simulated futures and uses the LLM in four roles: prior policy, system-response generator, user simulator, and value estimator. It returns both the next system dialogue act and a system utterance associated with that act.

    GDP-ZERO requires no task-specific model training or annotated policy data. Few-shot prompting supplies the behavioral knowledge needed for simulation, while Open-Loop MCTS prevents one accidentally generated system or user utterance from permanently defining an entire search subtree: utterances are regenerated when action sequences are revisited.

  2. Knowl 2 — Stochastic MDP formulation for dialogue policy planning

    model/method

    A dialogue policy is formulated as a Markov decision process ⟨S,A,R,P,γ⟩\langle S,A,R,P,\gamma\rangle. For a dialogue through turn tt, the history is

    ht=(a0sys,u1sys,u1usr,…,at−1sys,utsys,utusr),h_t=(a^{\mathrm{sys}}_0,u^{\mathrm{sys}}_1,u^{\mathrm{usr}}_1,\ldots,a^{\mathrm{sys}}_{t-1},u^{\mathrm{sys}}_t,u^{\mathrm{usr}}_t),

    where ajsysa^{\mathrm{sys}}_j is the system dialogue act at turn jj, and ujsysu^{\mathrm{sys}}_j and ujusru^{\mathrm{usr}}_j are the system and user utterances. The state at turn ii is the dialogue history si=(a0,u1sys,u1usr,…,ai−1,uisys,uiusr)∈Ss_i=(a_0,u^{\mathrm{sys}}_1,u^{\mathrm{usr}}_1,\ldots,a_{i-1},u^{\mathrm{sys}}_i,u^{\mathrm{usr}}_i)\in S, and a system dialogue act ai∈Aa_i\in A is the action. R(s,a)R(s,a) measures the likelihood of achieving the desired conversational outcome, P:S×A→SP:S\times A\rightarrow S is the transition distribution, and γ∈[0,1)\gamma\in[0,1) is the discount factor.

    Dialogue transitions are treated as stochastic because the same history and system action can produce many plausible user and system responses. Thus, the next state sampled during planning may differ across simulations and need not be the most probable response realization.

  3. Knowl 3 — Open-Loop tree representation and its purpose

    model/method

    GDP-ZERO represents each search node by an action sequence rather than by a particular dialogue transcript. A node at depth ii has the form sitr=(a0,…,ai)s_i^{\mathrm{tr}}=(a_0,\ldots,a_i), while a cache attached to that node stores several generated dialogue histories consistent with the action sequence. When the tree is traversed, GDP-ZERO samples a cached history if enough histories have already been generated; otherwise, it prompts the LLM to regenerate the system and user utterances needed to continue the simulation.

    This representation is designed for stochastic dialogue transitions. A closed-loop tree would attach one sampled transcript to a node, so an improbable response could contaminate every descendant simulation. The open-loop representation instead aggregates statistics over repeated executions of the same action sequence and continually regenerates utterances, reducing the propagation of errors from any single simulation.

  4. Knowl 4 — GDP-ZERO search, evaluation, and response selection

    algorithm

    GDP-ZERO takes a generative LLM MθM_\theta, the current dialogue history hih_i, a finite action set AA, the number of simulations nn, cache size kk, exploration coefficient cpc_p, and initial action value Q0Q_0. It maintains visit counts N(str,a)N(s^{\mathrm{tr}},a), action values Q(str,a)Q(s^{\mathrm{tr}},a), action-prior probabilities p(a∣str)p(a\mid s^{\mathrm{tr}}), cached histories, and cached-history values.

    At a tree node, selection chooses the action maximizing the PUCT score

    PUCT⁡(str,a)=Q(str,a)+cp p(a∣str)∑b∈AN(str,b)1+N(str,a),\operatorname{PUCT}(s^{\mathrm{tr}},a)=Q(s^{\mathrm{tr}},a)+c_p\,p(a\mid s^{\mathrm{tr}})\frac{\sqrt{\sum_{b\in A}N(s^{\mathrm{tr}},b)}}{1+N(s^{\mathrm{tr}},a)},

    where a,b∈Aa,b\in A, NN is the number of visits to an action edge, QQ is its running value estimate, pp is the LLM-derived prior, and cpc_p controls exploration. The selected action is appended to the open-loop action sequence. A cached history is sampled at each step, and a new history is generated until the node cache reaches size kk; traversal continues until a leaf is reached.

    At a leaf, the LLM is prompted to sample a prior distribution over next dialogue acts. Every newly expanded act is initialized with Q=Q0Q=Q_0 and N=0N=0. The LLM then estimates the value v(str)v(s^{\mathrm{tr}}) of the leaf by judging the simulated user’s current inclination to complete the task. Along the path back to the root, visit counts are incremented and action values are updated as running averages:

    N(str,a)←N(str,a)+1,Q(str,a)←Q(str,a)+v(str)−Q(str,a)N(str,a).N(s^{\mathrm{tr}},a)\leftarrow N(s^{\mathrm{tr}},a)+1, \qquad Q(s^{\mathrm{tr}},a)\leftarrow Q(s^{\mathrm{tr}},a)+\frac{v(s^{\mathrm{tr}})-Q(s^{\mathrm{tr}},a)}{N(s^{\mathrm{tr}},a)}.

    After nn simulations, GDP-ZERO selects the root action a∗=arg⁡max⁡a∈AN(s0tr,a)a^*=\arg\max_{a\in A}N(s^{\mathrm{tr}}_0,a). It then performs response selection: among cached histories produced after executing a∗a^*, it returns the system utterance from the history with the highest running value estimate. This avoids generating the chosen action’s response afresh and selects an utterance that was empirically favorable during search.

  5. Knowl 5 — Prompted LLM components and P4G value construction

    model/method

    For PersuasionForGood, GDP-ZERO uses one few-shot example while changing the prompt format for each operation. To generate a system response, the prompt includes the natural-language description of the selected dialogue act. To simulate a user response, the system and user roles are swapped and the LLM is instructed to produce one of five labeled user reactions: no donation, negative reaction, neutral, positive reaction, or donation.

    For value estimation, GDP-ZERO appends the question Would you be interested in donating to Save the Children? to a simulated history and samples the LLM l=10l=10 times at temperature 1.11.1. The reaction labels are mapped to numerical values: no donation =−1.0=-1.0, negative reaction =−0.5=-0.5, neutral =0.0=0.0, positive reaction =0.5=0.5, and donation =1.0=1.0. The mean of the sampled values is the state value in the range [−1,1][-1,1].

    For the prior policy, the LLM is sampled 1515 times at temperature 1.01.0 to obtain next-action labels. Their frequencies are converted to a probability distribution using add-one smoothing, so every action retains nonzero prior probability. In the reported implementation, cp=1.0c_p=1.0 and the tested initial values are Q0∈{0.0,0.25,0.5}Q_0\in\{0.0,0.25,0.5\}.

  6. Knowl 6 — PersuasionForGood action space and evaluation setting

    experimental setup

    The main evaluation uses PersuasionForGood, a dataset of 300 annotated human-human dialogues in which a persuader attempts to convince a persuadee to donate to Save the Children. Because some human-human strategies are unsuitable for a chatbot, GDP-ZERO plans over seven actions: logical appeal, emotion appeal, credibility appeal, task-related inquiry, proposition of donation, greeting, and other. The first four are persuasive strategies; proposition of donation, greeting, and other are non-strategy actions used to handle solicitation and unaccounted conversational situations.

    The generation backbone is ChatGPT, specifically the gpt-3.5-turbo version available in April 2023. The system plans one dialogue act at each turn and then prompts ChatGPT to produce the corresponding utterance. Evaluation covers both fixed-context static comparisons and end-to-end interactive conversations.

  7. Knowl 7 — Static persuasion results and search-budget effects

    data/table

    In the static evaluation, the first 20 PersuasionForGood dialogues supplied 154 evaluation turns. For each turn, ChatGPT judged which of two generated responses was more persuasive; each judgment was sampled five times and decided by majority vote, with response order randomly swapped to reduce a strong option-order bias. Results were repeated three times and reported as mean ±\pm standard deviation.

    Against the human ground-truth responses, direct ChatGPT prompting achieved a win rate of 88.84±0.75%88.84\pm0.75\%. GDP-ZERO achieved 87.22±0.61%87.22\pm0.61\% with n=5n=5, 90.69±1.60%90.69\pm1.60\% with n=10n=10, 88.86±1.24%88.86\pm1.24\% with n=20n=20, and 89.82±1.10%89.82\pm1.10\% with n=50n=50 simulations; the best GDP-ZERO setting therefore slightly exceeded direct prompting in this comparison.

    Against direct ChatGPT responses, the GDP-ZERO win rates and average runtimes were: 50.65±3.31%50.65\pm3.31\% and 18 seconds for n=5,k=3,Q0=0n=5,k=3,Q_0=0; 50.86±1.10%50.86\pm1.10\% and 36 seconds for n=10,k=3,Q0=0n=10,k=3,Q_0=0; 53.24±1.91%53.24\pm1.91\% and 75 seconds for n=20,k=3,Q0=0n=20,k=3,Q_0=0; and 59.32±1.84%59.32\pm1.84\% and 740 seconds for n=50,k=3,Q0=0n=50,k=3,Q_0=0. With n=10,k=3n=10,k=3, changing Q0Q_0 to 0.250.25 produced 57.79±2.95%57.79\pm2.95\% at 36 seconds, while Q0=0.50Q_0=0.50 produced 53.03±2.00%53.03\pm2.00\%. With n=10,Q0=0n=10,Q_0=0, cache sizes k=1k=1 and k=2k=2 produced 49.57±2.01%49.57\pm2.01\% and 51.30±1.59%51.30\pm1.59\%, respectively, with runtimes of 16 and 29 seconds. These results show that larger search budgets generally improved preference over direct ChatGPT, but at a substantial latency cost.

  8. Knowl 8 — Interactive human evaluation of end-to-end persuasion

    data/table

    The end-to-end evaluation used Amazon Mechanical Turk crowdworkers interacting with three ChatGPT-based systems: direct ChatGPT planning, GDP-ZERO planning, and RAP, a rule-based planner derived from expert knowledge. The study retained 40 surveys for GDP-ZERO, 35 for ChatGPT, and 36 for RAP. GDP-ZERO used n=10n=10, k=3k=3, and Q0=0.25Q_0=0.25; the same ChatGPT backbone generated the final responses. Scores were means ±\pm standard deviations; all rating scales were from 1 to 5 except donation probability, which was from 0 to 1.

    The results were as follows, listed in the order RAP, ChatGPT, GDP-ZERO: donation probability was 0.72±0.380.72\pm0.38, 0.73±0.380.73\pm0.38, and 0.79±0.370.79\pm0.37; increased donation intent was 4.08±0.684.08\pm0.68, 3.77±0.903.77\pm0.90, and 4.30±0.714.30\pm0.71; strong argument was 3.89±0.973.89\pm0.97, 3.91±0.993.91\pm0.99, and 4.28±0.744.28\pm0.74; convincingness was 4.11±0.744.11\pm0.74, 4.10±0.704.10\pm0.70, and 4.38±0.664.38\pm0.66; strategy diversity was 3.98±0.803.98\pm0.80, 3.83±1.033.83\pm1.03, and 3.95±0.823.95\pm0.82; perceived manipulation was 2.64±1.362.64\pm1.36, 2.96±1.382.96\pm1.38, and 2.29±1.332.29\pm1.33; naturalness was 4.25±0.684.25\pm0.68, 4.03±0.654.03\pm0.65, and 4.38±0.624.38\pm0.62; relevance was 4.64±0.544.64\pm0.54, 4.31±0.864.31\pm0.86, and 4.59±0.494.59\pm0.49; and coherence was 4.28±0.654.28\pm0.65, 4.06±0.894.06\pm0.89, and 4.42±0.494.42\pm0.49.

    GDP-ZERO received the strongest score on every reported persuasiveness-related criterion and was rated significantly better in increased donation intent, strong argument, convincingness, perceived manipulation, naturalness, and coherence according to the reported significance marks (p<0.05p<0.05 for one asterisk and p<0.01p<0.01 for two). RAP received the highest relevance score, while its strategy diversity was also competitive.

  9. Knowl 9 — Ablation evidence for Open-Loop search and response selection

    empirical result

    Ablations used the first 20 PersuasionForGood dialogues, n=20n=20, cp=1c_p=1, Q0=0Q_0=0, and k=3k=3, with ChatGPT judging the generated responses against ground truth. With Codex as the generation backbone, direct prompting achieved 38.09±2.00%38.09\pm2.00\%, GDP-ZERO achieved 45.46±2.95%45.46\pm2.95\%, GDP-ZERO without Open-Loop search achieved 39.16±3.42%39.16\pm3.42\%, and GDP-ZERO without response selection achieved 40.80±1.47%40.80\pm1.47\%.

    With ChatGPT as the backbone, direct prompting achieved 87.21±0.60%87.21\pm0.60\%, GDP-ZERO achieved 91.13±0.30%91.13\pm0.30\%, GDP-ZERO without Open-Loop search achieved 88.09±0.81%88.09\pm0.81\%, and GDP-ZERO without response selection achieved 91.03±0.75%91.03\pm0.75\%. For this evaluation, only the first three sentences of each ChatGPT generation were judged. Across both backbones, the comparisons support the contribution of Open-Loop MCTS, and they also indicate that selecting a cached response after planning is beneficial.

  10. Knowl 10 — Scope, latency, and simulation-quality limitations

    limitation

    The paper evaluates GDP-ZERO primarily on PersuasionForGood, so its effectiveness beyond persuasion is not established. The authors argue that it is most appropriate for tasks requiring long-horizon planning, lacking reliable optimal-action annotations, and permitting genuine interactive evaluation rather than hypothetical role-play. Applying it directly to hierarchical task-oriented dialogue could create a very large multi-label action space; the paper suggests that separate high- and low-level search trees may be needed.

    Runtime is a major limitation: larger simulation counts and cache sizes improve the opportunity to find useful policies but increase latency. With the OpenAI API rate limit, the interactive system planned over seven actions using n=10n=10 and k=3k=3, taking approximately 35 seconds per turn. The LLM can also generate utterances that do not match the planned action, and increasing kk helps compensate by providing multiple simulations but further increases runtime.

    The static evaluator may favor ChatGPT-generated responses. In an analysis of GDP-ZERO versus direct ChatGPT responses, response length correlated positively with ChatGPT’s persuasiveness preference (r=0.29r=0.29, p<0.001p<0.001), although interactive crowdworker evaluation still favored GDP-ZERO. Finally, because GDP-ZERO is goal-agnostic, the same planning capability could be applied to manipulative or unethical goals such as scams.

Coverage note — The detailed dialogue-act frequency plots, qualitative example conversations, complete prompt transcripts, and the annotated action-count table were omitted because they illustrate the method and results without adding separate load-bearing contributions.

References

  1. 1.James E Allen, Curry I Guinn, and Eric Horvitz. 1999. Mixed-initiative interaction. IEEE Intelligent Systems and their Applications, 14(5):14–23.
  2. 2.Sanghwan Bae, Donghyun Kwak, Sungdong Kim, Donghoon Ham, Soyoung Kang, Sang-Woo Lee, and Woomyoung Park. 2022. Building a role specified open-domain dialogue system leveraging large-scale language models. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2128–2150.
  3. 3.Paweł Budzianowski, Tsung-Hsien Wen, Bo-Hsiang Tseng, Iñigo Casanueva, Ultes Stefan, Ramadan Osman, and Milica Gašic. 2018. Multiwoz - a large-scale multi-domain wizard-of-oz dataset for task-oriented dialogue modelling. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (EMNLP).
  4. 4.Yan Cao, Keting Lu, Xiaoping Chen, and Shiqi Zhang. 2020. Adaptive dialog policy learning with hindsight and user modeling. In Proceedings of the 21th Annual Meeting of the Special Interest Group on Discourse and Dialogue, pages 329–338, 1st virtual meeting. Association for Computational Linguistics.
  5. 5.Guillaume MJ B Chaslot, Mark HM Winands, and H Jaap van Den Herik. 2008. Parallel monte-carlo tree search. In Computers and Games: 6th International Conference, CG 2008, Beijing, China, September 29-October 1, 2008. Proceedings 6, pages 60–71. Springer.
  6. 6.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, Dave Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert-Voss, William Hebgen Guss, Alex Nichol, Alex Paino, Nikolas Tezak, Jie Tang, Igor Babuschkin, Suchir Balaji, Shantanu Jain, William Saunders, Christopher Hesse, Andrew N. Carr, Jan Leike, Josh Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, and Wojciech Zaremba. 2021. Evaluating large language models trained on code.
  7. 7.Maximillian Chen, Alexandros Papangelis, Chenyang Tao, Seokhwan Kim, Andy Rosenbaum, Yang Liu, Zhou Yu, and Dilek Hakkani-Tur. 2023a. PLACES: Prompting language models for social conversation synthesis. Findings of the Association for Computational Linguistics: EACL 2023.
  8. 8.Maximillian Chen, Weiyan Shi, Feifan Yan, Ryan Hou, Jingwen Zhang, Saurav Sahay, and Zhou Yu. 2022. Seamlessly integrating factual information and social content with persuasive dialogue. In Proceedings of the 2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 12th International Joint Conference on Natural Language Processing, pages 399–413.
  9. 9.Maximillian Chen, Xiao Yu, Weiyan Shi, Urvi Awasthi, and Zhou Yu. 2023b. Controllable mixed-initiative dialogue generation through prompting. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 951–966, Toronto, Canada. Association for Computational Linguistics.
  10. 10.Yi Cheng, Wenge Liu, Wenjie Li, Jiashuo Wang, Ruihui Zhao, Bang Liu, Xiaodan Liang, and Yefeng Zheng. 2022. Improving multi-turn emotional support dialogue generation with lookahead strategy planning. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 3014–3026, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  11. 11.Fabrizio Gilardi, Meysam Alizadeh, and Maël Kubli. 2023. Chatgpt outperforms crowd-workers for text-annotation tasks.
  12. 12.Biyang Guo, Xin Zhang, Ziyuan Wang, Minqi Jiang, Jinran Nie, Yuxuan Ding, Jianwei Yue, and Yupeng Wu. 2023. How close is chatgpt to human experts? comparison corpus, evaluation, and detection.
  13. 13.Prakhar Gupta, Cathy Jiao, Yi-Ting Yeh, Shikib Mehri, Maxine Eskenazi, and Jeffrey P Bigham. 2022. Instructdial: Improving zero and few-shot generalization in dialogue through instruction tuning. EMNLP.
  14. 14.He He, Derek Chen, Anusha Balakrishnan, and Percy Liang. 2018. Decoupling strategy and generation in negotiation dialogues.
  15. 15.Wanwei He, Yinpei Dai, Yinhe Zheng, Yuchuan Wu, Zheng Cao, Dermot Liu, Peng Jiang, Min Yang, Fei Huang, Luo Si, et al. 2022. Galaxy: A generative pre-trained model for task-oriented dialog with semi-supervised learning and explicit policy injection. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, pages 10749–10757.
  16. 16.Xingwei He, Zhenghao Lin, Yeyun Gong, A Jin, Hang Zhang, Chen Lin, Jian Jiao, Siu Ming Yiu, Nan Duan, Weizhu Chen, et al. 2023. Annollm: Making large language models to be better crowdsourced annotators. arXiv preprint arXiv:2303.16854.
  17. 17.Ronald A Howard. 1960. Dynamic programming and markov processes.
  18. 18.Youngsoo Jang, Jongmin Lee, and Kee-Eung Kim. 2020. Bayes-adaptive monte-carlo planning and learning for goal-oriented dialogues. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 7994–8001.
  19. 19.Hyunwoo Kim, Jack Hessel, Liwei Jiang, Ximing Lu, Youngjae Yu, Pei Zhou, Ronan Le Bras, Malihe Alikhani, Gunhee Kim, Maarten Sap, et al. 2022. Soda: Million-scale dialogue distillation with social commonsense contextualization. arXiv preprint arXiv:2212.10465.
  20. 20.Esther Levin, Roberto Pieraccini, and Wieland Eckert. 1997. Learning dialogue strategies within the markov decision process framework. In 1997 IEEE Workshop on Automatic Speech Recognition and Understanding Proceedings, pages 72–79. IEEE.
  21. 21.Mike Lewis, Denis Yarats, Yann Dauphin, Devi Parikh, and Dhruv Batra. 2017. Deal or no deal? end-to-end learning of negotiation dialogues. In Conference on Empirical Methods in Natural Language Processing.
  22. 22.Yu Li, Josh Arnold, Feifan Yan, Weiyan Shi, and Zhou Yu. 2021. Legoeval: An open-source toolkit for dialogue system evaluation via crowdsourcing. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing: System Demonstrations, pages 317–324.
  23. 23.Bing Liu and Ian Lane. 2017. Iterative policy learning in end-to-end trainable task-oriented neural dialog models. In 2017 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), pages 482–489. IEEE.
  24. 24.Bing Liu, Gökhan Tür, Dilek Hakkani-Tur, Pararth Shah, and Larry Heck. 2018. Dialogue learning with human teaching and feedback in end-to-end trainable task-oriented dialogue systems. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 2060–2069.
  25. 25.Siyang Liu, Chujie Zheng, Orianna Demasi, Sahand Sabour, Yu Li, Zhou Yu, Yong Jiang, and Minlie Huang. 2021. Towards emotional support dialog systems. In Proceedings of the 59th annual meeting of the Association for Computational Linguistics.
  26. 26.Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023a. Gpтеval: Nlg evaluation using gpt-4 with better human alignment. arXiv preprint arXiv:2303.16634.
  27. 27.Yiheng Liu, Tianle Han, Siyuan Ma, Jiayue Zhang, Yuanyuan Yang, Jiaming Tian, Hao He, Antong Li, Mengshen He, Zhengliang Liu, Zihao Wu, Dajiang Zhu, Xiang Li, Ning Qiang, Dingang Shen, Tianming Liu, and Bao Ge. 2023b. Summary of chatgpt/gpt-4 research and perspective towards the future of large language models.
  28. 28.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach.
  29. 29.Yiren Liu and Halil Kilicoglu. 2023. Commonsense-aware prompting for controllable empathetic dialogue generation. arXiv preprint arXiv:2302.01441.
  30. 30.Zihan Liu, Mostofa Patwary, Ryan Prenger, Shrimai Prabhumoye, Wei Ping, Mohammad Shoeybi, and Bryan Catanzaro. 2022. Multi-stage prompting for knowledgeable dialogue generation. In Findings of the Association for Computational Linguistics: ACL 2022, pages 1317–1337.
  31. 31.Andrea Madotto, Zhaojiang Lin, Genta Indra Winata, and Pascale Fung. 2021. Few-shot bot: Prompt-based learning for dialogue systems. arXiv preprint arXiv:2110.08118.
  32. 32.Shikib Mehri and Maxine Eskenazi. 2021. Schema-guided paradigm for zero-shot dialog. In Proceedings of the 22nd Annual Meeting of the Special Interest Group on Discourse and Dialogue, pages 499–508, Singapore and Online. Association for Computational Linguistics.
  33. 33.OpenAI. 2022. Openai: Introducing chatgpt.
  34. 34.Alexander Pan, Chan Jun Shern, Andy Zou, Nathaniel Li, Steven Basart, Thomas Woodside, Jonathan Ng, Hanlin Zhang, Scott Emmons, and Dan Hendrycks. 2023. Do the rewards justify the means? measuring trade-offs between rewards and ethical behavior in the machiavelli benchmark. arXiv preprint arXiv:2304.03279.
  35. 35.Baolin Peng, Xiujun Li, Jianfeng Gao, Jingjing Liu, and Kam-Fai Wong. 2018. Deep Dyna-Q: Integrating planning for task-completion dialogue policy learning. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2182–2192, Melbourne, Australia. Association for Computational Linguistics.
  36. 36.Diego Perez Liebana, Jens Dieskau, Martin Hunermund, Sanaz Mostaghim, and Simon Lucas. 2015. Open loop search for general video game playing. In Proceedings of the 2015 Annual Conference on Genetic and Evolutionary Computation, GECCO ’15, page 337–344, New York, NY, USA. Association for Computing Machinery.
  37. 37.Christopher D Rosin. 2011. Multi-armed bandits with episode context. Annals of Mathematics and Artificial Intelligence, 61(3):203–230.
  38. 38.Weiyan Shi, Kun Qian, Xuewei Wang, and Zhou Yu. 2019. How to build user simulators to train RL-based dialog systems. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 1990–2000, Hong Kong, China. Association for Computational Linguistics.
  39. 39.David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis. 2016. Mastering the game of go with deep neural networks and tree search. Nature, 529(7587):484–489.
  40. 40.David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, Yutian Chen, Timothy Lillicrap, Fan Hui, Laurent Sifre, George van den Driessche, Thore Graepel, and Demis Hassabis. 2017. Mastering the game of Go without human knowledge. Nature, 550(7676):354–359.
  41. 41.Richard S Sutton and Andrew G Barto. 2018. Reinforcement learning: An introduction. MIT press.
  42. 42.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023. Llama: Open and efficient foundation language models.
  43. 43.Dirk Väth, Lindsey Vanderlyn, and Ngoc Thang Vu. 2023. Conversational tree search: A new hybrid dialog task. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, pages 1264–1280, Dubrovnik, Croatia. Association for Computational Linguistics.
  44. 44.Shuohang Wang, Yang Liu, Yichong Xu, Chenguang Zhu, and Michael Zeng. 2021. Want to reduce labeling cost? gpt-3 can help. arXiv preprint arXiv:2108.13487.
  45. 45.Sihan Wang, Kaijie Zhou, Kunfeng Lai, and Jianping Shen. 2020. Task-completion dialogue policy learning via Monte Carlo tree search with dueling network. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 3461–3471, Online. Association for Computational Linguistics.
  46. 46.Xuewei Wang, Weiyan Shi, Richard Kim, Yoojung Oh, Sijia Yang, Jingwen Zhang, and Zhou Yu. 2019. Persuasion for good: Towards a personalized persuasive dialogue system for social good. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 5635–5649, Florence, Italy. Association for Computational Linguistics.
  47. 47.Yizhong Wang, Hamish Ivison, Pradeep Dasigi, Jack Hessel, Tushar Khot, Khyathi Raghavi Chandu, David Wadden, Kelsey MacMillan, Noah A. Smith, Iz Beltagy, and Hannaneh Hajishirzi. 2023. How far can camels go? exploring the state of instruction tuning on open resources.
  48. 48.Richard Weber. 2010. Optimization and control. University of Cambridge.
  49. 49.Jingxuan Yang, Si Li, and Jun Guo. 2021. Multi-turn target-guided topic prediction with Monte Carlo tree search. In Proceedings of the 18th International Conference on Natural Language Processing (ICON), pages 324–334, National Institute of Technology Silchar, Silchar, India. NLP Association of India (NLPAI).
  50. 50.Yuting Yang, Wenqiang Lei, Juan Cao, Jintao Li, and Tat-Seng Chua. 2022. Prompt learning for few-shot dialogue state tracking. arXiv preprint arXiv:2201.05780.
  51. 51.Cong Zhang, Huilin Jin, Jienan Chen, Jinkuan Zhu, and Jinting Luo. 2020a. A hierarchy mcts algorithm for the automated pcb routing. In 2020 IEEE 16th International Conference on Control & Automation (ICCA), pages 1366–1371.
  52. 52.Haodi Zhang, Zhichao Zeng, Keting Lu, Kaishun Wu, and Shiqi Zhang. 2022a. Efficient dialog policy learning by reasoning with contextual knowledge. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, pages 11667–11675.
  53. 53.Qiang Zhang, Jason Naradowsky, and Yusuke Miyao. 2023. Ask an expert: Leveraging language models to improve strategic reasoning in goal-oriented dialogue models.
  54. 54.Shuo Zhang, Junzhou Zhao, Pinghui Wang, Yu Li, Yi Huang, and Junlan Feng. 2022b. " think before you speak": Improving multi-action dialog policy by planning single-action dialogs. arXiv preprint arXiv:2204.11481.
  55. 55.Zheng Zhang, Lizi Liao, Xiaoyan Zhu, Tat-Seng Chua, Zitao Liu, Yan Huang, and Minlie Huang. 2020b. Learning goal-oriented dialogue policy with opposite agent awareness. In Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing, pages 122–132.
  56. 56.Tiancheng Zhao and Maxine Eskenazi. 2018. Zero-shot dialog generation with cross-domain latent actions. In Proceedings of the 19th Annual SIGdial Meeting on Discourse and Dialogue, pages 1–10.

Citation

MLA
Yu, X., et al. “Prompt-Based Monte-Carlo Tree Search for Goal-oriented Dialogue Policy Planning”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 7101–25, https://doi.org/10.18653/v1/2023.emnlp-main.439.
APA
Yu, X., Chen, M., & Yu, Z. (2023). Prompt-Based Monte-Carlo Tree Search for Goal-oriented Dialogue Policy Planning. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 7101–7125. https://doi.org/10.18653/v1/2023.emnlp-main.439
Chicago
Yu, X., M. Chen, and Z. Yu. 2023. “Prompt-Based Monte-Carlo Tree Search for Goal-oriented Dialogue Policy Planning”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 7101–25. https://doi.org/10.18653/v1/2023.emnlp-main.439.
Harvard
Yu, X., Chen, M. and Yu, Z. (2023) “Prompt-Based Monte-Carlo Tree Search for Goal-oriented Dialogue Policy Planning”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 7101–7125. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.439.
Vancouver
1. Yu X, Chen M, Yu Z (2023) Prompt-Based Monte-Carlo Tree Search for Goal-oriented Dialogue Policy Planning. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 7101–7125

BibTeX

@inproceedings{yu-etal-2023-prompt,
    title = "Prompt-Based {M}onte-{C}arlo Tree Search for Goal-oriented Dialogue Policy Planning",
    author = "Yu, Xiao  and
      Chen, Maximillian  and
      Yu, Zhou",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.439/",
    doi = "10.18653/v1/2023.emnlp-main.439",
    pages = "7101--7125"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/