Evaluating Large Language Models as Generative User Simulators for Conversational Recommendation

Se-eun YoonZhankui HeJessica Maria EchterhoffJulian J. McAuley

article2024NAACL79 citations

Establishes the first standardized evaluation protocol comprising five core conversational tasks to measure how accurately large language models simulate diverse human behavior in conversational recommender systems and provides practical strategies to minimize behavioral discrepancies.

Listen

Conversational recommendation systems require extensive evaluation to ensure they effectively understand user needs and provide helpful suggestions. Testing these systems directly with human participants is expensive, time-consuming, and carries operational risks, while conventional offline benchmarks fail to capture interactive, multi-turn conversations. To resolve this dilemma, researchers and system developers are turning to synthetic users powered by large language models as automated stand-ins. However, prior to this study, there was no standardized framework to measure how accurately language model simulations replicate genuine human preferences and conversational behaviors across a broader population.

The main objective of the article is to establish and demonstrate the first systematic evaluation protocol for measuring the degree to which large language models can realistically emulate human behavior in conversational recommendation. To achieve this, the authors designed a zero-shot framework composed of five core tasks: selecting items to discuss, expressing binary preferences, providing open-ended opinions, formulating recommendation requests, and delivering feedback on suggested items. The authors evaluated several baseline language models—including GPT-3.5, GPT-4, and text-davinci-003—using real-world movie datasets gathered from ReDial, Reddit, MovieLens, and IMDB containing tens of thousands of conversations and reviews.

The investigation revealed critical distortions in how standard language models simulate human behavior. First, baseline models show a heavy popularity bias, generating significantly less diverse item mentions than humans; however, conditioning models on a user's interaction history substantially restored item diversity. Second, standard models exhibited an overly optimistic bias and correlated poorly with human preferences across movie rating tiers, with correlations near zero or undefined. Incorporating explicit "pickiness" personality traits into prompts sharply improved preference alignment, raising correlation coefficients up to 0.75–0.76 with GPT-4. Third, synthetic recommendation requests lacked individual personalization; the most diverse model produced requests that were roughly 23% less diverse than human requests, frequently recycling generic buzzwords like "gripping" and "mind-bending." Finally, while simulators provided coherent feedback on recommendations between 65% and 91% of the time, they occasionally rejected valid recommendations due to a failure to grasp subtle contextual nuances.

These findings demonstrate that off-the-shelf language models cannot be assumed to represent human populations out of the box. Relying on uncalibrated synthetic users creates significant risks of building conversational recommendation systems tuned to artificial, overly optimistic, and generic behaviors. However, the results prove that targeted prompting strategies—specifically integrating past interaction histories and explicit personality traits such as pickiness—alongside capable base models can effectively bridge the gap between artificial simulators and human behavior.

Decision-makers and engineering teams should adopt structured multi-task benchmarking protocols before using synthetic agents to evaluate user-facing conversational systems. Organizations should implement enriched persona prompts that specify interaction histories and varied preference thresholds rather than relying on default model settings. Furthermore, synthetic testing should be positioned as an efficient pre-deployment screening tool rather than a total replacement for live user testing. Future work should expand this evaluation framework beyond movies into broader commercial domains, test open-source foundation models, and develop more sophisticated methods to help synthetic users interpret subtle linguistic nuances in user requests.

Cover for Evaluating Large Language Models as Generative User Simulators for Conversational Recommendation

Abstract

Synthetic users are cost-effective proxies for real users in the evaluation of conversational recommender systems. Large language models show promise in simulating human-like behavior, raising the question of their ability to represent a diverse population of users. We introduce a new protocol to measure the degree to which language models can accurately emulate human behavior in conversational recommendation. This protocol is comprised of five tasks, each designed to evaluate a key property that a synthetic user should exhibit: choosing which items to talk about, expressing binary preferences, expressing open-ended preferences, requesting recommendations, and giving feedback. Through evaluation of baseline simulators, we demonstrate these tasks effectively reveal deviations of language models from human behavior, and offer insights on how to reduce the deviations with model selection and prompting strategies.

Table of Contents

  • 1 Introduction
  • 2 Evaluation Tasks
  • 3 Methods
  • 4 Experiments
  • 5 Related Work
  • 6 Conclusion
  • Limitations
  • Ethics Statement
  • Acknowledgements
  • References
  • A Appendix
  • A.1 Dataset statistics
  • A.2 Prompts
  • A.2.1 ItemsTalk
  • A.2.2 BinPref
  • A.2.3 OpenPref
  • A.2.4 RecRequest
  • A.2.5 Feedback
  • A.3 More results
  • A.3.1 Results from ItemsTalk
  • A.3.2 Results from BinPref
  • A.3.3 More Feedback examples

Knowls

  1. Knowl 1 — Five-Task Evaluation Protocol for Conversational Recommender User Simulators

    model/method

    The evaluation protocol for measuring the degree to which Large Language Models (LLMs) can act as generative user simulators in conversational recommender systems (CRSs) decomposes evaluation into five zero-shot tasks. Simulators are treated as black-box natural language generators that take free-form text input and return free-form text output across a simulated population of users:

    1. ItemsTalk (Choosing items to talk about): Measures the diversity and distribution of items mentioned by a population of simulators against human distributions extracted from conversational datasets (ReDial, Reddit, IMDB).
    2. BinPref (Expressing binary preference): Evaluates how well simulators reflect human item ratings by querying simulated users on whether they liked specific movies (using frequent vs. infrequent items from MovieLens) and computing correlation against ground-truth user rating averages.
    3. OpenPref (Expressing open-ended preference): Evaluates fine-grained opinions on item aspects (e.g., plot, cast) by performing aspect-based sentiment analysis on generated open-ended text compared to real IMDB user reviews.
    4. RecRequest (Requesting recommendations): Assesses whether synthetic recommendation requests achieve the lexical, semantic, and granular diversity of real human requests on Reddit when conditioned on fixed anchor items and character lengths.
    5. Feedback (Giving feedback): Tests whether simulators produce coherent feedback to CRS recommendations by evaluating their acceptance/rejection of relevant versus irrelevant recommendations, and preference between pairs of recommendations with and without explanations.
  2. Knowl 2 — Prompting Configurations for Conversational User Simulators

    model/method

    To evaluate user simulator behavior under different conditioning paradigms without fine-tuning, four prompting configurations are defined:

    • Vanilla LLM: The language model is prompted without specialized persona attributes, relying purely on default generation sampling variability.
    • Demographic Information (DI): Simulates user demographic diversity by conditioning the simulator on randomly sampled honorific titles (Mr. or Ms.) and one of the 500 most frequent surnames across five racial groups from census records (e.g., Pretend to be Ms. Guzman.).
    • Demographic Information with Pickiness Personality (DI + PP): Extends DI by conditioning the simulator with an explicit movie pickiness trait randomly sampled from three levels: not picky, moderately picky, and extremely picky (e.g., Pretend to be Ms. Guzman. You are extremely picky about movies.).
    • Interaction History (IH): Conditions the simulator on past user interactions sampled directly from human profiles. This includes previous movie mentions in a dialogue (ReDial), movie mentions with timestamps (Reddit), or a set of 10 prior movie titles with user review titles (IMDB).
  3. Knowl 3 — Item Mention Distribution Entropy

    equation

    To quantify the diversity of items mentioned by human users or simulated user populations in conversational contexts, Shannon entropy over the item vocabulary distribution is defined as:

    H(X)=−∑i=1np(xi)log⁡p(xi)H(X) = - \sum_{i=1}^n p(x_i) \log p(x_i)

    where nn is the total number of distinct items mentioned, and p(xi)p(x_i) is the empirical probability that item xix_i is mentioned within the generated or real conversational corpus (excluding any anchor items provided in the prompt).

  4. Knowl 4 — Embedding Cosine Diversity Metric for Request Granularity

    equation

    To evaluate the semantic diversity of recommendation requests generated across a population of simulated users compared to real users, the cosine diversity of embedding vectors is defined as:

    Diversity=1−1N∑i=1Ns⃗i⋅μ^∥s⃗i∥⋅∥μ^∥\text{Diversity} = 1 - \frac{1}{N} \sum_{i=1}^N \frac{\vec{s}_i \cdot \hat{\mu}}{\|\vec{s}_i\| \cdot \|\hat{\mu}\|}

    where NN is the total number of generated or real requests, s⃗i∈Rd\vec{s}_i \in \mathbb{R}^d is the dd-dimensional embedding (e.g., Word2Vec word embedding or SBERT sentence embedding) of the ii-th request, and μ^∈Rd\hat{\mu} \in \mathbb{R}^d is the dataset embedding centroid defined by:

    μ^=1N∑i=1Ns⃗i\hat{\mu} = \frac{1}{N} \sum_{i=1}^N \vec{s}_i

  5. Knowl 5 — Item Mention Diversity and the Effect of Interaction History (ItemsTalk Results)

    data/table

    In the ItemsTalk task, baseline LLM simulators conditioned only on demographic information (DI) exhibit severe popularity bias, producing significantly lower entropy over mentioned items compared to real human users. Conditioning simulators with user Interaction History (IH) substantially increases item diversity across all models (gpt-3.5-turbo, gpt-4, text-davinci-003) and benchmark datasets (IMDB, Reddit, ReDial).

    Generator IMDB Reddit ReDial
    Human 12.61 11.73 9.71
    Demographic information (DI)
    gpt-3.5 4.79 3.97 4.00
    gpt-4 5.29 4.78 4.18
    text-davinci 6.42 6.69 6.66
    Interaction history (IH)
    gpt-3.5 7.96 7.14 7.68
    gpt-4 8.59 9.50 9.03
    text-davinci 10.79 9.97 8.63

    The human baseline demonstrates high entropy (9.71 to 12.61). Prompting with demographic information results in entropy drops exceeding 50% in most cases. Prompting with interaction history restores item mention entropy to levels close to or exceeding human reference distributions in specific conversational sets (e.g., text-davinci-003 achieves 10.79 on IMDB and 9.97 on Reddit).

  6. Knowl 6 — Binary Preference Alignment via Pickiness Persona Conditioning (BinPref Results)

    data/table

    In the BinPref task, simulated user populations (100 runs per item) evaluate 200 frequent (≥5000\ge 5000 MovieLens ratings) and 200 infrequent (50≤ratings≤50050 \le \text{ratings} \le 500) movies. Pearson correlation coefficients (rr) measure the linear alignment between the proportion of positive 'Yes' answers and ground-truth human average ratings (scale 1–5). Without explicit personality traits, simulators display strong positive bias with low correlation to human ratings (text-davinci-003 with DI yielded undefined correlation for frequent items as it answered 'Yes' 100% of the time). Prompting with Pickiness Personality (DI + PP) substantially aligns simulator approval rates with human ratings.

    Generator Frequent items Infrequent items
    Demographic information (DI)
    gpt-3.5 0.18 0.12
    gpt-4 0.24 0.53
    text-davinci Undefined 0.29
    Demographic information + Pickiness (DI + PP)
    gpt-3.5 0.45 0.36
    gpt-4 0.75 0.76
    text-davinci 0.49 0.64

    All reported correlations have p<0.05p < 0.05. Higher item frequency does not automatically yield better preference alignment. Adding pickiness allows simulators (especially gpt-4, reaching r=0.75r = 0.75 and r=0.76r = 0.76) to penalize poorly rated items accurately.

  7. Knowl 7 — Aspect and Sentiment Distribution in Open-Ended Preferences (OpenPref Results)

    data/table

    The OpenPref task evaluates free-form movie feedback generated by LLMs against human IMDB reviews using PyABSA for aspect-based sentiment analysis. Evaluated metrics include the total number of distinct extracted aspects, the Shannon entropy of aspect distributions, and the Shannon entropy of sentiment distributions.

    Generator # aspects Aspect entropy Sentiment entropy
    Human 85 5.85 1.19
    Demographic information (DI)
    gpt-3.5 71 4.86 0.29
    gpt-4 97 5.57 1.11
    text-davinci 194 5.63 0.18
    Demographic information + Pickiness (DI + PP)
    gpt-3.5 101 5.20 1.09
    gpt-4 97 5.59 1.34
    text-davinci 232 5.47 0.48

    Simulators prompted with DI alone produce overly optimistic reviews, reflected in very low sentiment entropy (e.g., 0.18 for text-davinci and 0.29 for gpt-3.5). Adding pickiness personality (DI + PP) raises sentiment entropy toward human levels (1.09 to 1.34 vs. 1.19 human). While simulators generate more nominal aspects (up to 232 aspects for text-davinci), their aspect entropy is lower than humans (5.85), indicating that simulators repetitively evaluate a predictable set of explicit features (e.g., cast, plot) rather than contextual, nuanced impressions.

  8. Knowl 8 — Personalization Gap and Semantic Diversity in Synthetic Recommendation Requests (RecRequest Results)

    data/table

    In the RecRequest task, simulators generate recommendation requests conditioned on fixed seed movies and target character lengths from Reddit recommendation posts. Diversity is measured via word type-token ratio (Word), Word2Vec cosine embedding diversity (Word emb.), and SBERT sentence cosine embedding diversity (Sentence emb.).

    Generator Word (Type-Token) Word emb. Sentence emb.
    Human 0.65 0.427 0.391
    gpt-3.5 0.50 0.433 0.295
    gpt-4 0.61 0.436 0.300
    text-davinci 0.49 0.418 0.288

    Simulators exhibit a severe deficit in sentence-level request diversity, with gpt-4 achieving a sentence diversity of 0.300 compared to 0.391 for humans (a 23% reduction). Although word embedding diversity remains comparable to human levels (0.418–0.436 vs. 0.427), the type-token ratio is lower (0.49–0.61 vs. 0.65). LLMs repeatedly reuse stock adjectives (e.g., 'gripping', 'mind-bending', 'compelling') and produce generic, broad category requests rather than highly idiosyncratic, personal criteria (e.g., 'extreme loneliness or depression', 'rock climbing', 'growth mindset versus fixed mindset').

  9. Knowl 9 — Feedback Coherence and Explanation Effects in Simulator Responses (Feedback Results)

    data/table

    The Feedback task evaluates whether simulated users can reliably accept relevant recommendations and reject irrelevant negative recommendations across Reddit conversational pairs. In the accept/reject setting, feedback is coherent if the agent accepts positive recommendations and rejects negative recommendations. In the comparison setting, feedback is coherent if the simulator prefers the positive recommendation over the negative recommendation.

    Setting Generator Prop. coherent Prop. neither
    Items only gpt-3.5 0.8264 0.0087
    gpt-4 0.9096 0.0120
    text-davinci 0.8108 0
    Items + explain gpt-3.5 0.8039 0.0049
    gpt-4 0.9047 0
    text-davinci 0.6567 0

    Simulators are coherent in 65% to 91% of cases. Adding textual explanations to recommendations slightly decreases comparison coherence across all models (e.g., text-davinci drops from 0.8108 to 0.6567), because explanations provide persuasive rationales for negative/irrelevant recommendations, making them harder for simulators to distinguish. Qualitative error analysis reveals that incoherent simulator rejections often stem from failing to capture subtle request semantics (e.g., rejecting a movie featuring a loner character because the plot is not strictly about a loner character).

  10. Knowl 10 — Limitations of the Conversational Recommender Simulation Protocol

    limitation

    The five evaluation tasks define necessary but not sufficient conditions for synthetic agents to fully represent real user populations in conversational recommender systems. The protocol has several primary limitations:

    1. Task Scope: The protocol does not evaluate dynamic multi-turn interactions, user question-asking behaviors during recommendation, or simulator handling of novel/evolving items outside the LLM's pre-training corpus.
    2. Domain Specificity: The empirical validation is restricted entirely to movie recommendation domains (ReDial, Reddit, MovieLens, IMDB), which may not generalize directly to domains with different decision criteria such as e-commerce.
    3. Model & Hyperparameter Exploration: The baseline evaluations focus primarily on OpenAI models (gpt-3.5-turbo, gpt-4, text-davinci-003) using default sampling temperatures and basic prompting strategies, leaving open-source LLMs, temperature sensitivity, and complex agent architectures for future study.

Coverage note — None was omitted. All contributed tasks, mathematical formulations, prompting strategies, empirical results across Tables 1-8, and stated limitations are fully covered.

References

  1. 1.Gati V Aher, Rosa I Arriaga, and Adam Tauman Kalai. 2023. Using large language models to simulate multiple humans and replicate human subject studies. In ICML.
  2. 2.Ashton Anderson, Lucas Maystre, Ian Anderson, Rishabh Mehrotra, and Mounia Lalmas. 2020. Algorithmic effects on the diversity of consumption on spotify. In WWW.
  3. 3.Lisa P Argyle, Ethan C Busby, Nancy Fulda, Joshua R Gubler, Christopher Rytting, and David Wingate. 2023. Out of one, many: Using language models to simulate human samples. Political Analysis, 31(3):337–351.
  4. 4.Zheng Chen. 2023. Palr: Personalization aware llms for recommendation. arXiv preprint arXiv:2305.07622.
  5. 5.Konstantina Christakopoulou, Filip Radlinski, and Katja Hofmann. 2016. Towards conversational recommender systems. In KDD.
  6. 6.Wenqi Fan, Zihuai Zhao, Jiatong Li, Yunqing Liu, Xiaowei Mei, Yiqi Wang, Jiliang Tang, and Qing Li. 2023. Recommender systems in the era of large language models (llms). arXiv preprint arXiv:2307.02046.
  7. 7.Chen Gao, Xiaochong Lan, Zhihong Lu, Jinzhu Mao, Jinghua Piao, Huandong Wang, Depeng Jin, and Yong Li. 2023. S3: Social-network simulation system with large language model-empowered agents. arXiv preprint arXiv:2307.14984.
  8. 8.Chongming Gao, Wenqiang Lei, Xiangnan He, Maarten de Rijke, and Tat-Seng Chua. 2021. Advances and challenges in conversational recommender systems: A survey. AI Open.
  9. 9.Xiaoxiao Guo, Hui Wu, Yu Cheng, Steven Rennie, Gerald Tesauro, and Rogerio Feris. 2018. Dialog-based interactive image retrieval. NeurIPS.
  10. 10.Perttu Hämäläinen, Mikke Tavast, and Anton Kunnari. 2023. Evaluating large language models in generating synthetic hci research data: a case study. In CHI.
  11. 11.F Maxwell Harper and Joseph A Konstan. 2015. The movielens datasets: History and context. ACM TIIS.
  12. 12.Zhankui He, Zhouhang Xie, Rahul Jha, Harald Steck, Dawen Liang, Yesu Feng, Bodhisattwa Prasad Majumder, Nathan Kallus, and Julian McAuley. 2023. Large language models as zero-shot conversational recommenders. In CIKM.
  13. 13.Yupeng Hou, Junjie Zhang, Zihan Lin, Hongyu Lu, Ruobing Xie, Julian McAuley, and Wayne Xin Zhao. 2023. Large language models are zero-shot rankers for recommender systems. arXiv preprint arXiv:2305.08845.
  14. 14.Wang-Cheng Kang, Jianmo Ni, Nikhil Mehta, Maheswaran Sathiamoorthy, Lichan Hong, Ed Chi, and Derek Zhiyuan Cheng. 2023. Do llms understand user preferences? evaluating llms on user rating prediction. arXiv preprint arXiv:2305.06474.
  15. 15.Wenqiang Lei, Xiangnan He, Yisong Miao, Qingyun Wu, Richang Hong, Min-Yen Kan, and Tat-Seng Chua. 2020a. Estimation-action-reflection: Towards deep interaction between conversational and recommender systems. In WSDM.
  16. 16.Wenqiang Lei, Gangyi Zhang, Xiangnan He, Yisong Miao, Xiang Wang, Liang Chen, and Tat-Seng Chua. 2020b. Interactive path reasoning on graph for conversational recommendation. In KDD.
  17. 17.Jinming Li, Wentao Zhang, Tian Wang, Guanglei Xiong, Alan Lu, and Gerard Medioni. 2023. Gpt4rec: A generative framework for personalized recommendation and user interests interpretation. arXiv preprint arXiv:2304.03879.
  18. 18.Raymond Li, Samira Ebrahimi Kahou, Hannes Schulz, Vincent Michalski, Laurent Charlin, and Chris Pal. 2018. Towards deep conversational recommendations. In NeurIPS.
  19. 19.Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781.
  20. 20.Ida Momennejad, Hosein Hasanbeig, Felipe Vieira, Hiteshi Sharma, Robert Osazuwa Ness, Nebojsa Jojic, Hamid Palangi, and Jonathan Larson. 2023. Evaluating cognitive maps and planning in large language models with cogeval. arXiv preprint arXiv:2309.15129.
  21. 21.Seungwhan Moon, Pararth Shah, Anuj Kumar, and Rajen Subba. 2019. Opendialkg: Explainable conversational reasoning with attention-based walks over knowledge graphs. In ACL.
  22. 22.OpenAI. 2021. About openai.
  23. 23.Paul Owoicho, Ivan Sekulic, Mohammad Aliannejadi, Jeffrey Dalton, and Fabio Crestani. 2023. Exploiting simulated user feedback for conversational search: Ranking, rewriting, and beyond. In SIGIR.
  24. 24.Wenbo Pan, Qiguang Chen, Xiao Xu, Wanxiang Che, and Libo Qin. 2023. A preliminary evaluation of chatgpt for zero-shot dialogue understanding. arXiv preprint arXiv:2304.04256.
  25. 25.Joon Sung Park, Joseph C O’Brien, Carrie J Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein. 2023. Generative agents: Interactive simulacra of human behavior. In UIST.
  26. 26.Chen Qian, Xin Cong, Cheng Yang, Weize Chen, Yusheng Su, Juyuan Xu, Zhiyuan Liu, and Maosong Sun. 2023. Communicative agents for software development. arXiv preprint arXiv:2307.07924.
  27. 27.Nils Reimers and Iryna Gurevych. 2019. Sentence-bert: Sentence embeddings using siamese bert-networks. EMNLP.
  28. 28.Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto. 2023. Whose opinions do language models reflect? In ICML.
  29. 29.Yueming Sun and Yi Zhang. 2018. Conversational recommender system. In SIGIR.
  30. 30.Lei Wang, Jingsen Zhang, Xu Chen, Yankai Lin, Ruihua Song, Wayne Xin Zhao, and Ji-Rong Wen. 2023a. Recagent: A novel simulation paradigm for recommender systems. arXiv preprint arXiv:2306.02552.
  31. 31.Xiaolei Wang, Xinyu Tang, Wayne Xin Zhao, Jingyuan Wang, and Ji-Rong Wen. 2023b. Rethinking the evaluation for conversational recommendation in the era of large language models. In EMNLP.
  32. 32.Zhenduo Wang, Zhichao Xu, Qingyao Ai, and Vivek Srikumar. 2023c. An in-depth investigation of user response simulation for conversational search. arXiv preprint arXiv:2304.07944.
  33. 33.Yu Xia, Junda Wu, Tong Yu, Sungchul Kim, Ryan A Rossi, and Shuai Li. 2023. User-regulation deconfounded conversational recommender system with bandit feedback. In KDD.
  34. 34.Heng Yang, Chen Zhang, and Ke Li. 2023. Pyabsa: A modularized framework for reproducible aspect-based sentiment analysis. In CIKM.
  35. 35.Junjie Zhang, Yupeng Hou, Ruobing Xie, Wenqi Sun, Julian McAuley, Wayne Xin Zhao, Leyu Lin, and Ji-Rong Wen. 2023. Agentcf: Collaborative learning with autonomous language agents for recommender systems. arXiv preprint arXiv:2310.09233.
  36. 36.Shuo Zhang and Krisztian Balog. 2020. Evaluating conversational recommender systems via user simulation. In KDD.
  37. 37.Shuo Zhang, Mu-Chun Wang, and Krisztian Balog. 2022. Analyzing and simulating user utterance reformulation in conversational recommender systems. In SIGIR.
  38. 38.Weixiang Zhao, Yanyan Zhao, Xin Lu, Shilong Wang, Yanpeng Tong, and Bing Qin. 2023. Is chatgpt equipped with emotional dialogue capabilities? arXiv preprint arXiv:2304.09582.

Citation

MLA
Yoon, S.-. eun ., et al. “Evaluating Large Language Models as Generative User Simulators for Conversational Recommendation”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2024, pp. 1490–504, https://doi.org/10.18653/v1/2024.naacl-long.83.
APA
Yoon, S.-. eun ., He, Z., Echterhoff, J., & McAuley, J. (2024). Evaluating Large Language Models as Generative User Simulators for Conversational Recommendation. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 1490–1504. https://doi.org/10.18653/v1/2024.naacl-long.83
Chicago
Yoon, S.-. eun ., Z. He, J. Echterhoff, and J. McAuley. 2024. “Evaluating Large Language Models as Generative User Simulators for Conversational Recommendation”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 1490–1504. https://doi.org/10.18653/v1/2024.naacl-long.83.
Harvard
Yoon, S.-. eun . et al. (2024) “Evaluating Large Language Models as Generative User Simulators for Conversational Recommendation”, Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp. 1490–1504. Available at: https://doi.org/10.18653/v1/2024.naacl-long.83.
Vancouver
1. Yoon S-eun, He Z, Echterhoff J, McAuley J (2024) Evaluating Large Language Models as Generative User Simulators for Conversational Recommendation. In: Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp 1490–1504

BibTeX

@inproceedings{yoon-etal-2024-evaluating,
    title = "Evaluating Large Language Models as Generative User Simulators for Conversational Recommendation",
    author = "Yoon, Se-eun  and
      He, Zhankui  and
      Echterhoff, Jessica  and
      McAuley, Julian",
    editor = "Duh, Kevin  and
      Gomez, Helena  and
      Bethard, Steven",
    booktitle = "Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
    month = jun,
    year = "2024",
    address = "Mexico City, Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.naacl-long.83/",
    doi = "10.18653/v1/2024.naacl-long.83",
    pages = "1490--1504"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/