Evaluating Large Language Models as Generative User Simulators for Conversational Recommendation
Se-eun YoonZhankui HeJessica Maria EchterhoffJulian J. McAuley
Establishes the first standardized evaluation protocol comprising five core conversational tasks to measure how accurately large language models simulate diverse human behavior in conversational recommender systems and provides practical strategies to minimize behavioral discrepancies.
Conversational recommendation systems require extensive evaluation to ensure they effectively understand user needs and provide helpful suggestions. Testing these systems directly with human participants is expensive, time-consuming, and carries operational risks, while conventional offline benchmarks fail to capture interactive, multi-turn conversations. To resolve this dilemma, researchers and system developers are turning to synthetic users powered by large language models as automated stand-ins. However, prior to this study, there was no standardized framework to measure how accurately language model simulations replicate genuine human preferences and conversational behaviors across a broader population.
The main objective of the article is to establish and demonstrate the first systematic evaluation protocol for measuring the degree to which large language models can realistically emulate human behavior in conversational recommendation. To achieve this, the authors designed a zero-shot framework composed of five core tasks: selecting items to discuss, expressing binary preferences, providing open-ended opinions, formulating recommendation requests, and delivering feedback on suggested items. The authors evaluated several baseline language models—including GPT-3.5, GPT-4, and text-davinci-003—using real-world movie datasets gathered from ReDial, Reddit, MovieLens, and IMDB containing tens of thousands of conversations and reviews.
The investigation revealed critical distortions in how standard language models simulate human behavior. First, baseline models show a heavy popularity bias, generating significantly less diverse item mentions than humans; however, conditioning models on a user's interaction history substantially restored item diversity. Second, standard models exhibited an overly optimistic bias and correlated poorly with human preferences across movie rating tiers, with correlations near zero or undefined. Incorporating explicit "pickiness" personality traits into prompts sharply improved preference alignment, raising correlation coefficients up to 0.75–0.76 with GPT-4. Third, synthetic recommendation requests lacked individual personalization; the most diverse model produced requests that were roughly 23% less diverse than human requests, frequently recycling generic buzzwords like "gripping" and "mind-bending." Finally, while simulators provided coherent feedback on recommendations between 65% and 91% of the time, they occasionally rejected valid recommendations due to a failure to grasp subtle contextual nuances.
These findings demonstrate that off-the-shelf language models cannot be assumed to represent human populations out of the box. Relying on uncalibrated synthetic users creates significant risks of building conversational recommendation systems tuned to artificial, overly optimistic, and generic behaviors. However, the results prove that targeted prompting strategies—specifically integrating past interaction histories and explicit personality traits such as pickiness—alongside capable base models can effectively bridge the gap between artificial simulators and human behavior.
Decision-makers and engineering teams should adopt structured multi-task benchmarking protocols before using synthetic agents to evaluate user-facing conversational systems. Organizations should implement enriched persona prompts that specify interaction histories and varied preference thresholds rather than relying on default model settings. Furthermore, synthetic testing should be positioned as an efficient pre-deployment screening tool rather than a total replacement for live user testing. Future work should expand this evaluation framework beyond movies into broader commercial domains, test open-source foundation models, and develop more sophisticated methods to help synthetic users interpret subtle linguistic nuances in user requests.
- Paper: Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies, Gati V. Aher et al. (2023). This work establishes foundational methodologies for prompting LLMs with demographic personas to simulate diverse human populations in behavioral studies.
- Paper: Out of One, Many: Using Language Models to Simulate Human Samples, Lisa P. Argyle et al. (2022). This paper introduces silicon sampling to simulate nuanced human sub-population attitudes, providing the core conceptual basis for user simulation with LLMs.
- Paper: Whose Opinions Do Language Models Reflect?, Shibani Santurkar et al. (2023). This study analyzes how LLM opinions systematically diverge from human population distributions, motivating the need to evaluate and calibrate preference biases in simulators.
- Paper: Personalizing Dialogue Agents: I have a dog, do you have pets too?, Saizheng Zhang et al. (2018). This foundational paper shows that conditioning dialogue models on explicit textual persona profiles improves conversational consistency and engagement.
- Paper: How NOT To Evaluate Your Dialogue System: An Empirical Study of Unsupervised Evaluation Metrics for Dialogue Response Generation, Chia-Wei Liu et al. (2016). This benchmark paper demonstrates the severe limitations of standard automated metrics in conversational dialogue, justifying the turn toward interactive simulation-based evaluations.
- Paper: Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, Lianmin Zheng et al. (2023). This seminal benchmark validates using LLMs to approximate human conversational judgments and identifies key model biases in multi-turn interactive evaluation.
- Paper: Is this the real life? Is this just fantasy? The Misleading Success of Simulating Social Interactions With LLMs, Xuhui Zhou et al. (2024). This study critically investigates the validity of LLM-simulated social interactions under realistic information asymmetry, directly complementing conversational user simulation findings.
- Paper: Evaluating Very Long-Term Conversational Memory of LLM Agents, Adyasha Maharana et al. (2024). This work extends multi-turn conversational agent evaluation by benchmarking very long-term memory, temporal reasoning, and causal consistency across extended interaction histories.
- Paper: Self-Alignment of Large Language Models via Monopolylogue-based Social Scene Simulation, Xianghe Pang et al. (2024). This paper leverages multi-agent and social scene simulations to self-align language models, applying generative simulation techniques toward autonomous behavior refinement.
- Paper: Can Large Language Models be Good Emotional Supporter? Mitigating Preference Bias on Emotional Support Conversation, Dongjin Kang et al. (2024). This research analyzes and mitigates strategy preference biases in multi-stage conversational interactions, extending the study of behavioral alignment in conversational AI.
