Quantifying the Persona Effect in LLM Simulations

Tiancheng HuNigel Collier

article2024ACL260 citations

Quantifies the limits and efficacy of persona prompting across subjective NLP tasks, establishing that demographic variables explain under ten percent of annotation variance yet enable large language models to recover most predictable human variation when strong correlations exist.

Listen

As organizations increasingly explore artificial intelligence to simulate human opinions, user perspectives, and content evaluations, questions remain regarding how reliably large language models can mirror diverse demographic groups. Deploying AI systems that assume synthetic personas carries significant operational risks, including misrepresenting target audiences and failing to capture true stakeholder consensus. The article evaluates how effectively persona prompting—the practice of prepending demographic, social, and behavioral traits to model instructions—enables language models to simulate human perspectives on subjective language tasks and survey questions.

To establish an objective baseline, the researchers first applied mixed-effect linear regression across 10 subjective language processing datasets (such as toxicity, offensiveness, and sentiment labeling) and a benchmark national election survey to measure how much human variation persona traits actually explain. They then conducted zero-shot simulation experiments using various models, including GPT-4, GPT-3.5, and 70-billion-parameter open-source models, testing predictions with and without persona descriptions. The evaluation encompassed 600 sampled instances per dataset, analyzing performance across varying levels of annotator disagreement and testing robustness against variations in prompt wording and attribute order.

Key findings show that persona variables explain very little of the variance in human subjective annotations—typically between 1.4% and 10.6%—whereas text-specific differences account for up to 70% of variation, and 25% to 70% remains entirely unexplained. Consequently, incorporating persona prompts into language models yields modest, though occasionally statistically significant, performance gains; for instance, in tasks where personas explained 9% of human variance, prompting yielded only an average 1% improvement. The technique proved most useful on borderline cases where human annotators broadly disagreed within a narrow margin (high entropy and low standard deviation), allowing models to make subtle calibration adjustments. In structured survey environments where persona traits strongly determine responses, top 70-billion-parameter models captured 81% of the predictable variance, but accuracy dropped to zero whenever the underlying explanatory power of persona traits fell below 10%.

These results demonstrate that persona prompting cannot reliably simulate authentic human perspectives in standard subjective natural language tasks because demographic variables alone lack sufficient explanatory power. For decision-makers, relying on automated persona simulations to replace human panels introduces substantial compliance, safety, and operational risks. Models often default to representing groups as uniform monoliths rather than capturing true individual heterogeneity.

Decision-makers should exercise strict caution and avoid using zero-shot persona prompting as a direct substitute for human participants in subjective evaluations, policy assessments, or market research. If the objective is simply to increase the general diversity of generated text, basic persona prompting may suffice; however, high-fidelity human behavioral simulation requires extensive empirical validation, strategic dataset design capturing deeper personal attitudes and values, and targeted model fine-tuning. Because the underlying research relied primarily on English-language, United States-centric datasets with group-level demographic traits, organizations should exercise additional caution when deploying these models in cross-cultural or multilingual environments without independent testing.

No sufficiently relevant recommendations were found.

Cover for Quantifying the Persona Effect in LLM Simulations

Abstract

Large language models (LLMs) have shown remarkable promise in simulating human language and behavior. This study investigates how integrating persona variables—demographic, social, and behavioral factors—impacts LLMs’ ability to simulate diverse perspectives. We find that persona variables account for <10% variance in annotating existing subjective NLP datasets. Nonetheless, incorporating persona variables via prompting in LLMs provides modest but statistically significant improvements. Persona prompting is most effective in samples where many annotators disagree, but their disagreements are relatively minor. Notably, we find a linear relationship in our setting: the stronger the correlation between persona variables and human annotations, the more accurate the LLM predictions are using persona prompting. In a zero-shot setting, a powerful 70b model with persona prompting captures 81% of the annotation variance achievable by linear regression trained on ground truth annotations. However, for most subjective NLP datasets, where persona variables have limited explanatory power, the benefits of persona prompting are limited.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 The Relationship between Persona Variables and Annotation Outcome
  • 2.2 Modeling Persona Variables and LLM for Simulation
  • 2.3 Persona Prompting and AI Alignment
  • 3 RQ1: How much variance in human annotation could persona variables explain?
  • 4 RQ2: Can incorporating persona variables via prompting improve LLMs' predictions?
  • 5 RQ3: For what types of samples is persona prompting most useful?
  • 6 RQ4: How effectively can LLMs simulate personas when the importance of persona variables varies?
  • 7 Conclusion and Recommendation
  • 8 Limitations
  • 9 Ethical Considerations
  • Acknowledgements
  • References
  • A Implementation Detail
  • B Supplementary Results for Section 4
  • B.1 Prompt Template
  • B.2 Persona Variables
  • B.3 Results from All Models
  • C Supplementary Results for Section 6
  • D Robustness Test

Knowls

  1. Knowl 1 — Persona variables explain little annotation variance in most NLP datasets

    data/table

    The study estimated how much variation in human judgments was associated with annotator persona variables across 10 subjective NLP datasets, while accounting for the text being judged. Persona variables explained a small share of annotation variance in these datasets: marginal R2R^2 ranged from 0.005 to 0.106. In contrast, the ANES presidential-vote survey outcome had marginal R2=0.719R^2=0.719. The conditional R2R^2 values for NLP tasks include both persona variables and text-specific variation; marginal R2R^2 reflects persona variables alone.

    Task and datasetText samples, NNMean annotators per sample, AAConditional R2R^2Marginal R2R^2
    Toxicity, annWithAttitudes6265.50.6110.045
    Offensiveness, POPQUORN1,5008.70.3190.029
    Politeness, POPQUORN3,7186.70.4540.014
    Toxicity, Kumar et al. (2021)106,0355.10.3490.106
    Sentiment, Diaz et al. (2018)14,0714.20.3290.036
    Social acceptability, Social-Chem-1019,7406.10.4320.097
    Social acceptability, NLPositionality29150.20.5130.005
    Toxicity, NLPositionality29929.60.4320.017
    Social bias, SBIC35,5043.20.7580.031
    Irony, EPIC2,9944.70.2890.091
    Presidential vote, ANES 2012—2,728 respondents—0.719

    The text-specific random effect accounts for substantial additional variation in the NLP tasks, reaching roughly 70% in some datasets. ANES has no text-specific effect because respondents answer the same survey question. Thus, the strong ANES result does not characterize most of the subjective NLP annotation settings studied.

  2. Knowl 2 — LLM simulation accuracy rises with persona predictability but remains below a regression benchmark

    empirical result

    In a U.S. survey case study, the authors compared each question’s ground-truth predictability from persona variables (target R2R^2) with the R2R^2 obtained from persona-prompted LLM predictions of respondents’ answers (predicted R2R^2). Across the tested question/profile settings, predicted R2R^2 increased with target R2R^2. The strongest reported relationship was for Tulu-2-dpo-70b: its fitted line was y=0.81x−0.09y=0.81x-0.09, with fitted R2=0.74R^2=0.74, where xx is target R2R^2 and yy is predicted R2R^2. The authors characterize this as capturing 81% of the target variance-explanation level.

    No tested model exceeded the perfect-prediction line y=xy=x, so persona-prompted simulation remained below the linear-regression benchmark fitted to human responses. The largest models, especially preference-tuned 70b models, performed best; smaller 7b and 13b models generally performed worse. When target R2R^2 was below 0.1, predicted R2R^2 was often near zero. The evidence therefore supports a positive association between persona predictability and simulation accuracy in this setting, not reliable simulation when the available persona variables weakly predict human answers.

  3. Knowl 3 — Persona prompting gives modest, inconsistent gains in individual annotation prediction

    empirical result

    The study compared zero-shot predictions with and without persona prompts on four subjective NLP datasets: annWithAttitudes, Kumar et al. (2021) toxicity, EPIC irony, and POPQUORN politeness. The reported average changes aggregate six strong models (GPT-4, GPT-3.5, Llama-2-70b, Llama-2-70b-chat, Tulu-2-70b, and Tulu-2-dpo-70b). Positive changes are improvements for R2R^2, Cohen’s kappa (κ\kappa), and macro F1; a negative change is an improvement for mean absolute error (MAE). Asterisks indicate statistically significant aggregate changes under 1,000 bootstrap replications.

    DatasetChange in R2R^2Change in κ\kappaChange in macro F1 or MAE
    annWithAttitudes+0.06*+0.02MAE −0.14*
    Kumar et al. (2021)−0.01−0.01MAE −0.22*
    EPIC+0.01*+0.04*macro F1 +0.02*
    POPQUORN politeness−0.02+0.00MAE −0.05*

    At least one metric improved significantly in aggregate for each dataset, but gains were generally small and did not occur uniformly across models or metrics. GPT-4 was the strongest overall model in these comparisons. Llama-2-70b was notably sensitive to persona prompting: on annWithAttitudes its R2R^2 rose from 0.17 to 0.40, despite a sampled-data target marginal R2R^2 of only 0.03. Persona information can therefore change predictions substantially for some models without ensuring consistent gains.

  4. Knowl 4 — The largest persona-prompting gains occur with frequent but small disagreements

    empirical result

    Across toxicity ratings from Kumar et al. (2021) and politeness ratings from POPQUORN, the greatest mean MAE improvement from persona prompting occurred for samples with high annotation entropy but low annotation standard deviation. This category represents frequent disagreement among annotators whose ratings nevertheless stay within a relatively narrow range. The authors interpret the pattern as evidence that persona prompts can help refine predictions among nearby response options, rather than reliably producing large shifts across the rating scale.

    For both datasets, bootstrapping with 1,000 replications found significant MAE improvements in all four entropy-by-standard-deviation categories. A one-way ANOVA also found differences among categories. Tukey comparisons showed that high-entropy/low-standard-deviation samples significantly outperformed the other three categories for Kumar et al. (2021). In POPQUORN, this category had the greatest improvement and was not significantly worse than any other category. Low-entropy/high-standard-deviation samples consistently had the least significant MAE improvement. These results indicate that disagreement structure, not simply the presence of disagreement, matters for persona prompting.

  5. Knowl 5 — Persona explanatory power is estimated with a text-controlled mixed-effects regression

    model/method

    The study operationalizes persona variables broadly: they include demographics and social attributes as well as attitudes, behaviors, lived experiences, and values that could describe an annotator. For each NLP dataset, the response is an individual human annotation; persona variables enter as fixed effects, while a random intercept for each text sample controls for text-specific differences. The model is specified as annotation ~ persona variables + (1 | text_id). For the ANES survey, which has no varying text sample, the text random effect is unnecessary.

    The authors also tested adding an annotator-level random effect and report that fixed-effect estimates were very similar, so the main analysis retained only the text random effect. Marginal R2R^2 is used to quantify variance associated with the persona fixed effects, and conditional R2R^2 includes both fixed effects and the text random effect. This is an interpretable linear baseline: it does not model interactions among persona variables and is not intended to measure every possible nonlinear or unobserved source of individual variation.

  6. Knowl 6 — Zero-shot evaluation compares persona and no-persona prompts on individual labels

    experimental setup

    For each of four datasets—annWithAttitudes, Kumar et al. (2021) toxicity, EPIC irony, and POPQUORN politeness—the researchers sampled 600 instances and prompted each model twice: once with annotator persona information and once without it. Prompts preserved the original persona descriptions as closely as possible, described the task and answer choices in a multiple-choice format, and used the model’s next-token response as its prediction. The evaluation targets were individual annotations, not aggregated labels.

    The compared systems included GPT-4, GPT-3.5, Llama-2, Llama-2-chat, Tulu-2, and Tulu-2-dpo; the main comparison reported the 70b versions of the open models alongside the API models. Performance was measured using R2R^2, Cohen’s kappa, and MAE, with macro F1 used for binary EPIC classification. Bootstrap comparisons used 1,000 replications to test whether including persona variables changed performance.

  7. Knowl 7 — Disagreement-stratified evaluation balances annotation patterns and rating levels

    experimental setup

    To test which annotation patterns benefit from persona prompts, the study used the Kumar et al. (2021) toxicity dataset and POPQUORN politeness. For each dataset, samples were split into four groups using median thresholds for annotation entropy and standard deviation: low entropy/low standard deviation, low entropy/high standard deviation, high entropy/low standard deviation, and high entropy/high standard deviation. Entropy captures how dispersed annotator choices are across categories, while standard deviation captures the magnitude of their numerical spread.

    Within each group, samples were further divided into four bins by mean annotation value, and 150 samples were randomly selected from each bin, producing 600 samples per group. This stratification reduces extreme imbalances in rating levels across groups. Llama-2-70b, Llama-2-70b-chat, Tulu-2-70b, and Tulu-2-dpo-70b were run with and without persona prompts, and the comparison used mean MAE improvement.

  8. Knowl 8 — The ANES case study varies question predictability while holding survey context stable

    experimental setup

    The survey case study used ANES 2012, treating each carefully designed survey question as an outcome to predict for individual respondents. After filtering respondents with missing answers and randomly downsampling, the study used 600 respondents and 21 questions. Each question was paired with two persona-prompt configurations, yielding 42 question/profile settings. The selected questions varied in how well persona variables predicted human responses; the reported target R2R^2 values for the richer prompt configuration ranged from 0.22 to 0.66.

    For each setting, the LLM received a respondent’s persona information and predicted that respondent’s answer. The resulting predicted R2R^2 was compared with target R2R^2, which represents the explainability of human responses by persona variables. This design reduces the influence of variable social-media text and allows the study to examine how simulation performance changes as persona variables become more predictive.

  9. Knowl 9 — Weak persona predictors limit simulation fidelity and motivate targeted data collection

    limitation

    The study’s results imply that zero-shot persona prompting is unlikely to simulate annotation perspectives reliably in tasks where available persona variables explain little human-response variance. The authors therefore recommend validating simulation fidelity, and potentially fine-tuning models, rather than treating unvalidated zero-shot outputs as dependable substitutes for human responses. They also recommend collecting persona variables for a defined purpose and, when behavioral prediction is the goal, including more targeted information about attitudes, beliefs, and behaviors rather than relying only on demographics.

    The evidence is limited by the available datasets, which are predominantly U.S.-based and English-language, and by the relatively coarse or incomplete persona information they contain. The authors propose—but do not establish—that group-level attributes may fail to represent individual identities and that LLMs may flatten within-group diversity. They also caution that unmeasured personal and contextual factors will likely leave some annotation variation unexplained, and that findings should not be assumed to transfer to other cultural or language settings without separate evaluation.

  10. Knowl 10 — Prompt order and paraphrasing produce only modest robustness variation

    empirical result

    Robustness checks on Kumar et al. (2021) toxicity and ANES changed the order of persona variables across five configurations or replaced the persona wording with five semantic paraphrases. The checks used Llama-2-70b, Llama-2-70b-chat, Tulu-2-70b, and Tulu-2-dpo-70b, and assessed prediction performance with R2R^2, Cohen’s kappa, MAE, and F1 where applicable. The authors report that results varied only modestly across ordering and wording choices. This supports some prompt-format robustness in these two settings, but does not establish robustness to other prompt designs or tasks.

Coverage note — The supplementary leave-one-variable-out feature-importance scores are omitted because they are secondary to the paper’s main findings about overall persona explanatory power and simulation performance.

References

  1. 1.Gati V. Aher, Rosa I. Arriaga, and Adam Tauman Kalai. 2023. Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies. In International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA, volume 202 of Proceedings of Machine Learning Research, pages 337–371. PMLR.
  2. 2.Kabir Ahuja, Harshita Diddee, Rishav Hada, Millicent Ochieng, Krithika Ramesh, Prachi Jain, Akshay Nambi, Tanuja Ganu, Sameer Segal, Mohamed Ahmed, Kalika Bali, and Sunayana Sitaram. 2023. MEGA: Multilingual evaluation of generative AI. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 4232–4267, Singapore. Association for Computational Linguistics.
  3. 3.ANES. The american national election studies 2012 time series study.
  4. 4.Lisa P. Argyle, Ethan C. Busby, Nancy Fulda, Joshua R. Gubler, Christopher Rytting, and David Wingate. 2023. Out of One, Many: Using Language Models to Simulate Human Samples. Political Analysis, 31(3):337–351. Publisher: Cambridge University Press.
  5. 5.April H. Bailey, Adina Williams, and Andrei Cimpian. 2022. Based on billions of words on the internet, PEOPLE = MEN. Science Advances, 8(13):eabm2463.
  6. 6.Larry M. Bartels. 2002. Beyond the running tally: Partisan bias in political perceptions. Political Behavior, 24(2):117–150.
  7. 7.Tilman Beck, Hendrik Schuff, Anne Lauscher, and Iryna Gurevych. 2024. Sensitivity, performance, robustness: Deconstructing the effect of sociodemographic prompting. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2589–2615, St. Julian’s, Malta. Association for Computational Linguistics.
  8. 8.Laura Biester, Vanita Sharma, Ashkan Kazemi, Naihao Deng, Steven R. Wilson, and Rada Mihalcea. 2022. Analyzing the effects of annotator gender across NLP tasks. In Proceedings of the 1st Workshop on Perspectivist Approaches to NLPerspectives@LREC 2022, Marseille, France, 20th June 2022, pages 10–19. European Language Resources Association.
  9. 9.James Bisbee, Joshua D. Clinton, Cassy Dorff, Brenton Kenkel, and Jennifer M. Larson. 2024. Synthetic replacements for human survey data? the perils of large language models. Political Analysis, page 1–16.
  10. 10.Lawrence Bobo and Frederick C Licari. 1989. Education and political tolerance: Testing the effects of cognitive sophistication and target group affect. Public Opinion Quarterly, 53(3):285–308.
  11. 11.Federico Cabitza, Andrea Campagner, and Valerio Basile. 2023. Toward a Perspectivist Turn in Ground Truthing for Predictive Computing. Proceedings of the AAAI Conference on Artificial Intelligence, 37(6):6860–6868. Number: 6.
  12. 12.Myra Cheng, Tiziano Piccardi, and Diyi Yang. 2023. CoMPosT: Characterizing and evaluating caricature in LLM simulations. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 10853–10875, Singapore. Association for Computational Linguistics.
  13. 13.Jacob Cohen. 1960. A coefficient of agreement for nominal scales. Educational and Psychological Measurement, 20(1):37–46.
  14. 14.Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2023. Qlora: Efficient finetuning of quantized llms. arXiv preprint arXiv:2305.14314.
  15. 15.Mark Diaz, Isaac Johnson, Amanda Lazar, Anne Marie Piper, and Darren Gergle. 2018. Addressing Age-Related Bias in Sentiment Analysis. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems, CHI ’18, pages 1–14, New York, NY, USA. Association for Computing Machinery.
  16. 16.Danica Dillion, Niket Tandon, Yuling Gu, and Kurt Gray. 2023. Can ai language models replace human participants? Trends in Cognitive Sciences, 27(7):597–600. Epub 2023 May 10.
  17. 17.Yi Ding, Jacob You, Tonja-Katrin Machulla, Jennifer Jacobs, Pradeep Sen, and Tobias Höllerer. 2022. Impact of Annotator Demographics on Sentiment Dataset Labeling. Proceedings of the ACM on Human-Computer Interaction, 6(CSCW2):519:1–519:22.
  18. 18.Jesse Dodge, Maarten Sap, Ana Marasovic, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner. 2021. Documenting large webtext corpora: A case study on the colossal clean crawled corpus. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 1286–1305, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  19. 19.Esin Durmus, Karina Nyugen, Thomas I. Liao, Nicholas Schiefer, Amanda Askell, Anton Bakhtin, Carol Chen, Zac Hatfield-Dodds, Danny Hernandez, Nicholas Joseph, Liane Lovitt, Sam McCandlish, Orowa Sikder, Alex Tamkin, Janel Thamkul, Jared Kaplan, Jack Clark, and Deep Ganguli. 2023. Towards Measuring the Representation of Subjective Global Opinions in Language Models. ArXiv:2306.16388 [cs].
  20. 20.Bradley Efron. 1992. Bootstrap Methods: Another Look at the Jackknife, pages 569–593. Springer New York, New York, NY.
  21. 21.Eve Fleisig, Rediet Abebe, and Dan Klein. 2023. When the majority is wrong: Modeling annotator disagreement for subjective tasks. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 6715–6726, Singapore. Association for Computational Linguistics.
  22. 22.Maxwell Forbes, Jena D. Hwang, Vered Shwartz, Maarten Sap, and Yejin Choi. 2020. Social chemistry 101: Learning to reason about social and moral norms. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 653–670, Online. Association for Computational Linguistics.
  23. 23.Simona Frenda, Alessandro Pedrani, Valerio Basile, Soda Marem Lo, Alessandra Teresa Cignarella, Raffaella Panizzon, Cristina Marco, Bianca Scarlini, Viviana Patti, Cristina Bosco, and Davide Bernardi. 2023. EPIC: Multi-perspective annotation of a corpus of irony. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 13844–13857, Toronto, Canada. Association for Computational Linguistics.
  24. 24.Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. 2020. The Pile: An 800GB Dataset of Diverse Text for Language Modeling. ArXiv:2101.00027 [cs].
  25. 25.Mitchell L. Gordon, Michelle S. Lam, Joon Sung Park, Kayur Patel, Jeff Hancock, Tatsunori Hashimoto, and Michael S. Bernstein. 2022. Jury Learning: Integrating Dissenting Voices into Machine Learning Models. In Conference on Human Factors in Computing Systems - Proceedings. Association for Computing Machinery. ArXiv: 2202.02950.
  26. 26.Jesse Graham and Jonathan Haidt. 2010. Beyond beliefs: Religions bind individuals into moral communities. Personality and social psychology review, 14(1):140–150.
  27. 27.Igor Grossmann, Matthew Feinberg, Dawn C. Parker, Nicholas A. Christakis, Philip E. Tetlock, and William A. Cunningham. 2023. Ai and the transformation of social science research. Science, 380(6650):1108–1109.
  28. 28.Perttu Hämäläinen, Mikke Tavast, and Anton Kunnari. 2023. Evaluating large language models in generating synthetic hci research data: a case study. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, CHI ’23, New York, NY, USA. Association for Computing Machinery.
  29. 29.Danula Hettiachchi, Indigo Holcombe-James, Stephanie Livingstone, Anjalee de Silva, Matthew Lease, Flora D. Salim, and Mark Sanderson. 2023. How crowd worker factors influence subjective annotations: A study of tagging misogynistic hate speech in tweets. Proceedings of the AAAI Conference on Human Computation and Crowdsourcing, 11(1):38–50.
  30. 30.John J Horton. 2023. Large language models as simulated economic agents: What can we learn from homo silicus? Technical report, National Bureau of Economic Research.
  31. 31.Tiancheng Hu, Yara Kyrychenko, Steve Rathje, Nigel Collier, Sander van der Linden, and Jon Roozenbeek. 2023. Generative language models exhibit social identity biases. arXiv preprint arXiv:2310.15819.
  32. 32.Olivia Huang, Eve Fleisig, and Dan Klein. 2023. Incorporating worker perspectives into MTurk annotation practices for NLP. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 1010–1028, Singapore. Association for Computational Linguistics.
  33. 33.Hamish Ivison, Yizhong Wang, Valentina Pyatkin, Nathan Lambert, Matthew Peters, Pradeep Dasigi, Joel Jang, David Wadden, Noah A Smith, Iz Beltagy, et al. 2023. Camels in a changing climate: Enhancing lm adaptation with tulu 2. arXiv preprint arXiv:2311.10702.
  34. 34.Junsol Kim and Byungkyu Lee. 2023. AI-Augmented Surveys: Leveraging Large Language Models for Opinion Prediction in Nationally Representative Surveys. ArXiv:2305.09620 [cs].
  35. 35.Grgur Kovac, Masataka Sawayama, Rémy Portelas, Cédric Colas, Peter Ford Dominey, and Pierre-Yves Oudeyer. 2023. Large Language Models as Superpositions of Cultural Perspectives. ArXiv:2307.07870 [cs].
  36. 36.Deepak Kumar, Patrick Gage Kelley, Sunny Consolvo, Joshua Mason, Elie Bursztein, Zakir Durumeric, Kurt Thomas, and Michael Bailey. 2021. Designing toxic content classification for a diversity of perspectives. In Proceedings of the Seventeenth USENIX Conference on Usable Privacy and Security, SOUPS’21, USA. USENIX Association.
  37. 37.Nayeon Lee, Chani Jung, Junho Myung, Jiho Jin, Juho Kim, and Alice Oh. 2023. Crehate: Cross-cultural re-annotation of english hate speech dataset. arXiv preprint arXiv:2308.16705.
  38. 38.Yinhong Liu, Yimai Fang, David Vandyke, and Nigel Collier. 2024. Toad: Task-oriented automatic dialogs with diverse response styles. arXiv preprint arXiv:2402.10137.
  39. 39.Daniel Lüdecke, Mattan S. Ben-Shachar, Indrajeet Patil, Philip Waggoner, and Dominique Makowski. 2021. performance: An R package for assessment, comparison and testing of statistical models. Journal of Open Source Software, 6(60):3139.
  40. 40.Nicola Marsden and Monika Pröbster. 2019. Personas and identity: Looking at multiple identities to inform the construction of personas. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems, CHI ’19, page 1–14, New York, NY, USA. Association for Computing Machinery.
  41. 41.Aida Mostafazadeh Davani, Mark Díaz, and Vinodkumar Prabhakaran. 2022. Dealing with disagreements: Looking beyond the majority vote in subjective annotations. Transactions of the Association for Computational Linguistics, 10:92–110.
  42. 42.Shinichi Nakagawa and Holger Schielzeth. 2013. A general and simple method for obtaining r2 from generalized linear mixed-effects models. Methods in Ecology and Evolution, 4(2):133–142.
  43. 43.OpenAI. 2023a. Gpt-4 technical report. arXiv preprint arXiv:2303.08774.
  44. 44.OpenAI. 2023b. Introducing ChatGPT.
  45. 45.Matthias Orlikowski, Paul Röttger, Philipp Cimiano, and Dirk Hovy. 2023. The ecological fallacy in annotation: Modeling human label variation goes beyond sociodemographics. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 1017–1029, Toronto, Canada. Association for Computational Linguistics.
  46. 46.Cecilia Ovesdotter Alm. 2011. Subjective natural language problems: Motivations, applications, characterizations, and implications. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, pages 107–112, Portland, Oregon, USA. Association for Computational Linguistics.
  47. 47.Joon Sung Park, Joseph C. O’Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. 2023. Generative Agents: Interactive Simulacra of Human Behavior. In In the 36th Annual ACM Symposium on User Interface Software and Technology (UIST ’23), UIST ’23, New York, NY, USA. Association for Computing Machinery. Event-place: San Francisco, CA, USA.
  48. 48.Peter S Park, Philipp Schoenegger, and Chongyang Zhu. 2024. Diminished diversity-of-thought in a standard large language model. Behavior Research Methods, pages 1–17.
  49. 49.Jiaxin Pei and David Jurgens. 2023. When do annotator demographics matter? measuring the influence of annotator demographics with the POPQUORN dataset. In Proceedings of the 17th Linguistic Annotation Workshop (LAW-XVII), pages 252–265, Toronto, Canada. Association for Computational Linguistics.
  50. 50.Pew Research Center. 2014. Political polarization in the american public. Pew Research Center. Accessed: Dec 19, 2023.
  51. 51.Barbara Plank. 2022. The “problem” of human label variation: On ground truth in data, modeling and evaluation. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 10671–10682, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  52. 52.Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto. 2023. Whose opinions do language models reflect? In Proceedings of the 40th International Conference on Machine Learning, ICML’23. JMLR.org.
  53. 53.Sebastin Santy, Jenny Liang, Ronan Le Bras, Katharina Reinecke, and Maarten Sap. 2023. NLPositionality: Characterizing design biases of datasets and models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 9080–9102, Toronto, Canada. Association for Computational Linguistics.
  54. 54.Maarten Sap, Saadia Gabriel, Lianhui Qin, Dan Jurafsky, Noah A. Smith, and Yejin Choi. 2020. Social bias frames: Reasoning about social and power implications of language. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5477–5490, Online. Association for Computational Linguistics.
  55. 55.Maarten Sap, Swabha Swayamdipta, Laura Vianna, Xuhui Zhou, Yejin Choi, and Noah A. Smith. 2022. Annotators with attitudes: How annotator beliefs and identities bias toxic language detection. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 5884–5906, Seattle, United States. Association for Computational Linguistics.
  56. 56.Amir Taubenfeld, Yaniv Dover, Roi Reichart, and Ariel Goldstein. 2024. Systematic biases in llm simulations of debates. arXiv preprint arXiv:2402.04049.
  57. 57.Petter Törnberg, Diliara Valeeva, Justus Uitermark, and Christopher Bail. 2023. Simulating social media using large language models to evaluate alternative news feed algorithms. arXiv preprint arXiv:2310.05984.
  58. 58.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023. Llama 2: Open Foundation and Fine-Tuned Chat Models. ArXiv:2307.09288 [cs].
  59. 59.Angelina Wang, Jamie Morgenstern, and John P Dickerson. 2024. Large language models cannot replace human participants because they cannot portray identity groups. arXiv preprint arXiv:2402.01908.
  60. 60.Peter West, Ximing Lu, Nouha Dziri, Faeze Brahman, Linjie Li, Jena D Hwang, Liwei Jiang, Jillian Fisher, Abhilasha Ravichander, Khyathi Chandu, et al. 2023. The generative ai paradox:" what it can create, it may not understand". arXiv preprint arXiv:2311.00059.
  61. 61.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online. Association for Computational Linguistics.

Citation

MLA
Hu, T., and N. Collier. “Quantifying the Persona Effect in LLM Simulations”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 10289–307, https://doi.org/10.18653/v1/2024.acl-long.554.
APA
Hu, T., & Collier, N. (2024). Quantifying the Persona Effect in LLM Simulations. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 10289–10307. https://doi.org/10.18653/v1/2024.acl-long.554
Chicago
Hu, T., and N. Collier. 2024. “Quantifying the Persona Effect in LLM Simulations”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 10289–307. https://doi.org/10.18653/v1/2024.acl-long.554.
Harvard
Hu, T. and Collier, N. (2024) “Quantifying the Persona Effect in LLM Simulations”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 10289–10307. Available at: https://doi.org/10.18653/v1/2024.acl-long.554.
Vancouver
1. Hu T, Collier N (2024) Quantifying the Persona Effect in LLM Simulations. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 10289–10307

BibTeX

@inproceedings{hu-collier-2024-quantifying,
    title = "Quantifying the Persona Effect in {LLM} Simulations",
    author = "Hu, Tiancheng  and
      Collier, Nigel",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.554/",
    doi = "10.18653/v1/2024.acl-long.554",
    pages = "10289--10307"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/