CoMPosT: Characterizing and Evaluating Caricature in LLM Simulations

Myra ChengTiziano PiccardiDiyi Yang

article2023EMNLP149 citations

Presents a four-dimension framework and quantitative metric to evaluate how demographic simulations in large language models reduce complex human personas to exaggerated, stereotypical caricatures.

Listen

Researchers and organizations increasingly use large language models (LLMs) to simulate human behavior across social science experiments, public opinion polling, and product testing. Despite rapid adoption, there are no established methodologies to evaluate the quality of open-ended simulated responses. This gap creates significant risk, as automated simulations can easily produce flattened, exaggerated caricatures of demographic groups rather than meaningful, topical dialogue.

The article introduces CoMPosT, a structured framework designed to define LLM simulations and systematically measure their susceptibility to caricature. The primary objective is to evaluate how simulated personas and discussion topics interact to generate distorted, one-dimensional depictions of specific human populations.

The CoMPosT framework categorizes any simulation along four core dimensions: Context, Model, Persona, and Topic. To detect caricature, the article operationalizes two sequential criteria: individuation, which tests whether a simulated response can be distinguished from a neutral baseline using a binary classifier, and exaggeration, which measures whether the text disproportionately amplifies identity-related keywords over topic-relevant content using semantic axes. The authors evaluated open-ended generations from GPT-4 across multiple realistic scenarios, including online discussion forums, structured interviews, and social media posting. The study analyzed 15 distinct personas spanning age, political ideology, race and ethnicity, and gender across 60 discussion topics, sampling 100 responses per configuration.

The findings reveal that GPT-4 is highly susceptible to producing caricatures under specific conditions. First, simulations of political groups (such as conservatives) and marginalized demographics (including nonbinary, Black, Hispanic, and Middle-Eastern personas) exhibit the highest rates of exaggeration. Second, topic specificity is inversely related to caricature: general, uncontroversial subjects (such as health, relationships, or general technology) generate substantially higher levels of caricature than highly specific or contentious prompts. In these general scenarios, simulated personas frequently abandon the core subject to output generic identity slogans and activism narratives. Third, binary gender groups (men and women) exhibited the lowest caricature scores, reflecting the model's tendency to rely on implicit default personas, even though qualitative stereotypes were still present in baseline prompts.

These results demonstrate that unvetted LLM simulations pose operational and ethical risks. Relying on synthetic personas for market research, policy development, or behavioral modeling can mislead decision-makers through false homogeneity and outdated stereotypes. Organizations risk basing real-world strategies on artificial caricatures that fail to represent the actual diversity and nuanced opinions of target demographics.

To mitigate these risks, practitioners should avoid using broad, generic topics when prompting LLMs and instead provide fine-grained, highly contextualized prompts. Organizations should systematically evaluate synthetic personas using differentiation and exaggeration metrics before deploying them in production or research. Furthermore, simulation workflows should incorporate multifaceted persona descriptions and document researcher positionality to prevent outgroup bias from skewing simulation designs.

These conclusions should be interpreted within certain boundaries. The evaluation metrics detect specific forms of semantic exaggeration and differentiation, meaning low caricature scores do not guarantee the total absence of subtle biases or ensure complete factual accuracy. Because the experiments primarily focused on one-round generations produced by GPT-4, further validation is recommended before applying these findings to multi-turn conversational agents or alternative model architectures.

Cover for CoMPosT: Characterizing and Evaluating Caricature in LLM Simulations

Abstract

Recent work has aimed to capture nuances of human behavior by using LLMs to simulate responses from particular demographics in settings like social science experiments and public opinion surveys. However, there are currently no established ways to discuss or evaluate the quality of such LLM simulations. Moreover, there is growing concern that these simulations are flattened caricatures of the personas that they aim to simulate, failing to capture the multidimensionality of people and perpetuating stereotypes. To bridge these gaps, we present CoMPosT, a framework to characterize LLM simulations using four dimensions: Context, Model, Persona, and Topic. We use this framework to measure open-ended LLM simulations’ susceptibility to caricature, defined via two criteria: individuation and exaggeration. We evaluate the level of caricature in scenarios from existing work on LLM simulations. We find that for GPT-4, simulations of certain demographics (political and marginalized groups) and topics (general, uncontroversial) are highly susceptible to caricature.

Table of Contents

  • 1 Introduction
  • 2 CoMPosT: Taxonomizing Simulations
  • 3 Background: Caricature
  • 3.1 Definition of Caricature
  • 3.2 Implications of Caricature
  • 4 Caricature Detection Method
  • 4.1 Defining Defaults
  • 4.2 Measuring Individuation
  • 4.3 Measuring Exaggeration
  • 5 Experiments
  • 5.1 Online Forum
  • 5.2 Interview
  • 6 Results and Discussion
  • 6.1 Simulations of all personas can be individuated from the default-persona
  • 6.2 Exaggeration scores reveal the personas and topics most susceptible to caricature
  • 6.2.1 Caricature ↑: Topic specificity ↓
  • 6.2.2 Caricature ↑: Political ideology, race, and marginalized personas
  • 6.3 Stereotypes
  • 7 Recommendations
  • 8 Positionality
  • 9 Ethical Considerations
  • 10 Limitations
  • Acknowledgments
  • References
  • A Examples of Caricatures in Simulation
  • B Robustness of Individuation Measure
  • C Internal and External Validation of Semantic Axes
  • D Experimental Details: Topics and Personas
  • D.1 Online Forum Context
  • D.1.1 Fine-Grained Specificity Experiment
  • D.2 Interview Context
  • E Power Analysis
  • F Influence of the Context Dimension
  • G Twitter Context
  • G.1 Experimental Details
  • G.2 Results
  • H Interview Setting Result Details

Knowls

  1. Knowl 1 — CoMPosT characterizes simulations along four dimensions

    definition

    CoMPosT is a framework for specifying and comparing a large-language-model (LLM) simulation along four dimensions: Context (where and when the imagined situation occurs, including its norms, prompt wording, and requested response format), Model (the LLM used), Persona (whose opinions or actions are being simulated), and Topic (the subject or event the simulated response addresses). Context, persona, and topic are specified in the prompt; the model is selected externally. A simulation scenario can therefore be represented by its persona pp, topic tt, context cc, and model.

  2. Knowl 2 — Caricature requires both individuation and exaggeration

    definition

    In an LLM simulation, a caricature is a response that exaggerates characteristics associated with the simulated persona in a way that makes the persona distinguishable, while giving less weight to the topic than a meaningful topical response would. The two required properties are individuation—the output differs from a default-persona response to the same topic—and exaggeration—the persona-associated characteristics are amplified relative to topic-associated characteristics. Individuation alone is not sufficient: a persona may appropriately influence an answer without the answer becoming a caricature. In the paper’s evaluation procedure, exaggeration is assessed only after individuation is established.

  3. Knowl 3 — Individuation is measured by classification against a default persona

    model/method

    For a target simulation Sp,t,cS_{p,t,c} with persona pp, topic tt, and context cc, the default-persona comparison S_,t,cS_{\_,t,c} uses the same topic and context but omits a specific persona, typically substituting an unmarked term such as “person.” The default is model- and context-specific, not a universal or socially neutral baseline. To measure individuation, the authors embed outputs from the target and default-persona simulations with Sentence-BERT all-mpnet-base-v2, then train a scikit-learn random-forest binary classifier to distinguish the two sets. They use a stratified 80/20 train/test split and report test accuracy as the individuation score. Accuracy above 0.50.5 indicates differentiation better than chance; without such differentiation, the target simulation does not meet the method’s necessary condition for caricature.

  4. Knowl 4 — Exaggeration is scored on a persona–topic semantic axis

    model/method

    The exaggeration measure compares a target simulation with a contextualized semantic axis whose persona and topic poles are built from default simulations. For persona pp, topic tt, and context cc, Sp,_,cS_{p,\_,c} is the default-topic simulation (persona retained, topic omitted), and S_,t,cS_{\_,t,c} is the default-persona simulation (topic retained, persona omitted). The authors use Fightin’ Words weighted log-odds to identify words distinguishing these two sets of outputs; they use additional texts involving persona/topic p/tp/t or the corresponding defaults as a prior, and retain words with zz-score >1.96>1.96 for the relevant pole. Let WpW_p and WtW_t be the resulting persona and topic seed-word sets, with sizes kk and mm. For each seed word ww, let μ(w)\mu(w) be the mean contextualized embedding of sentences containing ww across the two default-simulation corpora. The persona-minus-topic axis is

    Vp,t=1k∑w∈Wpμ(w)−1m∑w∈Wtμ(w).V_{p,t}=\frac{1}{k}\sum_{w\in W_p}\mu(w)-\frac{1}{m}\sum_{w\in W_t}\mu(w).

    For nn target outputs s1,…,sns_1,\ldots,s_n, let e(si)e(s_i) be the contextualized embedding of output sis_i, and let cos⁡(x,y)\cos(x,y) denote cosine similarity. The reported exaggeration score is

    Ep,t,c=1n∑i=1ncos⁡(e(si),Vp,t)−cos⁡(S_,t,c,Vp,t)cos⁡(Sp,_,c,Vp,t)−cos⁡(S_,t,c,Vp,t).E_{p,t,c}=\frac{\frac{1}{n}\sum_{i=1}^{n}\cos(e(s_i),V_{p,t})-\cos(S_{\_,t,c},V_{p,t})}{\cos(S_{p,\_,c},V_{p,t})-\cos(S_{\_,t,c},V_{p,t})}.

    Here, cosine similarity for a default simulation is averaged over its outputs in the same way as for the target. The normalization expresses the target’s similarity to the axis relative to the default-persona and default-topic simulations. A larger score indicates greater persona-oriented exaggeration under this method; it is used as a proxy for caricature when individuation is present.

  5. Knowl 5 — GPT-4 simulations were tested in forum and interview settings

    experimental setup

    The main experiments used GPT-4 and generated 100 outputs for each simulation setting. In the online-forum context, the prompt framed a response as a comment posted by a persona about a topic. The study used 15 personas spanning age, political ideology, race/ethnicity, and gender, plus a neutral “person” baseline, and 30 pairs of topics selected to vary in specificity and controversy. Topics came from WikiHow categories and ProCon.org debate topics. In the interview context, the prompt requested an identity description followed by an answer; 30 questions identified as contentious in Pew’s OpinionQA dataset were converted from multiple-choice to open-ended questions, and the same personas were used. Results were averaged across the 100 generated outputs per setting. The authors also evaluated a Twitter-posting context as a robustness study.

  6. Knowl 6 — Topic generality was associated with greater exaggeration

    empirical result

    In GPT-4 online-forum simulations, general topics produced higher exaggeration scores than more specific topics, with the strongest scores for general, uncontroversial WikiHow topics. The five topics with the highest mean exaggeration were Health, Philosophy and Religion, Education and Communications, Relationships, and Finance and Business. A focused experiment varying topic specificity across five levels for health-related topics found that exaggeration decreased as specificity increased; the paper also reports no correlation between topic length and caricature in the broader analysis. Interview-topic scores were broadly comparable to those for more-specific forum topics. The inverse association between specificity and exaggeration also appeared in the Twitter-context study.

  7. Knowl 7 — Political and several marginalized personas had high exaggeration scores

    empirical result

    Across the tested GPT-4 simulations, mean exaggeration scores varied by persona. In the online-forum context, the highest-scoring personas included nonbinary, Black, Hispanic, Middle-Eastern, and conservative personas. In the interview context, the highest included nonbinary, Hispanic, 80-year-old, conservative, and Middle-Eastern personas. Overall, political leanings, non-white race/ethnicity, and nonbinary gender were among the categories most susceptible to exaggeration, while the binary-gender personas had the lowest scores. Asian and woman personas also had relatively low scores despite being members of marginalized groups; the authors note that this may reflect implicit defaults in LLM outputs. These patterns describe the tested prompts and contexts and do not establish a universal ranking of demographic groups.

  8. Knowl 8 — All tested personas were distinguishable, especially in interviews

    empirical result

    For GPT-4 outputs in both the online-forum and interview contexts, every persona had a mean individuation accuracy above 0.50.5, and its 95% confidence interval was also above 0.50.5. Thus, every tested persona was distinguishable from the corresponding default-persona output better than chance. Interview-context mean accuracy exceeded 0.950.95 for every persona. In the online-forum context, the man and woman personas were less easy to distinguish than the other personas, though their mean scores still exceeded chance. The authors attribute the higher interview scores to lower variability in that context’s sample distributions, and caution that the interview scores are consequently less useful for comparing caricature susceptibility across personas.

  9. Knowl 9 — Topic-associated patterns persisted across contexts and in Twitter simulations

    empirical result

    In context-switching experiments, the authors paired online-forum topics with the interview context and interview topics with the online-forum context. Exaggeration patterns continued to track the topics across contexts rather than being explained by context alone, although scores were slightly higher in the switched settings. In the Twitter-context study, Republican and Democrat simulations had mean individuation scores of 0.880.88 and 0.940.94, respectively. Republican simulations had higher mean exaggeration than Democrat simulations; race and other group topics generally scored higher than specific public-figure topics. Topic–persona relationships were not uniform: for Democrat simulations, race-related topics had relatively high exaggeration, and Dr. Anthony Fauci had the lowest exaggeration among the public-figure topics.

  10. Knowl 10 — The caricature metric is not a comprehensive test of simulation quality

    limitation

    The proposed measure detects one failure mode—persona exaggeration relative to a topic—not overall simulation quality or all stereotypes. A low exaggeration score does not show that a simulation is unbiased, acceptable, or free of stereotyped content; for example, the authors observed domestic-task stereotypes in simulated woman responses despite low caricature scores. The semantic-axis persona pole captures words that distinguish a persona simulation in a particular context, not a universal account of how the model represents that demographic. The study focuses on subpopulation personas and one-round responses, although the framework could be extended to multi-round simulations by evaluating the full exchange or separate parts.

Coverage note — Standalone sample generations, full seed-word tables, and the paper’s positionality, ethical discussion, and practitioner recommendations are omitted because they illustrate or interpret the framework and findings rather than add a separate core method or result.

References

  1. 1.David Adkins, Bilal Alsallakh, Adeel Cheema, Narine Kokhlikyan, Emily McReynolds, Pushkar Mishra, Chavez Procope, Jeremy Sawruk, Erin Wang, and Polina Zvyagina. 2022. Prescriptive and descriptive approaches to machine-learning transparency. In CHI Conference on Human Factors in Computing Systems Extended Abstracts, pages 1–9.
  2. 2.Gati Aher, Rosa I. Arriaga, and Adam Tauman Kalai. 2023. Using large language models to simulate multiple humans and replicate human subject studies. In Proceedings of the 40th International Conference on Machine Learning, ICML’23. JMLR.org.
  3. 3.Anu Aneja. 1993. “Jasmine,” the sweet scent of exile. Pacific Coast Philology, pages 72–80.
  4. 4.Lisa P. Argyle, Ethan C. Busby, Nancy Fulda, Joshua R. Gubler, Christopher Rytting, and David Wingate. 2023. Out of one, many: Using language models to simulate human samples. Political Analysis, 31(3):337–351.
  5. 5.April H Bailey, Adina Williams, and Andrei Cimpian. 2022. Based on billions of words on the Internet, people= men. Science Advances, 8(13):eabm2463.
  6. 6.Michiel Bakker, Martin Chadwick, Hannah Sheahan, Michael Tessler, Lucy Campbell-Gillingham, Jan Balaguer, Nat McAleese, Amelia Glaese, John Aslanides, Matt Botvinick, et al. 2022. Fine-tuning language models to find agreement among humans with diverse preferences. Advances in Neural Information Processing Systems, 35:38176–38189.
  7. 7.David Bamman, Jacob Eisenstein, and Tyler Schnoebelen. 2014. Gender identity and lexical variation in social media. Journal of Sociolinguistics, 18(2):135–160.
  8. 8.Daniel Bar-Tal, Carl F Graumann, Arie W Kruglanski, and Wolfgang Stroebe. 2013. Stereotyping and prejudice: Changing conceptions. Springer Science & Business Media.
  9. 9.Caroline Bassett. 2019. The computational therapeutic: Exploring Weizenbaum’s ELIZA as a history of the present. AI & SOCIETY, 34:803–812.
  10. 10.Joseph Bates et al. 1994. The role of emotion in believable agents. Communications of the ACM, 37(7):122–125.
  11. 11.Emily M Bender and Batya Friedman. 2018. Data statements for natural language processing: Toward mitigating system bias and enabling better science. Transactions of the Association for Computational Linguistics, 6:587–604.
  12. 12.Marcel Binz and Eric Schulz. 2023. Using cognitive psychology to understand GPT-3. Proceedings of the National Academy of Sciences, 120(6):e2218523120.
  13. 13.Irene V Blair, Jennifer E Ma, and Alison P Lenton. 2001. Imagining stereotypes away: The moderation of implicit stereotypes through mental imagery. Journal of personality and social psychology, 81(5):828.
  14. 14.Luisa N Borrell, Jennifer R Elhawary, Elena Fuentes-Afflick, Jonathan Witonsky, Nirav Bhakta, Alan HB Wu, Kirsten Bibbins-Domingo, José R Rodríguez-Santana, Michael A Lenoir, James R Gavin III, et al. 2021. Race and genetic ancestry in medicine—a time for reckoning with racism. New England Journal of Medicine, 384(5):474–480.
  15. 15.Leslie Bow. 2019. Racist cute: Caricature, kawaii-style, and the Asian thing. American Quarterly, 71(1):29–58.
  16. 16.Robert Bowman, Camille Nadal, Kellie Morrissey, Anja Thieme, and Gavin Doherty. 2023. Using thematic analysis in healthcare HCI at CHI: A scoping review. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, pages 1–18.
  17. 17.James DJ Brown. 2010. A stereotype, wrapped in a cliché, inside a caricature: Russian foreign policy and orientalism. Politics, 30(3):149–159.
  18. 18.Yang Cao, Anna Sotnikova, Hal Daumé III, Rachel Rudinger, and Linda Zou. 2022. Theory-grounded measurement of US social stereotypes in english language models. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1276–1295.
  19. 19.Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, and Chiyuan Zhang. 2023. Quantifying memorization across neural language models. In The Eleventh International Conference on Learning Representations.
  20. 20.Eshwar Chandrasekharan, Mattia Samory, Shagun Jhaver, Hunter Charvat, Amy Bruckman, Cliff Lampe, Jacob Eisenstein, and Eric Gilbert. 2018. The internet’s hidden rules: An empirical study of Reddit norm violations at micro, meso, and macro scales. Proceedings of the ACM on Human-Computer Interaction, 2(CSCW):1–25.
  21. 21.Myra Cheng, Maria De-Arteaga, Lester Mackey, and Adam Tauman Kalai. 2023a. Social norm bias: Residual harms of fairness-aware algorithms. Data Mining and Knowledge Discovery, pages 1–27.
  22. 22.Myra Cheng, Esin Durmus, and Dan Jurafsky. 2023b. Marked personas: Using natural language prompts to measure stereotypes in language models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, pages 1276–1295.
  23. 23.Alexander M Czopp, Aaron C Kay, and Sapna Cheryan. 2015. Positive stereotypes are pervasive and powerful. Perspectives on Psychological Science, 10(4):451–463.
  24. 24.Eberhard Demm. 1993. Propaganda and caricature in the first world war. Journal of Contemporary History, 28(1):163–192.
  25. 25.Yann Dubois, Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Alpaca-Farm: A simulation framework for methods that learn from human feedback.
  26. 26.Aparna Elangovan, Jiayuan He, and Karin Verspoor. 2021. Memorization vs. generalization: Quantifying data leakage in NLP performance evaluation. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 1325–1335, Online. Association for Computational Linguistics.
  27. 27.Susan T Fiske, Amy JC Cuddy, Peter Glick, and Jun Xu. 2002. A model of (often mixed) stereotype content: competence and warmth respectively follow from perceived status and competition. Journal of personality and social psychology, 82(6):878.
  28. 28.Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé Iii, and Kate Crawford. 2021. Datasheets for datasets. Communications of the ACM, 64(12):86–92.
  29. 29.Peter Gottschalk and Gabriel Greenberg. 2011. From Muhammad to Obama: Caricatures, cartoons, and stereotypes of Muslims. Islamophobia: The challenge of pluralism in the 21st century, pages 191–210.
  30. 30.Jonathan Grudin. 2006. Why personas work: The psychological evidence. The persona lifecycle, 12:642–664.
  31. 31.Perttu Hämäläinen, Mikke Tavast, and Anton Kunnari. 2023. Evaluating large language models in generating synthetic HCI research data: a case study. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, pages 1–19.
  32. 32.Sil Hamilton. 2023. Blind judgement: Agent-based supreme court modelling with GPT. arXiv preprint arXiv:2301.05327.
  33. 33.Alex Hanna, Emily Denton, Andrew Smart, and Jamila Smith-Loud. 2020. Towards a critical race methodology in algorithmic fairness. In Proceedings of the 2020 conference on Fairness, Accountability, and Transparency, pages 501–512.
  34. 34.Madeline E Heilman. 2001. Description and prescription: How gender stereotypes prevent women’s ascent up the organizational ladder. Journal of Social Issues, 57(4):657–674.
  35. 35.Joseph Henrich, Steven J Heine, and Ara Norenzayan. 2010. The weirdest people in the world? Behavioral and brain sciences, 33(2-3):61–83.
  36. 36.Susan C Herring and John C Paolillo. 2006. Gender and genre variation in weblogs. Journal of Sociolinguistics, 10(4):439–459.
  37. 37.James D Hollan, Edwin L Hutchins, and Louis Weitzman. 1984. Steamer: An interactive inspectable simulation-based training system. AI magazine, 5(2):15–15.
  38. 38.Marjan Hosseinia, Eduard Dragut, and Arjun Mukherjee. 2020. Stance prediction for contemporary issues: Data and experiments. In Proceedings of the Eighth International Workshop on Natural Language Processing for Social Media, pages 32–40, Online. Association for Computational Linguistics.
  39. 39.Hang Jiang, Doug Beeferman, Brandon Roy, and Deb Roy. 2022. CommunityLM: Probing partisan worldviews from language models. In Proceedings of the 29th International Conference on Computational Linguistics, pages 6818–6826, Gyeongju, Republic of Korea. International Committee on Computational Linguistics.
  40. 40.Randolph M Jones, John E Laird, Paul E Nielsen, Karen J Coulter, Patrick Kenny, and Frank V Koss. 1999. Automated intelligent pilots for combat flight simulation. AI magazine, 20(1):27–27.
  41. 41.Lee Jussim, Jarret T Crawford, Stephanie M Anglin, Sean T Stevens, and Jose L Duarte. 2016. Interpretations and methods: Towards a more effectively self-correcting social psychology. Journal of Experimental Social Psychology, 66:116–133.
  42. 42.Gauri Kambhatla, Ian Stewart, and Rada Mihalcea. 2022. Surfacing racial stereotypes through identity portrayal. In 2022 ACM Conference on Fairness, Accountability, and Transparency, pages 1604–1615.
  43. 43.Os Keyes, Burren Peil, Rua M Williams, and Katta Spiel. 2020. Reimagining (women’s) health: HCI, gender and essentialised embodiment. ACM Transactions on Computer-Human Interaction (TOCHI), 27(4):1–42.
  44. 44.Mahnaz Koupaee and William Yang Wang. 2018. Wiki-how: A large scale text summarization dataset. arXiv preprint arXiv:1810.09305.
  45. 45.Kay Dian Kriz. 2008. Slavery, sugar, and the culture of refinement: Picturing the British West Indies, 1700–1840. Paul Mellon Centre.
  46. 46.Neha Kumar, Naveena Karusala, Azra Ismail, Marisol Wong-Villacres, and Aditya Vishwanath. 2019. Engaging feminist solidarity for comparative research, design, and practice. Proceedings of the ACM on Human-Computer Interaction, 3(CSCW):1–24.
  47. 47.Patrick Lewis, Pontus Stenetorp, and Sebastian Riedel. 2021. Question and answer test-train overlap in open-domain question answering datasets. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 1000–1008.
  48. 48.Calvin A Liang, Sean A Munson, and Julie A Kientz. 2021. Embracing four tensions in human-computer interaction research with marginalized people. ACM Transactions on Computer-Human Interaction (TOCHI), 28(2):1–47.
  49. 49.Ruibo Liu, Ruixin Yang, Chenyan Jia, Ge Zhang, Denny Zhou, Andrew M Dai, Diyi Yang, and Soroush Vosoughi. 2023. Training socially aligned language models in simulated human society. arXiv preprint arXiv:2305.16960.
  50. 50.Li Lucy, Divya Tadimeti, and David Bamman. 2022. Discovering differences in the representation of people using contextualized semantic axes. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 3477–3494, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  51. 51.Bohun Lynch. 1927. A history of caricature. Faber and Gwyer.
  52. 52.Julia M. Markel, Steven G. Opferman, James A. Landay, and Chris Piech. 2023. GPTeach: Interactive TA training with GPT-based students. In Proceedings of the Tenth ACM Conference on Learning @ Scale, L@S ’23, page 226–236, New York, NY, USA. Association for Computing Machinery.
  53. 53.Nicola Marsden and Maren Haag. 2016. Stereotypes and politics: Reflections on personas. In Proceedings of the 2016 CHI conference on human factors in computing systems, pages 4017–4031.
  54. 54.Nicola Marsden and Monika Pröbster. 2019. Personas and identity: Looking at multiple identities to inform the construction of personas. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems, pages 1–14.
  55. 55.Melissa McCracken, Miho Olsen, Moon S Chen Jr, Ahmedin Jemal, Michael Thun, Vilma Cokkinides, Dennis Deapen, and Elizabeth Ward. 2007. Cancer incidence, mortality, and associated risk factors among Asian Americans of Chinese, Filipino, Vietnamese, Korean, and Japanese ethnicities. CA: a cancer journal for clinicians, 57(4):190–205.
  56. 56.Amita Misra, Brian Ecker, and Marilyn Walker. 2016. Measuring the similarity of sentential arguments in dialogue. In Proceedings of the 17th Annual Meeting of the Special Interest Group on Discourse and Dialogue, pages 276–287, Los Angeles. Association for Computational Linguistics.
  57. 57.Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. 2019. Model cards for model reporting. In Proceedings of the Conference on Fairness, Accountability, and Transparency, pages 220–229.
  58. 58.Chandra Mohanty. 1988. Under Western eyes: Feminist scholarship and colonial discourses. Feminist review, 30(1):61–88.
  59. 59.Burt L Monroe, Michael P Colaresi, and Kevin M Quinn. 2008. Fightin’words: Lexical feature selection and evaluation for identifying the content of political conflict. Political Analysis, 16(4):372–403.
  60. 60.OpenAI. 2023. GPT-4 technical report. arXiv.
  61. 61.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744.
  62. 62.Joon Sung Park, Joseph C O’Brien, Carrie J Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein. 2023a. Generative agents: Interactive simulacra of human behavior. arXiv preprint arXiv:2304.03442.
  63. 63.Joon Sung Park, Lindsay Popowski, Carrie Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein. 2022. Social simulacra: Creating populated prototypes for social computing systems. In Proceedings of the 35th Annual ACM Symposium on User Interface Software and Technology, pages 1–18.
  64. 64.Peter S Park, Philipp Schoenegger, and Chongyang Zhu. 2023b. Artificial intelligence in psychology research. arXiv preprint arXiv:2302.07267.
  65. 65.Andi Peng, Besmira Nushi, Emre Kiciman, Kori Inkpen, and Ece Kamar. 2022. Investigations of performance and bias in human-AI teamwork in hiring. Proceedings of the AAAI Conference on Artificial Intelligence, 36(11):12089–12097.
  66. 66.David Perkins. 1975. A definition of caricature and caricature and recognition. Studies in Visual Communication, 2(1):1–24.
  67. 67.Scott Ed Plous. 2003. Understanding prejudice and discrimination. McGraw-Hill.
  68. 68.Jen’nan Ghazal Read, Scott M Lynch, and Jessica S West. 2021. Disaggregating heterogeneity among non-hispanic whites: evidence and implications for us racial/ethnic health disparities. Population research and policy review, 40:9–31.
  69. 69.Nils Reimers and Iryna Gurevych. 2019. Sentence-bert: Sentence embeddings using siamese bert-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). Association for Computational Linguistics.
  70. 70.Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto. 2023. Whose opinions do language models reflect? In Proceedings of the 40th International Conference on Machine Learning, ICML’23. JMLR.org.
  71. 71.Sebastin Santy, Jenny Liang, Ronan Le Bras, Katharina Reinecke, and Maarten Sap. 2023. NLPositionality: Characterizing design biases of datasets and models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 9080–9102, Toronto, Canada. Association for Computational Linguistics.
  72. 72.Johannes Schneider, Christian Meske, and Michalis Vlachos. 2020. Deceptive AI explanations: Creation and detection. In International Conference on Agents and Artificial Intelligence. SciTePress.
  73. 73.Marshall H Segall, Donald Thomas Campbell, and Melville Jean Herskovits. 1966. The influence of culture on visual perception, volume 310. Bobbs-Merrill Indianapolis.
  74. 74.Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Yarin Gal, Nicolas Papernot, and Ross Anderson. 2023. The curse of recursion: Training on generated data makes models forget. arXiv preprint arxiv:2305.17493.
  75. 75.Phillip R Slavney. 1984. Histrionic personality and antisocial personality: caricatures of stereotypes? Comprehensive Psychiatry, 25(2):129–141.
  76. 76.Irene Solaiman and Christy Dennison. 2021. Process for adapting language models to society (palms) with values-targeted datasets. Advances in Neural Information Processing Systems, 34:5861–5873.
  77. 77.Keita Takayama. 2017. Imagining east asian education otherwise: Neither caricature, nor scandalization. Asia Pacific Journal of Education, 37(2):262–274.
  78. 78.Kushal Tirumala, Aram Markosyan, Luke Zettlemoyer, and Armen Aghajanyan. 2022. Memorization without overfitting: Analyzing the training dynamics of large language models. Advances in Neural Information Processing Systems, 35:38274–38290.
  79. 79.Andrea Aler Tubella, Dimitri Coelho Mollo, Adam Dahlgren Lindström, Hannah Devinney, Virginia Dignum, Petter Ericson, Anna Jonsson, Timotheus Kampik, Tom Lenaerts, Julian Alfredo Mendez, et al. 2023. ACROCPoLis: A descriptive framework for making sense of fairness. In 2023 ACM Conference on Fairness, Accountability, and Transparency, pages 1604–1615.
  80. 80.Veniamin Veselovsky, Manoel Horta Ribeiro, and Robert West. 2023. Artificial artificial artificial intelligence: Crowd workers widely use large language models for text production tasks. arXiv preprint arXiv:2306.07899.
  81. 81.Kailas Vodrahalli, Roxana Daneshjou, Tobias Gerstenberg, and James Zou. 2022. Do humans trust advice more if it comes from AI? an analysis of human-AI interactions. In Proceedings of the 2022 AAAI/ACM Conference on AI, Ethics, and Society, pages 763–777.
  82. 82.Angelina Wang, Vikram V Ramaswamy, and Olga Russakovsky. 2022. Towards intersectionality in machine learning: Including more identities, handling underrepresentation, and performing evaluation. In 2022 ACM Conference on Fairness, Accountability, and Transparency, pages 336–349.
  83. 83.Joseph Weizenbaum. 1966. ELIZA—a computer program for the study of natural language communication between man and machine. Communications of the ACM, 9(1):36–45.
  84. 84.Robert Wolfe and Aylin Caliskan. 2022. Markedness in visual semantic AI. 2022 ACM Conference on Fairness, Accountability, and Transparency.
  85. 85.Fuzhao Xue, Yao Fu, Wangchunshu Zhou, Zangwei Zheng, and Yang You. 2023. To repeat or not to repeat: Insights from scaling llm under token-crisis. arXiv preprint arXiv:2305.13230.
  86. 86.Diyi Yang. 2019. Computational Social Roles. Ph.D. thesis, Carnegie Mellon University Pittsburgh, PA, USA.
  87. 87.JD Zamfirescu-Pereira, Richmond Y Wong, Bjoern Hartmann, and Qian Yang. 2023. Why Johnny can’t prompt: how non-AI experts try (and fail) to design llm prompts. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, pages 1–21.
  88. 88.Caleb Ziems, Jane Dwivedi-Yu, Yi-Chia Wang, Alon Halevy, and Diyi Yang. 2023. NormBank: A knowledge bank of situational social norms. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 7756–7776, Toronto, Canada. Association for Computational Linguistics.

Citation

MLA
Cheng, M., et al. “CoMPosT: Characterizing and Evaluating Caricature in LLM Simulations”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 10853–75, https://doi.org/10.18653/v1/2023.emnlp-main.669.
APA
Cheng, M., Piccardi, T., & Yang, D. (2023). CoMPosT: Characterizing and Evaluating Caricature in LLM Simulations. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 10853–10875. https://doi.org/10.18653/v1/2023.emnlp-main.669
Chicago
Cheng, M., T. Piccardi, and D. Yang. 2023. “CoMPosT: Characterizing and Evaluating Caricature in LLM Simulations”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 10853–75. https://doi.org/10.18653/v1/2023.emnlp-main.669.
Harvard
Cheng, M., Piccardi, T. and Yang, D. (2023) “CoMPosT: Characterizing and Evaluating Caricature in LLM Simulations”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 10853–10875. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.669.
Vancouver
1. Cheng M, Piccardi T, Yang D (2023) CoMPosT: Characterizing and Evaluating Caricature in LLM Simulations. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 10853–10875

BibTeX

@inproceedings{cheng-etal-2023-compost,
    title = "{C}o{MP}os{T}: Characterizing and Evaluating Caricature in {LLM} Simulations",
    author = "Cheng, Myra  and
      Piccardi, Tiziano  and
      Yang, Diyi",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.669/",
    doi = "10.18653/v1/2023.emnlp-main.669",
    pages = "10853--10875"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/