CoMPosT: Characterizing and Evaluating Caricature in LLM Simulations
Myra ChengTiziano PiccardiDiyi Yang
Presents a four-dimension framework and quantitative metric to evaluate how demographic simulations in large language models reduce complex human personas to exaggerated, stereotypical caricatures.
Researchers and organizations increasingly use large language models (LLMs) to simulate human behavior across social science experiments, public opinion polling, and product testing. Despite rapid adoption, there are no established methodologies to evaluate the quality of open-ended simulated responses. This gap creates significant risk, as automated simulations can easily produce flattened, exaggerated caricatures of demographic groups rather than meaningful, topical dialogue.
The article introduces CoMPosT, a structured framework designed to define LLM simulations and systematically measure their susceptibility to caricature. The primary objective is to evaluate how simulated personas and discussion topics interact to generate distorted, one-dimensional depictions of specific human populations.
The CoMPosT framework categorizes any simulation along four core dimensions: Context, Model, Persona, and Topic. To detect caricature, the article operationalizes two sequential criteria: individuation, which tests whether a simulated response can be distinguished from a neutral baseline using a binary classifier, and exaggeration, which measures whether the text disproportionately amplifies identity-related keywords over topic-relevant content using semantic axes. The authors evaluated open-ended generations from GPT-4 across multiple realistic scenarios, including online discussion forums, structured interviews, and social media posting. The study analyzed 15 distinct personas spanning age, political ideology, race and ethnicity, and gender across 60 discussion topics, sampling 100 responses per configuration.
The findings reveal that GPT-4 is highly susceptible to producing caricatures under specific conditions. First, simulations of political groups (such as conservatives) and marginalized demographics (including nonbinary, Black, Hispanic, and Middle-Eastern personas) exhibit the highest rates of exaggeration. Second, topic specificity is inversely related to caricature: general, uncontroversial subjects (such as health, relationships, or general technology) generate substantially higher levels of caricature than highly specific or contentious prompts. In these general scenarios, simulated personas frequently abandon the core subject to output generic identity slogans and activism narratives. Third, binary gender groups (men and women) exhibited the lowest caricature scores, reflecting the model's tendency to rely on implicit default personas, even though qualitative stereotypes were still present in baseline prompts.
These results demonstrate that unvetted LLM simulations pose operational and ethical risks. Relying on synthetic personas for market research, policy development, or behavioral modeling can mislead decision-makers through false homogeneity and outdated stereotypes. Organizations risk basing real-world strategies on artificial caricatures that fail to represent the actual diversity and nuanced opinions of target demographics.
To mitigate these risks, practitioners should avoid using broad, generic topics when prompting LLMs and instead provide fine-grained, highly contextualized prompts. Organizations should systematically evaluate synthetic personas using differentiation and exaggeration metrics before deploying them in production or research. Furthermore, simulation workflows should incorporate multifaceted persona descriptions and document researcher positionality to prevent outgroup bias from skewing simulation designs.
These conclusions should be interpreted within certain boundaries. The evaluation metrics detect specific forms of semantic exaggeration and differentiation, meaning low caricature scores do not guarantee the total absence of subtle biases or ensure complete factual accuracy. Because the experiments primarily focused on one-round generations produced by GPT-4, further validation is recommended before applying these findings to multi-turn conversational agents or alternative model architectures.
- Paper: Out of One, Many: Using Language Models to Simulate Human Samples, Lisa P. Argyle et al. (2022). This earlier study establishes persona-conditioned “silicon sampling” as a method for simulating human populations, the methodological starting point for CoMPosT’s analysis of persona-driven caricature.
- Paper: Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies, Gati V. Aher et al. (2023). Its Turing Experiment evaluates LLMs as multiple human subjects across behavioral studies, providing a direct precursor to CoMPosT’s framework for assessing the quality and limitations of human simulations.
- Paper: Whose Opinions Do Language Models Reflect?, Shibani Santurkar et al. (2023). Its measurement of how LLM opinions align with demographic groups grounds CoMPosT’s concern with whether simulated personas represent real populations rather than flattened stereotypes.
- Paper: Personalizing Dialogue Agents: I have a dog, do you have pets too?, Saizheng Zhang et al. (2018). This work introduces explicit persona profiles for dialogue agents, clarifying the persona-conditioning approach whose effects CoMPosT evaluates in simulated responses.
- Paper: Language (Technology) is Power: A Critical Survey of “Bias” in NLP, Su Lin Blodgett et al. (2020). Its conceptual distinction between representational harms and other forms of bias provides essential grounding for CoMPosT’s focus on stereotypes and caricature.
- Paper: "I'm sorry to hear that": Finding New Biases in Language Models with a Holistic Descriptor Dataset, Eric Michael Smith et al. (2022). Its systematic evaluation of demographic stereotypes in model outputs supplies a relevant foundation for CoMPosT’s criteria of exaggeration and individuation.
- Paper: Quantifying the Persona Effect in LLM Simulations, Tiancheng Hu et al. (2024). Building on CoMPosT’s warning that persona simulations can distort group representation, this study quantifies how little demographic traits explain human judgments and tests the limits of persona prompting.
- Paper: Is this the real life? Is this just fantasy? The Misleading Success of Simulating Social Interactions With LLMs, Xuhui Zhou et al. (2024). Extending CoMPosT’s critique of unrealistic human simulations, this study shows how omniscient role-play inflates apparent social competence compared with interactions under realistic information asymmetry.
- Paper: Evaluating Large Language Models as Generative User Simulators for Conversational Recommendation, Se-eun Yoon et al. (2024). Applying the challenge of evaluating simulated human behavior to conversational recommendation, this study tests whether LLM-generated users reproduce real preferences and interaction patterns.
