Built independently by an author, for readers. Read the story and support ChapterPal

keyword

Persona variables

Persona variables are specific demographic, social, psychological, and behavioral attributes used to define the identity, background, or profile of an individual or simulated agent. In the context of artificial intelligence and computational modeling, these variables encompass characteristics such as age, gender, geographic location, socioeconomic background, political ideology, cultural affiliation, and personal beliefs. They are commonly incorporated into system prompts or conditioning frameworks to steer large language models toward reflecting distinct human viewpoints. By parameterizing identity in this way, persona variables allow researchers and developers to simulate diverse annotator behaviors, analyze subjective interpretations of text, and evaluate how background attributes influence decision-making and language generation.

1 item

Quantifying the Persona Effect in LLM Simulations

Quantifying the Persona Effect in LLM Simulations

Tiancheng Hu, Nigel Collier

OrganizationsUniversity of Cambridge

Why you should read this

Quantifies the limits and efficacy of persona prompting across subjective NLP tasks, establishing that demographic variables explain under ten percent of annotation variance yet enable large language models to recover most predictable human variation when strong correlations exist.

Large language models (LLMs) have shown remarkable promise in simulating human language and behavior. This study investigates how integrating persona variables—demographic, social, and behavioral factors—impacts LLMs’ ability to simulate diverse perspectives. We find that persona variables account for <10% variance in annotating existing subjective NLP datasets. Nonetheless, incorporating persona variables via prompting in LLMs provides modest but statistically significant improvements. Persona prompting is most effective in samples where many annotators disagree, but their disagreements are relatively minor. Notably, we find a linear relationship in our setting: the stronger the correlation between persona variables and human annotations, the more accurate the LLM predictions are using persona prompting. In a zero-shot setting, a powerful 70b model with persona prompting captures 81% of the annotation variance achievable by linear regression trained on ground truth annotations. However, for most subjective NLP datasets, where persona variables have limited explanatory power, the benefits of persona prompting are limited.

Added

2026-09-30