Built independently by an author, for readers. Read the story and support ChapterPal

keyword

R²

R-squared, also known as the coefficient of determination, is a statistical measure that represents the proportion of variance in a dependent variable that is explained or predicted by one or more independent variables in a regression model. Ranging typically between 0 and 1, or 0% and 100%, a value of 0 indicates that the model explains none of the variability of the response data around its mean, whereas a value of 1 indicates that it explains all of the variability. In simple linear regression with a single predictor, it equals the square of the Pearson correlation coefficient between the two variables. It serves as a standard metric for evaluating model goodness of fit and quantifying the explanatory strength of statistical relationships.

1 item

Quantifying the Persona Effect in LLM Simulations

Quantifying the Persona Effect in LLM Simulations

Tiancheng Hu, Nigel Collier

OrganizationsUniversity of Cambridge

Why you should read this

Quantifies the limits and efficacy of persona prompting across subjective NLP tasks, establishing that demographic variables explain under ten percent of annotation variance yet enable large language models to recover most predictable human variation when strong correlations exist.

Large language models (LLMs) have shown remarkable promise in simulating human language and behavior. This study investigates how integrating persona variables—demographic, social, and behavioral factors—impacts LLMs’ ability to simulate diverse perspectives. We find that persona variables account for <10% variance in annotating existing subjective NLP datasets. Nonetheless, incorporating persona variables via prompting in LLMs provides modest but statistically significant improvements. Persona prompting is most effective in samples where many annotators disagree, but their disagreements are relatively minor. Notably, we find a linear relationship in our setting: the stronger the correlation between persona variables and human annotations, the more accurate the LLM predictions are using persona prompting. In a zero-shot setting, a powerful 70b model with persona prompting captures 81% of the annotation variance achievable by linear regression trained on ground truth annotations. However, for most subjective NLP datasets, where persona variables have limited explanatory power, the benefits of persona prompting are limited.

Added

2026-09-30