The steerability of large language models toward data-driven personas
Junyi LiCharith PerisNinareh MehrabiPalash GoyalKai-Wei ChangAram GalstyanRichard S. ZemelRahul Gupta
Proposes a collaborative filtering framework that maps survey responses into continuous embedding spaces to steer large language models toward distinct individual and group viewpoints more accurately than demographic-based prompting.
Large language models often exhibit viewpoint bias by over-representing certain populations while under-representing others, such as older adults or specific religious groups. Addressing this issue is critical as these systems are increasingly deployed across high-stakes domains like healthcare, finance, and education. Traditional attempts to steer model outputs rely on broad demographic categories like age, race, or political party, which fail to capture the nuanced, cross-cutting beliefs found within real populations.
The article demonstrates a data-driven approach to steer large language models toward distinct personas based on actual opinion patterns rather than static demographic traits. The primary objective is to evaluate whether defining personas through collaborative filtering improves the controllable generation of diverse viewpoints.
To accomplish this, the authors used survey data from the OpinionQA dataset, covering 18,339 participants and 1,476 multiple-choice questions across 23 topics. They applied collaborative filtering to map individual response histories into continuous 16-dimensional vectors, creating individual personas as well as six distinct cluster personas grouped by response similarity. A compact soft-prompting model—a two-layer neural network—was trained to convert these opinion vectors into virtual tokens that steer frozen language models without requiring full model retraining. The approach was benchmarked against standard prompting, demographic prompting, and context-based prompting across four language models (GPT-Neo-1.3B, GPT-Neo-2.7B, GPT-j-6B, and Falcon-7B-Instruct).
The analysis yielded several key findings. First, steering language models using individual data-driven personas achieved prediction accuracies between roughly 60% and 62%, representing a 57% to 77% relative improvement over the best baseline methods, which only achieved 30% to 39% accuracy. Second, cluster personas representing opinion centroids performed within 8% to 12% of individual embeddings and consistently outperformed demographic-based embeddings (such as political party traits) by up to 2.59%. Third, the method generalized effectively to unseen individuals: using just a single response to construct a user vector yielded roughly 39% to 41% accuracy, outperforming the baseline methods, with accuracy steadily rising above 57% as more individual responses were observed. Finally, demographic analysis revealed that the opinion clusters naturally contained diverse mixes of demographic traits, confirming that shared viewpoints do not strictly align with standard demographic boundaries.
These findings indicate that data-driven personas offer a far more accurate, scalable, and computationally efficient way to represent diverse viewpoints in artificial intelligence systems. By avoiding full model fine-tuning and instead using lightweight soft prompting, organizations can reduce training costs and compute overhead while preventing polarization and ensuring under-represented perspectives are faithfully simulated.
Organizations deploying conversational systems should consider replacing static demographic personas with data-driven opinion embeddings when simulating human perspectives. However, practitioners must conduct careful audits of source opinion datasets prior to deployment, as steering models toward specific user viewpoints can inadvertently amplify harmful biases present in the training data. Future work should validate this approach across additional datasets and survey formats, as well as test other parameter-efficient adaptation methods like low-rank adaptation.
- Paper: Whose Opinions Do Language Models Reflect?, Shibani Santurkar et al. (2023). Its survey-based measurement of whose opinions language models reflect establishes the representation gap that the source addresses with data-driven personas.
- Paper: Out of One, Many: Using Language Models to Simulate Human Samples, Lisa P. Argyle et al. (2022). Its “silicon sampling” approach shows how real people’s demographic and attitudinal profiles can condition LLM simulations, a direct precursor to the source’s opinion-based personas.
- Paper: Tailor: A Soft-Prompt-Based Approach to Attribute-Based Controlled Text Generation, Kexin Yang et al. (2023). Its soft-prompt method for steering a frozen language model provides the parameter-efficient control technique that helps explain the source’s virtual-token approach.
No sufficiently relevant recommendations were found.
