Out of One, Many: Using Language Models to Simulate Human Samples
Lisa P. ArgyleEthan C. BusbyNancy FuldaJoshua GublerChristopher RyttingDavid Wingate
Demonstrates that large language models conditioned on detailed demographic profiles can accurately replicate human survey response distributions across diverse subgroups, establishing artificial intelligence as a viable tool for computational social science research.
Social science research and public opinion polling face rising costs and operational hurdles when sampling human populations. At the same time, artificial intelligence systems are often criticized for exhibiting uniform algorithmic bias. The article evaluates whether large-scale generative language models, specifically GPT-3, can overcome skewed training data to serve as accurate proxies for targeted human sub-populations.
The article demonstrates that algorithmic bias within language models is fine-grained and demographically correlated rather than monolithic. By introducing a method called "silicon sampling," the researchers condition the model on first-person demographic and attitudinal profiles from real survey participants, testing the model's "algorithmic fidelity"—the degree to which simulated outputs mirror the nuanced attitudes and behavioral patterns of human groups.
To establish credibility across multiple contexts, the researchers conducted three empirical studies using data from thousands of real human respondents across the 2012, 2016, and 2020 American National Election Studies as well as a prior partisan polarization survey. In the first study, human evaluators assessed whether GPT-3 could generate free-form text descriptions of political parties matching human responses. In the second study, the model predicted presidential voting choices across three election cycles based on demographic inputs. In the third study, the model completed a virtual 12-question interview to evaluate complex inter-relationships across political attitudes, demographics, and social behaviors.
The key findings show high algorithmic fidelity across all evaluations. First, human evaluators could not distinguish GPT-3 text from human-written text: participants correctly guessed human-written lists were human 61.7% of the time and GPT-3 lists were human 61.2% of the time, while evaluations of tone, extremity, and content matched closely. Second, GPT-3 accurately predicted presidential voting patterns across all three election cycles, yielding high raw agreement (85% in 2012, 87% in 2016, and 89% in 2020) and strong tetrachoric correlations of 0.90 to 0.94 against actual survey choices across various subgroups. Third, in complex closed-ended multi-variable surveys, the model faithfully mirrored human correlation structures across eleven demographic and attitudinal variables, showing a minimal mean difference in association metrics of just -0.026. Fourth, the model maintained high predictive fidelity even for the 2020 election, which occurred after the model's 2019 training cutoff date.
These findings suggest language models capture deep structural associations between socio-demographic contexts and human attitudes, rather than just superficial text styles. This capability provides a cost-effective method to pilot experimental designs, test survey question wording, triage potential confounds, and refine theories before deploying expensive field research. However, the high fidelity also poses serious ethical and security risks, as actors could leverage targeted simulation for automated disinformation, targeted manipulation, or fraudulent campaigns.
Organizations and researchers should consider adopting silicon sampling as an exploratory pre-testing tool prior to human deployments, which can drastically cut preliminary research costs. Before strong decisions or public policies are based solely on silicon samples, researchers must establish algorithmic fidelity within each specific domain of study. Future work must extend evaluations beyond U.S. political opinion, optimize prompt templates, and develop community-wide ethical safeguards and oversight frameworks to detect and mitigate potential abuse.
Confidence in these findings is high for U.S. political attitudes and aggregate subgroup distributions within the examined surveys. However, readers should exercise caution because the model does not predict individual human responses deterministically. Performance weakens significantly among politically unaligned individuals (such as pure political independents), and high missing data rates or prompt non-compliance can occur depending on temperature settings and question formatting.
- Paper: Language Models are Few-Shot Learners, T. B. Brown et al. (2020). Introduces GPT-3 and its few-shot conditioning capabilities, providing the core model and prompting paradigm that the source uses to generate 'silicon samples'.
- Paper: Semantics derived automatically from language corpora contain human-like biases, Aylin Caliskan et al. (2016). Establishes how standard text embeddings absorb nuanced human cultural attitudes and demographic associations, laying the conceptual foundation for treating language model bias as reflecting human social spectra.
- Paper: Language (Technology) is Power: A Critical Survey of “Bias” in NLP, Su Lin Blodgett et al. (2020). Provides a critical taxonomy of representational harms and social biases in NLP, establishing the literature on algorithmic bias that the source reconceptualizes as fine-grained algorithmic fidelity.
- Paper: Personalizing Dialogue Agents: I have a dog, do you have pets too?, Saizheng Zhang et al. (2018). Demonstrates the technique of conditioning dialogue agents on explicit textual persona profiles to maintain consistent persona-driven responses, directly informing demographic backstory prompting.
- Paper: StereoSet: Measuring stereotypical bias in pretrained language models, Moin Nadeem et al. (2020). Systematically quantifies stereotypical demographic associations across pretrained language models, which the source builds upon to test demographic emulation.
- Paper: On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? 🦜, Emily M. Bender et al. (2021). Critically analyzes how massive language models encode and reflect the disparate viewpoints of their training corpora, framing the debate around using model outputs to represent human groups.
- Paper: Whose Opinions Do Language Models Reflect?, Shibani Santurkar et al. (2023). Critiques and directly tests the representativeness of language models across 60 demographic subgroups on Pew survey data, evaluating how well models actually reflect diverse human opinions.
- Paper: Generative Agents: Interactive Simulacra of Human Behavior, Joon Sung Park et al. (2023). Extends the concept of silicon human proxies by placing generative agents with detailed backstories into an interactive sandbox to simulate individual and emergent collective social behavior.
- Paper: ChatGPT outperforms crowd workers for text-annotation tasks, Fabrizio Gilardi et al. (2023). Applies the premise of language models acting as effective substitutes for human subjects to social science text-annotation and coding tasks.
- Paper: A survey on large language model based autonomous agents, Lei Wang et al. (2023). Surveys the broader architectural framework and applications of LLM-based autonomous agents simulating human personas, planning, and behaviors.
- Paper: Towards Understanding Sycophancy in Language Models, Mrinank Sharma et al. (2023). Investigates sycophancy in LLMs, uncovering behavioral biases that complicate the reliable simulation of genuine, unconditioned human viewpoints.
