Out of One, Many: Using Language Models to Simulate Human Samples

Lisa P. ArgyleEthan C. BusbyNancy FuldaJoshua GublerChristopher RyttingDavid Wingate

article2022Political Analysis1,364 citations

Demonstrates that large language models conditioned on detailed demographic profiles can accurately replicate human survey response distributions across diverse subgroups, establishing artificial intelligence as a viable tool for computational social science research.

Listen

Social science research and public opinion polling face rising costs and operational hurdles when sampling human populations. At the same time, artificial intelligence systems are often criticized for exhibiting uniform algorithmic bias. The article evaluates whether large-scale generative language models, specifically GPT-3, can overcome skewed training data to serve as accurate proxies for targeted human sub-populations.

The article demonstrates that algorithmic bias within language models is fine-grained and demographically correlated rather than monolithic. By introducing a method called "silicon sampling," the researchers condition the model on first-person demographic and attitudinal profiles from real survey participants, testing the model's "algorithmic fidelity"—the degree to which simulated outputs mirror the nuanced attitudes and behavioral patterns of human groups.

To establish credibility across multiple contexts, the researchers conducted three empirical studies using data from thousands of real human respondents across the 2012, 2016, and 2020 American National Election Studies as well as a prior partisan polarization survey. In the first study, human evaluators assessed whether GPT-3 could generate free-form text descriptions of political parties matching human responses. In the second study, the model predicted presidential voting choices across three election cycles based on demographic inputs. In the third study, the model completed a virtual 12-question interview to evaluate complex inter-relationships across political attitudes, demographics, and social behaviors.

The key findings show high algorithmic fidelity across all evaluations. First, human evaluators could not distinguish GPT-3 text from human-written text: participants correctly guessed human-written lists were human 61.7% of the time and GPT-3 lists were human 61.2% of the time, while evaluations of tone, extremity, and content matched closely. Second, GPT-3 accurately predicted presidential voting patterns across all three election cycles, yielding high raw agreement (85% in 2012, 87% in 2016, and 89% in 2020) and strong tetrachoric correlations of 0.90 to 0.94 against actual survey choices across various subgroups. Third, in complex closed-ended multi-variable surveys, the model faithfully mirrored human correlation structures across eleven demographic and attitudinal variables, showing a minimal mean difference in association metrics of just -0.026. Fourth, the model maintained high predictive fidelity even for the 2020 election, which occurred after the model's 2019 training cutoff date.

These findings suggest language models capture deep structural associations between socio-demographic contexts and human attitudes, rather than just superficial text styles. This capability provides a cost-effective method to pilot experimental designs, test survey question wording, triage potential confounds, and refine theories before deploying expensive field research. However, the high fidelity also poses serious ethical and security risks, as actors could leverage targeted simulation for automated disinformation, targeted manipulation, or fraudulent campaigns.

Organizations and researchers should consider adopting silicon sampling as an exploratory pre-testing tool prior to human deployments, which can drastically cut preliminary research costs. Before strong decisions or public policies are based solely on silicon samples, researchers must establish algorithmic fidelity within each specific domain of study. Future work must extend evaluations beyond U.S. political opinion, optimize prompt templates, and develop community-wide ethical safeguards and oversight frameworks to detect and mitigate potential abuse.

Confidence in these findings is high for U.S. political attitudes and aggregate subgroup distributions within the examined surveys. However, readers should exercise caution because the model does not predict individual human responses deterministically. Performance weakens significantly among politically unaligned individuals (such as pure political independents), and high missing data rates or prompt non-compliance can occur depending on temperature settings and question formatting.

Cover for Out of One, Many: Using Language Models to Simulate Human Samples

Abstract

We propose and explore the possibility that language models can be studied as effective proxies for specific human sub-populations in social science research. Practical and research applications of artificial intelligence tools have sometimes been limited by problematic biases (such as racism or sexism), which are often treated as uniform properties of the models. We show that the "algorithmic bias" within one such tool -- the GPT-3 language model -- is instead both fine-grained and demographically correlated, meaning that proper conditioning will cause it to accurately emulate response distributions from a wide variety of human subgroups. We term this property "algorithmic fidelity" and explore its extent in GPT-3. We create "silicon samples" by conditioning the model on thousands of socio-demographic backstories from real human participants in multiple large surveys conducted in the United States. We then compare the silicon and human samples to demonstrate that the information contained in GPT-3 goes far beyond surface similarity. It is nuanced, multifaceted, and reflects the complex interplay between ideas, attitudes, and socio-cultural context that characterize human attitudes. We suggest that language models with sufficient algorithmic fidelity thus constitute a novel and powerful tool to advance understanding of humans and society across a variety of disciplines.

Table of Contents

  • 1 Introduction
  • 2 The GPT-3 Language Model
  • 3 Algorithmic Fidelity
  • 4 Silicon Sampling: Correcting Skewed Marginals
  • 5 Study 1: Free-form Partisan Text
  • 6 Study 2: Vote Prediction
  • 7 Study 3: Closed-ended Questions and Complex Correlations in Human Data
  • 8 Where do we go from here?
  • 9 Discussion
  • A General details on GPT-3 usage
  • B Details on Study 1
  • B.1 Details on Human and GPT-3 samples
  • B.2 Lucid survey design
  • B.3 Lucid results analysis
  • C Details on Study 2
  • C.1 Data generation
  • C.2 Data analysis
  • C.3 Ablation analysis
  • C.4 Model comparison
  • D Details on Study 3
  • D.1 Data generation
  • D.2 Data analysis
  • D.2.1 Missing Data
  • D.2.2 Descriptive Statistics
  • D.3 Alternative Specifications
  • D.3.1 Completely Synthetic Data
  • D.3.2 GPT-3 Temperature Variation
  • E Cost Analysis
  • References

Knowls

  1. Knowl 1 — Concept and Definition of Algorithmic Fidelity

    definition

    Algorithmic fidelity is defined as the degree to which the complex patterns of relationships between ideas, attitudes, and socio-cultural contexts within a language model accurately mirror those observed within a range of human sub-populations.

    The foundational premise of algorithmic fidelity is that a generative language model does not produce text from a single monolithic distribution. Rather, its output space represents a mixture of diverse conditional distributions corresponding to varied human socio-cultural experiences. By systematically providing conditioning context (such as demographic, attitudinal, and experiential background information), the model can be steered to sample from specific sub-distributions that correlate with the response patterns of distinct human groups and their intersections.

  2. Knowl 2 — Four Criteria for Assessing Algorithmic Fidelity

    definition

    To evaluate whether a generative language model possesses sufficient algorithmic fidelity to serve as a surrogate for human survey respondents in social science research, four criteria are established:

    1. Criterion 1 (Social Science Turing Test): Text responses generated by the model are indistinguishable from parallel texts produced by human respondents.
    2. Criterion 2 (Backward Continuity): Generated outputs are consistent with the input conditioning context such that human observers reading the outputs can accurately infer key socio-demographic and attitudinal attributes of the simulated individual.
    3. Criterion 3 (Forward Continuity): Generated outputs proceed naturally and plausibly from the provided conditioning context, reliably preserving the designated tone, format, and topical content.
    4. Criterion 4 (Pattern Correspondence): Generated outputs replicate the underlying inter-relationships and correlations between demographics, attitudes, and reported behaviors observed in real human reference populations.
  3. Knowl 3 — Silicon Sampling Framework for Correcting Skewed Marginal Distributions

    model/method

    Large language models trained on massive internet corpora capture joint distributions over responses VV (such as voting choices) and demographic backstories BLLMB_{\text{LLM}}, written as: P(V,BLLM)=P(V∣BLLM)P(BLLM)P(V, B_{\text{LLM}}) = P(V \mid B_{\text{LLM}}) P(B_{\text{LLM}})

    Because internet demographic distributions P(BLLM)P(B_{\text{LLM}}) are heavily unrepresentative of real-world target populations (such as eligible voters P(BTrue)P(B_{\text{True}})), unconditioned marginal estimates P(V)=∫BP(V,BLLM)dBP(V) = \int_B P(V, B_{\text{LLM}}) dB suffer from severe marginal skew, analogous to Simpson's Paradox.

    Silicon sampling corrects this skew by conditioning the language model on backstories drawn from a known, representative sample distribution P(Btarget)P(B_{\text{target}}) (such as the American National Election Studies, ANES). The marginal distribution is estimated by integrating over the representative sample: P(V)=∫BP(V∣B)P(Btarget)dBP(V) = \int_B P(V \mid B) P(B_{\text{target}}) dB

    Provided the conditional probability P(V∣B)P(V \mid B) demonstrates high algorithmic fidelity, silicon sampling enables accurate simulation of arbitrary target populations and sub-populations.

  4. Knowl 4 — Backstory Conditioning via First-Person Narratives and Virtual Interviews

    model/method

    Silicon respondents are instantiated from real human survey records using two structured conditioning methods:

    1. First-Person Declarative Backstories: Demographic and attitudinal survey responses are translated into continuous first-person statements concatenated in a fixed sequence (e.g., ideology, 7-point partisanship, race, gender, income, age). Continuous variables are mapped to categorical descriptors (e.g., ages 18–24 map to 'young', 25–39 to 'middle-aged', 40–60 to 'old', 61+ to 'very old'; annual household income under $15k maps to 'very poor', $15k–$50k to 'poor', $50k–$150k to 'middle-class', $150k+ to 'upper-class'). The conditioning text concludes with a prompt designed to elicit specific target responses.

    2. Virtual Interview Dialogues: Multi-item surveys are simulated as two-party transcripts between an Interviewer: and Me:. The interviewer presents exact survey questions with response options, and the simulated persona provides the true human responses for k−1k-1 items. The kk-th item is left open for the language model to complete zero-shot, eliciting closed-ended categorical choices with minimal extraneous generation.

  5. Knowl 5 — Human Indistinguishability and Partisan Evaluation of Silicon Text Generations

    empirical result

    In a replication of the 'Pigeonholing Partisans' study, GPT-3 (175B Davinci, temperature 0.7) was conditioned on 2,107 human demographic backstories to generate 4-word descriptions of Democrats and Republicans, yielding 4,083 unique silicon lists alongside 3,592 human lists. A sample of 2,873 human evaluators on Lucid assessed the lists across content and authenticity dimensions:

    • Turing Test Performance (Criterion 1): Evaluators classified 61.7% of human lists and 61.2% of GPT-3 lists as human-generated (p=0.44p = 0.44, two-tailed), demonstrating that evaluators could not distinguish between the two sources.
    • Content Alignment: Evaluators judged human and GPT-3 lists as having highly similar content profiles in terms of positivity (52.8% human vs. 51.5% GPT-3), extremity (39.8% human vs. 41.0% GPT-3), and prominence of personality traits (72.3% human vs. 66.5% GPT-3), mirroring the specific qualitative dominance of traits over issues and groups observed in real human data.
    • Backward Continuity (Criterion 2): Human evaluators correctly guessed the author's political party significantly above chance (33.3%) for both human-authored lists (60.1%) and GPT-3-authored lists (52.8%, p<0.001p < 0.001).
  6. Knowl 6 — Subgroup Presidential Vote Choice Prediction Across Time

    empirical result

    GPT-3 was evaluated on predicting presidential vote choice across the 2012, 2016, and 2020 American National Election Studies (ANES) by computing the conditional log-probabilities of candidate token sets (e.g., Romney vs. Obama in 2012, Trump vs. Clinton in 2016, Trump vs. Biden in 2020) given 10-variable demographic and political backstories.

    • Overall Sample Accuracy: Across all respondents, the tetrachoric correlation between human self-reports and dichotomized GPT-3 predictions (>0.50>0.50 Republican) was 0.90 in 2012, 0.92 in 2016, and 0.94 in 2020. Proportion agreement was 0.85 in 2012, 0.87 in 2016, and 0.89 in 2020.
    • Subgroup Pattern Correspondence (Criterion 4): Tetrachoric correlations exceeded 0.90 for the majority of demographic and political subgroups across all three cycles (e.g., strong partisans: 0.99–1.00; individuals with high political interest: 0.95–0.97; churchgoers: 0.91–0.94; Hispanics: 0.86–0.93).
    • Pure Independents Anomaly: Pure political independents were the sole subgroup where GPT-3 failed to align with human data, showing weak tetrachoric correlation (0.31 in 2012, 0.41 in 2016, and 0.02 in 2020), matching political science findings that pure independents exhibit low ideological consistency and volatile voting choices.
    • Temporal Generalization: Algorithmic fidelity remained equally high on the 2020 ANES dataset despite the GPT-3 pretraining corpus ending in 2019.
  7. Knowl 7 — Presidential Vote Choice Correspondence Metrics by Demographic Subgroup

    data/table

    The following table compares the correspondence between GPT-3 predicted vote choice (dichotomized at 0.50 probability) and ANES human vote choice across 2012, 2016, and 2020 for overall samples and key demographic/political subgroups using tetrachoric correlation ('Tetra') and raw proportion agreement ('Prop. Agree').

    Subgroup 2012 2016 2020
    Tetra Prop. Agree Tetra Prop. Agree Tetra Prop. Agree
    Whole sample 0.90 0.85 0.92 0.87 0.94 0.89
    Men 0.90 0.85 0.93 0.88 0.95 0.88
    Women 0.91 0.86 0.92 0.86 0.94 0.90
    Strong partisans 0.99 0.97 1.00 0.97 1.00 0.97
    Weak partisans 0.73 0.74 0.71 0.74 0.84 0.82
    Leaners 0.90 0.85 0.93 0.87 0.95 0.89
    Independents 0.31 0.59 0.41 0.62 0.02 0.53
    Conservatives 0.84 0.84 0.88 0.86 0.91 0.89
    Moderates 0.65 0.77 0.76 0.78 0.71 0.77
    Liberals 0.81 0.95 0.73 0.95 0.86 0.97
    Whites 0.87 0.82 0.91 0.85 0.94 0.89
    Blacks 0.71 0.97 0.87 0.96 0.81 0.94
    Hispanics 0.86 0.86 0.93 0.90 0.88 0.83
    Attends church 0.91 0.86 0.93 0.88 0.94 0.88
    Doesn't attend church 0.88 0.85 0.90 0.85 0.93 0.90
    High political interest 0.95 0.90 0.97 0.93 0.97 0.92
    Low political interest 0.71 0.74 0.75 0.75 0.83 0.81
    Discusses politics 0.92 0.87 0.94 0.88 0.95 0.90
    Doesn't discuss politics 0.83 0.82 0.81 0.79 0.80 0.79
    18 to 30 years old 0.90 0.87 0.90 0.86 0.90 0.87
    31 to 45 years old 0.90 0.85 0.92 0.87 0.94 0.90
    46 to 60 years old 0.90 0.86 0.92 0.86 0.92 0.87
    Over 60 0.90 0.85 0.93 0.87 0.96 0.91

    The data demonstrate consistent pattern correspondence across three distinct national election cycles, across demographics, political engagement levels, and religious participation.

  8. Knowl 8 — Complex Multidimensional Association Alignment Across Survey Variables

    empirical result

    In a 12-variable simulation using the 2016 ANES dataset (covering gender, race/ethnicity, age, education, church attendance, patriotism toward the flag, discussing politics, political interest, ideology, party ID, voting turnout, and candidate vote choice), GPT-3 was queried to predict each variable conditioned on the other 11 via virtual interview transcripts across 1,782 complete cases.

    • Bivariate Association Matching: Cramer's VV was calculated for all pairwise combinations of variables within human data and compared to the corresponding Cramer's VV between human conditioning inputs and GPT-3 outputs.
    • Statistical Alignment: The mean difference between human-human Cramer's VV and human-GPT-3 Cramer's VV was −0.026-0.026 (SD=0.068SD = 0.068). Strong human relationships (e.g., Party ID and 2016 Vote: human V=0.48V = 0.48, GPT-3 V=0.37V = 0.37; Ideology and Party ID: human V=0.37V = 0.37, GPT-3 V=0.32V = 0.32) were reproduced with high strength, while weak human associations (e.g., Gender and Discussing Politics: human V=0.00V = 0.00, GPT-3 V=0.01V = 0.01) showed negligible association in GPT-3 data.
    • Full Synthetic Replication: Estimating Cramer's VV purely between synthetic GPT-3 vectors without using human ANES inputs yielded an equivalent pattern of inter-correlations.
  9. Knowl 9 — Backstory Component Ablation and Language Model Parameter Scaling

    empirical result

    Systematic ablation and model comparison on the 2016 ANES vote prediction task established how conditioning elements and model architecture determine simulation fidelity:

    • Backstory Ablation: No single demographic or political attribute accounted for the full predictive performance of GPT-3. In single-element tests, party ID alone yielded a proportion agreement of ≈0.82\approx 0.82, and ideology yielded ≈0.79\approx 0.79, while single demographic factors (such as church attendance or gender) yielded 0.42−0.550.42-0.55. When both party ID and ideology were removed simultaneously from the template, the combination of the remaining 8 demographic factors produced a proportion agreement of ≈0.70\approx 0.70, higher than any single demographic factor alone. This confirms that the model synthesizes multifaceted demographic profiles.
    • Model Parameter Scaling: Comparative testing across GPT-2, GPT-Neo, GPT-J, Jurassic, and GPT-3 showed that algorithmic fidelity increases with model parameter count. Notably, the 6B-parameter GPT-Neo model achieved vote prediction agreement rivaling that of the 175B-parameter GPT-3 model.
  10. Knowl 10 — Impact of Generation Temperature on Survey Association Alignment

    empirical result

    Varying the sampling temperature TT of GPT-3 (Davinci) in the multidimensional survey simulation (2016 ANES) demonstrated that mid-range temperatures provide optimal algorithmic fidelity:

    • Temperature T=0.7T = 0.7: Produced a mean error in Cramer's VV of −0.026-0.026 (SD=0.068SD = 0.068, range: [−0.241,0.168][-0.241, 0.168], N=1,782N = 1,782 complete cases).
    • Temperature T=1.0T = 1.0: Produced a mean error in Cramer's VV of −0.031-0.031 (SD=0.070SD = 0.070, range: [−0.250,0.119][-0.250, 0.119], N=1,022N = 1,022 complete cases).
    • Temperature T=0.001T = 0.001 (Greedy Sampling): Produced a substantially higher mean error in Cramer's VV of +0.059+0.059 (SD=0.141SD = 0.141, range: [−0.123,0.700][-0.123, 0.700], N=2,518N = 2,518).

    At T=0.001T = 0.001, deterministic mode collapse occurred: GPT-3 predicted 100% of respondents as white, entirely eliminating demographic variation on race and systematically overstating correlations between human inputs and model outputs.

Coverage note — No substantial contributed material was omitted; monetary cost details and auxiliary OLS control tables from the appendix were summarized within the respective empirical knowls.

References

  1. 1.Pedro Rodriguez and Arthur Spirling. Word embeddings: What works, what doesn’t, and how to tell the difference for applied research. 2021.
  2. 2.Kenneth Benoit, Kevin Munger, and Arthur Spirling. Measuring and explaining political sophistication through textual complexity. American Journal of Political Science, 63(2):491–508, 2019.
  3. 3.Pablo Barberá, Amber E Boydstun, Suzanna Linn, Ryan McMahon, and Jonathan Nagler. Automated text classification of news articles: A practical guide. Political Analysis, 29(1):19–42, 2021.
  4. 4.Ludovic Rheault and Christopher Cochrane. Word embeddings for the analysis of ideological placement in parliamentary corpora. Political Analysis, 28(1):112–133, 2020.
  5. 5.Justin Grimmer, Margaret E Roberts, and Brandon M Stewart. Machine learning for social science: An agnostic approach. Annual Review of Political Science, 24:395–419, 2021.
  6. 6.Kevin T Greene, Baekkwan Park, and Michael Colaresi. Machine learning human rights and wrongs: How the successes and failures of supervised learning algorithms can inform the debate about information effects. Political Analysis, 27(2):223–230, 2019.
  7. 7.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsupervised multitask learners. OpenAI Blog, 1(8):9, 2019.
  8. 8.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67, 2020. URL http://jmlr.org/papers/v21/20-074.html.
  9. 9.Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. arXiv:2005.14165, 2020.
  10. 10.Daniel Adiwardana, Minh-Thang Luong, David R So, Jamie Hall, Noah Fiedel, Romal Thoppilan, Zi Yang, Apoorv Kulshreshtha, Gaurav Nemade, Yifeng Lu, et al. Towards a human-like open-domain chatbot. arXiv preprint arXiv:2001.09977, 2020.
  11. 11.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsupervised multitask learners. OpenAI Blog, 1(8):9, 2019.
  12. 12.Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov. Transformer-xl: Attentive language models beyond a fixed-length context. arXiv preprint arXiv:1901.02860, 2019.
  13. 13.Trishan Panch, Heather Mattie, and Rifat Atun. Artificial intelligence and algorithmic bias: Implications for health systems. Journal of Global Health, 9(2):010318, December 2019. ISSN 2047-2978, 2047-2986. doi: 10.7189/jogh.09.020318.
  14. 14.Sandra G. Mayson. Bias in, Bias out. Yale Law Journal, 128:2218, 2018.
  15. 15.Solon Barocas and Andrew D. Selbst. Big Data’s Disparate Impact. California Law Review, 104:671, 2016.
  16. 16.See https://electionstudies.org/about-us/.
  17. 17.Jacob E. Rothschild, Adam J. Howat, Richard M. Shafranek, and Ethan C. Busby. Pigeonholing partisans: Stereotypes of party supporters and partisan polarization. Political Behavior, 41(2):423–443, 2019.
  18. 18.Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, pages 610–623, 2021.
  19. 19.Gary Marcus. The next decade in ai: four steps towards robust artificial intelligence. arXiv preprint arXiv:2002.06177, 2020.
  20. 20.Nikhil Garg, Londa Schiebinger, Dan Jurafsky, and James Zou. Word embeddings quantify 100 years of gender and ethnic stereotypes. Proceedings of the National Academy of Sciences, 115(16):E3635–E3644, 2018.
  21. 21.Aylin Caliskan, Joanna J Bryson, and Arvind Narayanan. Semantics derived automatically from language corpora contain human-like biases. Science, 356(6334):183–186, 2017.
  22. 22.Bernard Berelson, Paul F. Lazarsfeld, and William N. McPhee. Voting: a study of opinion formation in a presidential campaign. University of Chicago Press, 1954.
  23. 23.Angus Campbell, Philip E. Converse, Warren E. Miller, and Donald E. Stokes. The American Voter. University of Chicago Press, 1960.
  24. 24.Vincent L. Hutchings and Nicholas A. Valentino. The centrality of race in american politics. Annual Review of Political Science, 7(1):383–408, 2004.
  25. 25.Nancy Burns and Katherine Gallagher. Public opinion on gender issues: The politics of equity and roles. Annual Review of Political Science, 13(1):425–443, 2010.
  26. 26.James N. Druckman and Arthur Lupia. Preference change in competitive political environments. Annual Review of Political Science, 19(1):13–31, 2016.
  27. 27.Katherine Cramer. Understanding the role of racism in contemporary us public opinion. Annual Review of Political Science, 23(1):153–169, 2020.
  28. 28.Edward H. Simpson. The interpretation of interaction in contingency tables. Journal of the Royal Statistical Society, Ser. B, 13:238–241, 1951.
  29. 29.Shanto Iyengar, Gaurav Sood, and Yphtach Lelkes. Affect, not ideology a social identity perspective on polarization. Public Opinion Quarterly, 76(3):405–431, 2012.
  30. 30.Lilliana Mason. Uncivil Agreement. University of Chicago Press, 2018.
  31. 31.Alexander Coppock and Oliver A. McClellan. Validating the demographic, political, psychological, and experimental results obtained from a new source of online survey respondents. Research & Politics, 6(1):1–14, 2019.
  32. 32.Janet M. Box-Steffensmeier, Suzanna De Boef, and Tse min Lin. The dynamics of the partisan gender gap. American Political Science Review, 98(3):515–528, 2004.
  33. 33.Katherine Tate. From Protest to Politics: The New Black Voters in American Elections. Harvard University Press, 1994.
  34. 34.Katherine J. Cramer. The politics of resentment: Rural consciousness in Wisconsin and the rise of Scott Walker. University of Chicago Press, 2016.
  35. 35.Ashley Jardina. White Identity Politics. Cambridge University Press, 2019.
  36. 36.Bruce E. Keith, David B. Magleby, Candice J. Nelson, Elizabeth Orr, and Mark C. Westyle. The myth of the independent voter. University of California Press, 1992.
  37. 37.David B. Magleby, Candice J. Nelson, and Mark C. Westlye. The myth of the independent voter revisited. In Paul M. Sniderman and Benjamin Highton, editors, Facing the challenge of democracy: Explorations in the analysis of public opinion and political participation, pages 238–266. Princeton University Press, 2011.
  38. 38.Samara Klar and Yanna Krupnikov. Independent Politics: How American Disdain for Parties Leads to Political Inaction. Cambridge University Press, 2016.
  39. 39.Harald Cramér. Mathematical Methods of Statistics. Princeton University Press, 1946.
  40. 40.Ronald S Ross. Guide for conducting risk assessments (nist sp-800-30rev1). The National Institute of Standards and Technology (NIST), Gaithersburg, 2012.
  41. 41.Matthew J. Salganik. Bit by Bit: Social Research in the Digital Age. Princeton University Press, Princeton, NJ, open review edition edition, 2017.

Citation

MLA
Argyle, L. P., et al. “Out of One, Many: Using Language Models to Simulate Human Samples”. Political Analysis, vol. 31, no. 3, 2023, pp. 337–51, https://doi.org/10.1017/pan.2023.2.
APA
Argyle, L. P., Busby, E. C., Fulda, N., Gubler, J. R., Rytting, C., & Wingate, D. (2023). Out of One, Many: Using Language Models to Simulate Human Samples. Political Analysis, 31(3), 337–351. https://doi.org/10.1017/pan.2023.2
Chicago
Argyle, L. P., E. C. Busby, N. Fulda, J. R. Gubler, C. Rytting, and D. Wingate. 2023. “Out of One, Many: Using Language Models to Simulate Human Samples”. Political Analysis 31 (3): 337–51. https://doi.org/10.1017/pan.2023.2.
Harvard
Argyle, L.P. et al. (2023) “Out of One, Many: Using Language Models to Simulate Human Samples”, Political Analysis, 31(3), pp. 337–351. Available at: https://doi.org/10.1017/pan.2023.2.
Vancouver
1. Argyle LP, Busby EC, Fulda N, Gubler JR, Rytting C, Wingate D (2023) Out of One, Many: Using Language Models to Simulate Human Samples. Political Analysis 31:337–351

BibTeX

@article{Argyle_2023, title={Out of One, Many: Using Language Models to Simulate Human Samples}, volume={31}, ISSN={1476-4989}, url={http://dx.doi.org/10.1017/pan.2023.2}, DOI={10.1017/pan.2023.2}, number={3}, journal={Political Analysis}, publisher={Cambridge University Press (CUP)}, author={Argyle, Lisa P. and Busby, Ethan C. and Fulda, Nancy and Gubler, Joshua R. and Rytting, Christopher and Wingate, David}, year={2023}, month=Feb, pages={337–351} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by-sa/4.0/