Improving Diversity of Demographic Representation in Large Language Models via Collective-Critiques and Self-Voting
Preethi LahotiNicholas BlummXiao MaRaghavendra KotikalapudiSahitya PotluriQijun TanHansa SrinivasanBen PackerAhmad BeiramiAlex Beutel
Introduces a collective-critique and self-voting prompting framework that enables large language models to self-correct demographic and cultural underrepresentation in open-ended text generation without requiring prompt tuning or handcrafted examples.
Large language models often rely on implicit default assumptions when answering open-ended or under-specified questions. Consequently, their outputs frequently display severe demographic homogenization, systematically under-representing or erasing diverse demographic groups. As foundational artificial intelligence models become integral to downstream business applications, addressing this lack of representation is a vital fairness and responsibility priority.
The article formalizes diversity of demographic representation in generative language models and introduces an in-context prompting technique called collective-critiques and self-voting to enhance people and cultural diversity without requiring model retraining, fine-tuning, or handcrafted examples.
The analysis evaluates language model generations across newly created datasets spanning 105 occupations as well as cultural topics such as literature, music, and travel. Diversity was measured automatically along gender and ethnicity dimensions using entropy (to capture distributional spread) and max-gap (to capture disparities between most- and least-represented groups), complemented by side-by-side human evaluations of diversity and helpfulness. The authors benchmarked standard zero-shot, instruction-following, chain-of-thought, and constitutional approaches against the proposed technique using a 540-billion parameter model.
The key findings reveal that baseline language models produce nearly homogeneous responses, with approximately 99% of generated entities belonging to the same gender and 98% to the same ethnicity. While standard instruction prompting and step-by-step reasoning fail to improve diversity, the proposed zero-shot collective-critiques and self-voting method significantly enhances demographic balance, raising ethnicity entropy from 0.04 to 0.76 and gender entropy from 0.02 to 0.33 while boosting helpfulness from 26% to 93%. In human evaluations, raters preferred the zero-shot proposed method over the baseline in 89.5% of comparisons for diversity and 91.8% for helpfulness. Furthermore, the approach generalizes effectively to cultural topics and respects user-defined constraints, successfully knowing when not to diversify a restricted attribute while expanding representation across other dimensions.
These findings demonstrate that large language models inherently possess the conceptual reasoning necessary to identify and correct representation gaps in their own outputs. Aggregating multiple sampled critiques and selecting among candidate revisions via self-voting dramatically improves output quality and diversity in a single operational step, eliminating the need for expensive, brittle, prompt-tuned demonstration libraries.
Organizations developing or deploying conversational agents should integrate collective-critiquing and self-selection mechanisms to mitigate algorithmic bias and improve response quality. To manage increased computational latency and decoding costs during real-time inference, engineering teams should consider using the proposed approach offline to generate high-diversity synthetic training data for fine-tuning smaller, more cost-effective production models.
Confidence in these findings is strong for large English-language foundation models, supported by statistically significant alignment between automated metrics and human evaluations. Readers should note that current evaluations rely on template-based prompts and knowledge-graph attribute classifications, meaning further validation is recommended before generalizing to multilingual environments, highly specialized domains, or resource-constrained small models.
- Paper: "I'm sorry to hear that": Finding New Biases in Language Models with a Holistic Descriptor Dataset, Eric Michael Smith et al. (2022). Its broad, community-informed demographic taxonomy and bias evaluation establish the measurement context for understanding how this paper detects and addresses under-representation.
- Paper: Prompt Programming for Large Language Models: Beyond the Few-Shot Paradigm, Laria Reynolds et al. (2021). Its account of zero-shot prompting as a way to elicit learned capabilities provides useful grounding for the source’s inference-time prompting intervention.
- Paper: Self-Generated Critiques Boost Reward Modeling for Language Models, Yue Yu 0009 et al. (2025). It extends the use of model-generated critiques by training them to improve reward modeling, carrying critique-based refinement into preference alignment.
- Paper: CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation, Pei Ke et al. (2024). It develops critique generation into an evaluation framework, extending the source’s use of critiques from revising outputs to assessing their quality.
