Improving Diversity of Demographic Representation in Large Language Models via Collective-Critiques and Self-Voting

Preethi LahotiNicholas BlummXiao MaRaghavendra KotikalapudiSahitya PotluriQijun TanHansa SrinivasanBen PackerAhmad BeiramiAlex Beutel

article2023EMNLP60 citations

Introduces a collective-critique and self-voting prompting framework that enables large language models to self-correct demographic and cultural underrepresentation in open-ended text generation without requiring prompt tuning or handcrafted examples.

Listen

Large language models often rely on implicit default assumptions when answering open-ended or under-specified questions. Consequently, their outputs frequently display severe demographic homogenization, systematically under-representing or erasing diverse demographic groups. As foundational artificial intelligence models become integral to downstream business applications, addressing this lack of representation is a vital fairness and responsibility priority.

The article formalizes diversity of demographic representation in generative language models and introduces an in-context prompting technique called collective-critiques and self-voting to enhance people and cultural diversity without requiring model retraining, fine-tuning, or handcrafted examples.

The analysis evaluates language model generations across newly created datasets spanning 105 occupations as well as cultural topics such as literature, music, and travel. Diversity was measured automatically along gender and ethnicity dimensions using entropy (to capture distributional spread) and max-gap (to capture disparities between most- and least-represented groups), complemented by side-by-side human evaluations of diversity and helpfulness. The authors benchmarked standard zero-shot, instruction-following, chain-of-thought, and constitutional approaches against the proposed technique using a 540-billion parameter model.

The key findings reveal that baseline language models produce nearly homogeneous responses, with approximately 99% of generated entities belonging to the same gender and 98% to the same ethnicity. While standard instruction prompting and step-by-step reasoning fail to improve diversity, the proposed zero-shot collective-critiques and self-voting method significantly enhances demographic balance, raising ethnicity entropy from 0.04 to 0.76 and gender entropy from 0.02 to 0.33 while boosting helpfulness from 26% to 93%. In human evaluations, raters preferred the zero-shot proposed method over the baseline in 89.5% of comparisons for diversity and 91.8% for helpfulness. Furthermore, the approach generalizes effectively to cultural topics and respects user-defined constraints, successfully knowing when not to diversify a restricted attribute while expanding representation across other dimensions.

These findings demonstrate that large language models inherently possess the conceptual reasoning necessary to identify and correct representation gaps in their own outputs. Aggregating multiple sampled critiques and selecting among candidate revisions via self-voting dramatically improves output quality and diversity in a single operational step, eliminating the need for expensive, brittle, prompt-tuned demonstration libraries.

Organizations developing or deploying conversational agents should integrate collective-critiquing and self-selection mechanisms to mitigate algorithmic bias and improve response quality. To manage increased computational latency and decoding costs during real-time inference, engineering teams should consider using the proposed approach offline to generate high-diversity synthetic training data for fine-tuning smaller, more cost-effective production models.

Confidence in these findings is strong for large English-language foundation models, supported by statistically significant alignment between automated metrics and human evaluations. Readers should note that current evaluations rely on template-based prompts and knowledge-graph attribute classifications, meaning further validation is recommended before generalizing to multilingual environments, highly specialized domains, or resource-constrained small models.

arXiv: 2310.16523
Cover for Improving Diversity of Demographic Representation in Large Language Models via Collective-Critiques and Self-Voting

Abstract

A crucial challenge for generative large language models (LLMs) is diversity: when a user's prompt is under-specified, models may follow implicit assumptions while generating a response, which may result in homogenization of the responses, as well as certain demographic groups being under-represented or even erased from the generated responses. In this paper, we formalize diversity of representation in generative LLMs. We present evaluation datasets and propose metrics to measure diversity in generated responses along people and culture axes. We find that LLMs understand the notion of diversity, and that they can reason and critique their own responses for that goal. This finding motivated a new prompting technique called collective-critique and self-voting (CCSV) to self-improve people diversity of LLMs by tapping into its diversity reasoning capabilities, without relying on handcrafted examples or prompt tuning. Extensive empirical experiments with both human and automated evaluations show that our proposed approach is effective at improving people and culture diversity, and outperforms all baseline methods by a large margin.

Table of Contents

  • 1 Introduction
  • 2 Background & Related Work
  • 3 Mitigation Design
  • 3.1 Method
  • 3.2 Design Rationale
  • 4 Evaluation Design
  • 4.1 Evaluation Datasets
  • 4.2 Automated Eval & Metrics
  • 4.3 Human Eval and Metrics
  • 5 Experiments
  • 5.1 Results
  • 6 Analysis, Insights & Ablations
  • 6.1 Robustness of Diversity Methods
  • 6.2 Generalization to Cultural Diversity Task
  • 6.3 Ablation study
  • 7 Conclusion
  • Limitations & Broader Impact
  • Ethics Statement
  • Acknowledgements
  • References
  • A Appendix
  • A.1 Illustrative Result
  • A.2 Implementation of Evaluation Datasets.
  • A.2.1 People-Diversity Dataset
  • A.2.2 Culture-Diversity Dataset
  • A.3 Implementation of Baselines and Methods
  • A.3.1 Visualizations of Baseline and Methods
  • A.4 Few-shot Prompts
  • A.5 Human Evaluation Supplement Materials
  • A.6 Additional Results

Knowls

  1. Knowl 1 — Collective-critique and self-voting (CCSV)

    model/method

    CCSV improves an initial language-model answer through critique, revision, and selection, using the same model for every step. Given a user prompt, the model first generates an answer. It then samples five critiques that identify ways the answer lacks diversity and suggest improvements. The model receives all five critiques together and generates five revised answers addressing them. Finally, the candidate revisions are presented to the model, which judges which is most diverse while still answering the user’s question; the most-voted revision is returned. The process can be repeated using the selected answer as the new starting point, but the reported experiments use one round. The zero-shot form uses no hand-written examples or task-specific prompt tuning.

  2. Knowl 2 — Automated measures of representation diversity

    definition

    For an evaluated response yy, let pv(y)p_v(y) be the fraction of identified people entities assigned attribute value vv by an entity extractor and knowledge graph. For an attribute with value set AA, the response-level entropy is −∑v∈Apv(y)log⁡2pv(y)-\sum_{v\in A}p_v(y)\log_2 p_v(y); the reported entropy is averaged over evaluated responses. It is unnormalized, measured in bits, and ranges from 00 to log⁡2∣A∣\log_2|A|; higher values indicate more diverse representation. The complementary max-gap is the average, over responses, of the largest difference between any two attribute-value proportions: 1∣Y∣∑y∈Ymax⁡u,v∈A∣pu(y)−pv(y)∣\frac{1}{|Y|}\sum_{y\in Y}\max_{u,v\in A}|p_u(y)-p_v(y)|, where YY is the set of evaluated responses. It ranges from 00 to 11, with lower values indicating less disparity between the most- and least-represented values. The paper reports each metric separately for gender and ethnicity. Responses containing no people entities are assigned entropy 00 and max-gap 11; their fraction is reported separately as “Is helpful,” meaning the fraction of prompts for which the response contains people entities.

  3. Knowl 3 — Evaluation datasets for people and cultural diversity

    experimental setup

    The people-diversity dataset consists of hand-crafted people-seeking prompts about 105 professions drawn from Wikipedia’s occupation lists. Ten prompt templates are populated with profession names and, where applicable, adjectives such as “famous,” “influential,” and “popular.” The cultural-diversity dataset uses five hand-crafted prompt templates populated with cultural topics, including cities, countries, cuisines, books, languages, music, traditions, and tourist locations. These datasets evaluate diversity in open-ended generated responses rather than in ranked candidate lists. Both datasets rely on hand-crafted templates and populations, so their coverage is limited and not comprehensive.

  4. Knowl 4 — Experimental model and inference conditions

    experimental setup

    Experiments use the instruction-tuned Flan-PaLM 540B model, initialized as a conversational agent with a preamble asking it to answer users fluently and comprehensively. All methods are implemented by inference-time prompting. Decoding uses top-kk sampling at temperature 0.70.7 with a maximum of 1,024 decode steps; the paper does not specify the value of kk. CCSV samples five outputs at its critique and revision stages. Its reported results use one critique–revision–voting round. Comparisons include ordinary zero-shot prompting, zero-shot diversity-instruction and chain-of-thought variants, and five-shot standard, chain-of-thought, and Constitutional AI prompting; five-shot methods use hand-crafted in-context examples.

  5. Knowl 5 — Automated people-diversity results

    empirical result

    On the people-diversity benchmark, Flan-PaLM 540B’s standard zero-shot answers have very low entropy and high max-gap. The five-shot CCSV variant has the strongest overall automated results in the comparison, while zero-shot CCSV substantially improves on ordinary zero-shot prompting and leads the zero-shot alternatives. Entropy is higher-is-more-diverse, max-gap is lower-is-more-diverse, and “Is helpful” is the fraction of responses containing people entities. Entropy values are unnormalized, so gender and ethnicity entropy values should not be compared directly.

    MethodEntropy (ethnicity) ↑Entropy (gender) ↑Gap (ethnicity) ↓Gap (gender) ↓Is helpful
    Baseline0.040.020.980.990.26
    0-shot IF0.100.030.960.990.24
    0-shot CoT0.050.030.980.990.34
    Standard 5-shot0.770.250.730.910.80
    5-shot CoT0.600.270.790.890.86
    5-shot CAI0.380.230.860.910.56
    0-shot CCSV0.760.330.720.890.93
    5-shot CCSV1.080.490.640.830.96

    The simple zero-shot diversity instruction (IF) and “Let’s think step by step” (CoT) produce little change from the baseline. Zero-shot CCSV achieves entropy gains of more than 72 percentage points for gender and 31 percentage points for ethnicity over baseline, and five-shot CCSV reaches 1.08 ethnicity entropy and 0.49 gender entropy.

  6. Knowl 6 — Human evaluation of people-diversity responses

    empirical result

    Human raters compared each method’s response with the same baseline response for a prompt, judging diversity and helpfulness side by side. Three raters evaluated each task. Each judgment used a seven-option scale mapped to scores from −1.5-1.5 to 1.51.5 in increments of 0.50.5; positive mean scores favor the method being compared with the baseline. Both CCSV variants receive the highest scores among methods in their respective prompting groups. The automated diversity measures were also reported to correlate with human judgments at p<.05p<.05 across the human-evaluation trials.

    Method compared with baselineDiversity SxS ↑Helpfulness SxS ↑
    0-shot IF0.0290.027
    0-shot CoT0.0660.060
    Standard 5-shot0.5880.591
    5-shot CoT0.5760.529
    5-shot CAI0.4550.422
    0-shot CCSV0.8370.892
    5-shot CCSV0.7080.663

    For zero-shot CCSV, 89.50% of diversity ratings and 91.83% of helpfulness ratings preferred CCSV to baseline; the corresponding preferences for five-shot CCSV were 92.67% and 93.50%.

  7. Knowl 7 — Diversity under user-specified group constraints

    empirical result

    The paper tests whether methods can respect a requested demographic constraint while diversifying along another attribute. For example, prompts can request people of a particular gender, while the response is assessed for ethnicity diversity. The comparison tracks the fraction of responses satisfying the requested group constraint and entropy on the other attribute. CCSV variants occupy the strongest region of this trade-off: they largely satisfy the requested constraint while attaining the highest reported diversity on the unconstrained attribute. Ordinary zero-shot baselines satisfy the constraint but show low diversity on the other attribute. Standard five-shot prompting is comparatively brittle, satisfying the constraint less reliably. The paper also reports the same qualitative pattern for prompts requesting African-American people while measuring gender entropy.

  8. Knowl 8 — Generalization to cultural-diversity prompts

    empirical result

    The authors applied the people-diversity methods to hand-crafted cultural prompts without changing their prompts, instructions, or few-shot examples. Because the automated metrics did not generalize to this task, results are based on human side-by-side scores against the standard-prompting baseline. Both CCSV variants outperform the other tested methods on diversity and helpfulness; the five-shot CCSV comparison has the largest scores. Positive scores indicate preference for the listed method over baseline on the −1.5-1.5 to 1.51.5 human rating scale.

    Method compared with baselineDiversity SxS ↑Helpfulness SxS ↑
    0-shot IF0.0320.012
    0-shot CoT-0.0210.001
    Standard 5-shot0.0770.056
    5-shot CoT0.0270.049
    5-shot CAI0.3560.453
    0-shot CCSV0.4730.760
    5-shot CCSV1.0870.941

    The five-shot standard and chain-of-thought methods show little improvement in this setting, whereas five-shot CAI improves over those baselines but remains below CCSV.

  9. Knowl 9 — Effects of collective critiques and self-voting

    empirical result

    An ablation compares greedy critique–revision, collective critiques without self-voting, and collective critiques with self-voting. Each component improves the reported diversity measures relative to the simpler variants. Aggregating multiple critiques produces the largest gain; asking the model to select among multiple revised drafts adds a smaller but notable gain. The advantage of collective critiques is especially useful after a single critique–revision round, and the gap between variants narrows with additional rounds. Thus, aggregating critique outputs can achieve gains with fewer recursive rounds than greedy critique–revision.

  10. Knowl 10 — Scope and limitations of the evaluation and method

    limitation

    The evaluation uses hand-crafted English templates and covers only one language-model family, Flan-PaLM; it therefore does not establish performance for other models, languages, or the full range of diversity in language generation. The automated measures depend on entity extraction and knowledge graphs, whose demographic labels may have incomplete coverage or be incorrect or outdated. Automated evaluation is limited to gender and ethnicity, and the available gender categories include “other,” which the authors note is not an adequate representation of genders beyond male and female. Human raters’ demographic diversity and how rater backgrounds shape diversity judgments remain concerns. CCSV also assumes that a model has sufficiently diverse knowledge to critique its own responses; the authors caution that it may be less effective for smaller models with weaker reasoning and critique abilities. Its multiple critique, revision, and voting steps increase inference cost, although the reported experiments show gains after one round.

Coverage note — The paper’s illustrative example outputs and full few-shot prompt exemplars are omitted because they instantiate the method rather than add a distinct result; detailed prompt-template inventories are summarized at the dataset-design level.

References

  1. 1.Pranjal Aggarwal, Aman Madaan, Yiming Yang, et al. 2023. Let’s sample step by step: Adaptive-consistency for efficient reasoning with llms. arXiv preprint arXiv:2305.11860.
  2. 2.Rakesh Agrawal, Sreenivas Gollapudi, Alan Halverson, and Samuel Ieong. 2009. Diversifying search results. In Proceedings of the Second ACM International Conference on Web Search and Data Mining, WSDM ’09, page 5–14, New York, NY, USA. Association for Computing Machinery.
  3. 3.Rohan Anil, Andrew M Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al. 2023. Palm 2 technical report. arXiv preprint arXiv:2305.10403.
  4. 4.Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, Ben Mann, and Jared Kaplan. 2022a. Training a helpful and harmless assistant with reinforcement learning from human feedback.
  5. 5.Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al. 2022b. Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073.
  6. 6.Rishi Bommasani, Kathleen A Creel, Ananya Kumar, Dan Jurafsky, and Percy S Liang. 2022. Picking on the same person: Does algorithmic monoculture lead to outcome homogenization? Advances in Neural Information Processing Systems, 35:3663–3678.
  7. 7.Keith Bradley and Barry Smyth. 2001. Improving recommendation diversity.
  8. 8.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners.
  9. 9.Jaime Carbonell and Jade Goldstein. 1998. The use of mmr, diversity-based reranking for reordering documents and producing summaries. In Proceedings of the 21st Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR ’98, page 335–336, New York, NY, USA. Association for Computing Machinery.
  10. 10.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.
  11. 11.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. 2022. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416.
  12. 12.Ameet Deshpande, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, and Karthik Narasimhan. 2023. Toxicity in chatgpt: Analyzing persona-assigned language models. arXiv preprint arXiv:2304.05335.
  13. 13.Eve Fleisig, Rediet Abebe, and Dan Klein. 2023. When the majority is wrong: Leveraging annotator disagreement for subjective tasks. arXiv preprint arXiv:2305.06626.
  14. 14.Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A Smith. 2020. Realtoxiciprompts: Evaluating neural toxic degeneration in language models. arXiv preprint arXiv:2009.11462.
  15. 15.Sahin Cem Geyik, Stuart Ambler, and Krishnaram Kenthapadi. 2019. Fairness-aware ranking in search and recommendation systems with application to linkedin talent search. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. ACM.
  16. 16.Lou Jost. 2006. Entropy and diversity. Oikos, 113(2):363–375.
  17. 17.Matthew Kay, Cynthia Matuszek, and Sean A Munson. 2015. Unequal representation and gender stereotypes in image search results for occupations. In Proceedings of the 33rd annual acm conference on human factors in computing systems, pages 3819–3828.
  18. 18.Hannah Rose Kirk, Yennie Jun, Filippo Volpin, Haider Iqbal, Elias Benussi, Frederic Dreyer, Aleksandar Shtedritski, and Yuki Asano. 2021. Bias out-of-the-box: An empirical analysis of intersectional occupational biases in popular generative language models. Advances in neural information processing systems, 34:2611–2624.
  19. 19.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. arXiv preprint arXiv:2205.11916.
  20. 20.Matevž Kunaver and Tomaž Požrl. 2017. Diversity in recommender systems – a survey. Knowledge-Based Systems, 123:154–162.
  21. 21.Paul Pu Liang, Chiyu Wu, Louis-Philippe Morency, and Ruslan Salakhutdinov. 2021. Towards understanding and mitigating social biases in language models. In International Conference on Machine Learning, pages 6565–6576. PMLR.
  22. 22.Lists of occupations. 2023. Lists of occupations — Wikipedia, the free encyclopedia. [Online; accessed 09-June-2023].
  23. 23.Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al. 2023. Self-refine: Iterative refinement with self-feedback. arXiv preprint arXiv:2303.17651.
  24. 24.Lucie Charlotte Magister, Jonathan Mallinson, Jakub Adamek, Eric Malmi, and Aliaksei Severyn. 2022. Teaching small language models to reason. arXiv preprint arXiv:2212.08410.
  25. 25.Moin Nadeem, Anna Bethke, and Siva Reddy. 2020. Stereoset: Measuring stereotypical bias in pretrained language models. arXiv preprint arXiv:2004.09456.
  26. 26.OpenAI. 2023. GPT-4 Technical Report.
  27. 27.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744.
  28. 28.Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel R Bowman. 2021. Bbq: A hand-built bias benchmark for question answering. arXiv preprint arXiv:2110.08193.
  29. 29.Vinodkumar Prabhakaran, Rida Qadri, and Ben Hutchinson. 2022. Cultural incongruencies in artificial intelligence. arXiv preprint arXiv:2211.13069.
  30. 30.Rachel Rudinger, Jason Naradowsky, Brian Leonard, and Benjamin Van Durme. 2018. Gender bias in coreference resolution. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), pages 8–14.
  31. 31.Timo Schick, Sahana Udupa, and Hinrich Schütze. 2021. Self-diagnosis and self-debiasing: A proposal for reducing corpus-based bias in nlp. Transactions of the Association for Computational Linguistics, 9:1408–1424.
  32. 32.Omar Shaikh, Hongxin Zhang, William Held, Michael Bernstein, and Diyi Yang. 2022. On second thought, let’s not think step by step! bias and toxicity in zero-shot reasoning. arXiv preprint arXiv:2212.08061.
  33. 33.Pedro Silva, Bhawna Juneja, Shloka Desai, Ashudeep Singh, and Nadia Fawaz. 2023. Representation online matters: Practical end-to-end diversification in search and recommender systems. arXiv preprint arXiv:2305.15534.
  34. 34.Eric Michael Smith, Melissa Hall, Melanie Kambadur, Eleonora Presani, and Adina Williams. 2022. “i’m sorry to hear that”: Finding new biases in language models with a holistic descriptor dataset. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 9180–9211.
  35. 35.Irene Solaiman, Miles Brundage, Jack Clark, Amanda Askell, Ariel Herbert-Voss, Jeff Wu, Alec Radford, Gretchen Krueger, Jong Wook Kim, Sarah Kreps, et al. 2019. Release strategies and the social impacts of language models. arXiv preprint arXiv:1908.09203.
  36. 36.Rui Wang, Hongru Wang, Fei Mi, Yi Chen, Ruifeng Xu, and Kam-Fai Wong. 2023a. Self-critique prompting with large language models for inductive instructions. arXiv preprint arXiv:2305.13733.
  37. 37.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2022. Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171.
  38. 38.Zihao Wang, Lin Gui, Jeffery Negrea, and Victor Veitch. 2023b. Concept algebra for text-controlled vision models. arXiv preprint arXiv:2302.03693.
  39. 39.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed H Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems.
  40. 40.Jieyu Zhao, Daniel Khashabi, Tushar Khot, Ashish Sabharwal, and Kai-Wei Chang. 2021. Ethical-advice taker: Do language models understand natural language interventions? In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 4158–4164.
  41. 41.Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. 2018. Gender bias in coreference resolution: Evaluation and debiasing methods. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), pages 15–20.
  42. 42.Xiaojin Zhu, Andrew Goldberg, Jurgen Van Gael, and David Andrzejewski. 2007. Improving diversity in ranking using absorbing random walks. In Human Language Technologies 2007: The Conference of the North American Chapter of the Association for Computational Linguistics; Proceedings of the Main Conference, pages 97–104, Rochester, New York. Association for Computational Linguistics.

Citation

MLA
Lahoti, P., et al. “Improving Diversity of Demographic Representation in Large Language Models via Collective-Critiques and Self-Voting”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 10383–405, https://doi.org/10.18653/v1/2023.emnlp-main.643.
APA
Lahoti, P., Blumm, N., Ma, X., Kotikalapudi, R., Potluri, S., Tan, Q., Srinivasan, H., Packer, B., Beirami, A., Beutel, A., & Chen, J. (2023). Improving Diversity of Demographic Representation in Large Language Models via Collective-Critiques and Self-Voting. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 10383–10405. https://doi.org/10.18653/v1/2023.emnlp-main.643
Chicago
Lahoti, P., N. Blumm, X. Ma, et al. 2023. “Improving Diversity of Demographic Representation in Large Language Models via Collective-Critiques and Self-Voting”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 10383–405. https://doi.org/10.18653/v1/2023.emnlp-main.643.
Harvard
Lahoti, P. et al. (2023) “Improving Diversity of Demographic Representation in Large Language Models via Collective-Critiques and Self-Voting”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 10383–10405. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.643.
Vancouver
1. Lahoti P, Blumm N, Ma X, et al (2023) Improving Diversity of Demographic Representation in Large Language Models via Collective-Critiques and Self-Voting. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 10383–10405

BibTeX

@inproceedings{lahoti-etal-2023-improving,
    title = "Improving Diversity of Demographic Representation in Large Language Models via Collective-Critiques and Self-Voting",
    author = "Lahoti, Preethi  and
      Blumm, Nicholas  and
      Ma, Xiao  and
      Kotikalapudi, Raghavendra  and
      Potluri, Sahitya  and
      Tan, Qijun  and
      Srinivasan, Hansa  and
      Packer, Ben  and
      Beirami, Ahmad  and
      Beutel, Alex  and
      Chen, Jilin",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.643/",
    doi = "10.18653/v1/2023.emnlp-main.643",
    pages = "10383--10405"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/