"Thinking" Fair and Slow: On the Efficacy of Structured Prompts for Debiasing Language Models

Shaz FurniturewalaSurgan JandialAbhinav JavaPragyan BanerjeeSimra ShahidSumit BhatiaKokil Jaidka

article2024EMNLP45 citations

Presents a System 2-inspired prompting framework that enables end users to significantly reduce social biases in closed-source and open-source language models without retraining weights or accessing internal probability distributions.

Listen

Large language models frequently inherit and perpetuate societal stereotypes present in their training data. While traditional bias-mitigation methods rely on model retraining, internal parameter adjustments, or access to output probabilities, many state-of-the-art models are deployed as closed systems accessible only through interfaces that restrict internal access. Even when open-source models are available, curating data and fine-tuning models is computationally expensive and risks degrading general task performance. Consequently, end-users requiring fairer text generation need black-box mitigation techniques that work exclusively through prompt design.

The article evaluates whether structured prompting strategies can mitigate bias in closed and open language models without internal modifications or loss in general performance. Specifically, it organizes and tests prompting methods across three structured categories inspired by human cognitive decision-making processes: direct prefix instructions, iterative self-refinement, and multi-step implication prompting that explicitly reasons about underlying stereotypes.

To assess these strategies, the authors evaluated four prominent open-access language models (GPT-J 6B, Mistral 7B, MPT-Instruct 7B, and Llama-2 13B) across established benchmarks measuring stereotypical bias across gender, race, religion, and profession, social regard differences, and completion toxicity. The evaluation compared single-step prefixing, multi-iteration self-refinement, and dynamic implication prompting against baseline models and existing internal modification techniques. Downstream question-answering benchmarks were also evaluated to verify whether general model capability was preserved.

The findings show that structured, reasoning-based prompting significantly reduces bias while matching or exceeding the performance of internal debiasing techniques. Implication Prompting proved the most effective across all benchmarks, reducing stereotype scores by an average of about 4% and bias on social perception metrics by roughly 27% compared to other prompting methods, while decreasing output toxicity by approximately 7%. Iterative Self-Refinement delivered strong improvements over single-step prefixes, though running more than one refinement step yielded diminishing returns, improving stereotype metrics by less than 0.5%. Across all methods, assigning an unbiased persona or role outperformed simple negative instructions, and smaller language models could effectively generate the required implications without degrading overall debiasing quality.

These results demonstrate that end-users can successfully eliminate bias in third-party and proprietary language models without costly fine-tuning or specialized infrastructure. Because these prompting pipelines preserve underlying language modeling capabilities—showing parity on factual and general question-answering tasks—organizations can adopt them without risking operational performance. When implementing prompt-based fairness controls, teams should prioritize dynamically generated implications or single-iteration role-based refinements rather than basic instruction prefixes or compute-intensive multiple refinement loops.

Decision-makers should note certain limitations: the empirical tests were bounded by resource constraints to models up to 13 billion parameters, and they depend on the model's pre-existing internal data containing sufficient context to navigate toward fairer outputs. Furthermore, structured prompts direct stochastic search rather than instill genuine moral reasoning. Nevertheless, the high statistical confidence across diverse benchmarks indicates that structured prompt frameworks are a practical and robust tool for end-user bias mitigation.

arXiv: 2405.10431

No sufficiently relevant recommendations were found.

Cover for "Thinking" Fair and Slow: On the Efficacy of Structured Prompts for Debiasing Language Models

Abstract

Existing debiasing techniques are typically training-based or require access to the model’s internals and output distributions, so they are inaccessible to end-users looking to adapt LLM outputs for their particular needs. In this study, we examine whether structured prompting techniques can offer opportunities for fair text generation. We evaluate a comprehensive end-user-focused iterative framework of debiasing that applies System 2 thinking processes for prompts to induce logical, reflective, and critical text generation, with single, multi-step, instruction, and role-based variants. By systematically evaluating many LLMs across many datasets and different prompting strategies, we show that the more complex System 2-based Implicative Prompts significantly improve over other techniques demonstrating lower mean bias in the outputs with competitive performance on the downstream tasks. Our work offers research directions for the design and the potential of end-user-focused evaluative frameworks for LLM use.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Prompting Framework
  • 3.1 Prefix Prompting (PP)
  • 3.2 Self Refinement (SR)
  • 3.3 Implication Prompting (IP)
  • 4 Models and Metrics
  • 5 Results and Discussion
  • 6 Ablations and Analysis
  • 7 Conclusion
  • 8 Limitations and Future Work
  • References
  • A Additional Details and Results
  • B Comparing prompting methods with the other debiasing methods
  • C Utilizing a Fixed Generic Implication
  • D Measuring Language Model’s Performance on downstream Question answering tasks

Knowls

  1. Knowl 1 — Implication prompting uses model-generated explanations to revise biased text

    algorithm

    Implication Prompting (IP) is a three-generation procedure for revising a language model output. Given a user prompt CC, the generation model MM first produces an initial text SS. An implication generator then analyzes SS using an instruction that asks it to identify the implication or stereotype in the text. Finally, MM receives the initial text, its generated implication, an instruction to avoid discrimination, and the original prompt, and produces the revised text. By default, the same model generates both the implication and the final answer; the authors also tested other implication generators.

    Input: User prompt C, generation model M, implication generator G, implication instruction I_impl, fairness instruction I_IP
    Output: Revised text S_debiased
    1. S <- M(C)
    2. C_impl <- concatenate(S, I_impl)
    3. S_impl <- G(C_impl)
    4. C_IP <- concatenate(S, S_impl, I_IP, C)
    5. S_debiased <- M(C_IP)
    6. return S_debiased

    The tested implication instructions were a direct request to generate the implication, a zero-shot chain-of-thought request to think step by step, and a few-shot instruction with examples identifying stereotypes. The final instruction asks the model to use the implication and original sentence to generate unbiased text.

  2. Knowl 2 — Implication prompting has the strongest reported aggregate debiasing results

    empirical result

    Across the evaluated models and prompting variants, the authors report that Implication Prompting (IP) performs better overall than Prefix Prompting (PP) and Self-Refinement (SR) on StereoSet, Regard, and toxicity. Averaged across IP variants and models, IP has a 4.05% lower StereoSet stereotype score and a 6.80% higher ICAT score than the other methods; it also shows a 26.85% improvement on Regard and a 6.98% decrease in mean toxicity. These are aggregate comparisons reported by the authors, not a claim that every IP variant wins on every individual model and metric. Their overall ordering is IP, then SR, then PP.

    For single-step SR compared with PP, the reported averages are a 6.85% lower stereotype score, an 11.65% higher ICAT score, and a 21.64% improvement on Regard, alongside a 1.11% increase in toxicity. Role-based PP also outperforms instruction-based PP in the authors’ aggregate comparison: 2.14% lower stereotype score, 5.08% higher ICAT, nearly 39.47% better Regard performance, and 4.34% lower toxicity.

  3. Knowl 3 — Self-Refinement conditions each revision on the previous output

    algorithm

    Self-Refinement (SR) repeatedly asks a language model to revise text while referring to its previous output. Given a user prompt CC, a refinement instruction or role prefix ISRI_{SR}, and an iteration count kk, the model produces an initial text and then generates kk revisions. Each revision is prompted with the previous output, the fairness instruction, and the original user prompt; the last revision is returned. The paper’s prose describes the initial generation as using Prefix Prompting, whereas its generic pseudocode initializes the first output with M(C)M(C); both descriptions agree that later steps use the prior output and original prompt.

    Input: User prompt C, refinement prefix I_SR, number of iterations k, language model M
    Output: Revised text S_debiased
    1. S_0 <- M(C)
    2. for i from 1 to k do
    3. C_SR <- concatenate(I_SR, S_(i-1), C)
    4. S_i <- M(C_SR)
    5. end for
    6. return S_k

    The refinement instruction can be phrased as a direct instruction to be unbiased or as a role asking the model to respond as an unbiased person. The experiments include k=1k=1 and k=2k=2 in the main comparisons and test additional iteration counts in an ablation.

  4. Knowl 4 — Prefix Prompting prepends an instruction or role to the user prompt

    model/method

    Prefix Prompting (PP) adapts a language model without changing its weights or requiring access to its internal probabilities. For a user prompt CC, the method forms a new prompt by prepending a debiasing prefix and then generates text from that combined prompt. The instruction variant tells the model to avoid discrimination based on sensitive attributes; the role variant asks it to adopt the persona of an unbiased person. In the study, both variants address gender, race, religion, and other sensitive attributes. The method is single-step: it does not use a previous model output or a generated explanation.

  5. Knowl 5 — Evaluation covers stereotypical associations, social regard, toxicity, and four language models

    experimental setup

    The evaluation uses GPT-J, Mistral, MPT-Instruct, and Llama-2. The methods description lists Llama-2 as 7B, while the result tables label it as 13B. Models were run from Hugging Face weights and implementations on one 32 GB NVIDIA V100 GPU. Generation temperature was 1.0; StereoSet additionally used a repetition penalty of 1.3, and beam search was the default decoding strategy.

    StereoSet measures gender, race, religion, and profession associations. The experiments use about 53% of its sentence clusters, restricted to prompts with the blank at the end so a causal language model can generate the missing word; the authors report that base-model results on this subset are similar to results on the full dataset. The stereotype score (SS) is the fraction of comparisons in which the stereotypical completion is more probable than the anti-stereotypical one, with 50% ideal. The language-modeling score (LM) measures how often the unrelated completion has the lowest probability, with 100% ideal. ICAT combines SS and LM to reflect their tradeoff, with 100% ideal.

    Regard estimates social perception rather than relying only on sentiment. It compares generated outputs for paired demographics in gender, race, and sexual orientation; the ideal paired-group Regard difference is zero. Toxicity is evaluated on 1,000 randomly sampled RealToxicityPrompts using a fine-tuned hate-speech classifier. The reported measures are mean toxicity and percentage change from the base model.

  6. Knowl 6 — Prompt-based debiasing is competitive on TruthfulQA and BoolQ

    empirical result

    On TruthfulQA and BoolQ, the authors report that Self-Refinement and Implication Prompting generally preserve downstream question-answering performance relative to the base models. Averaged across models, Prefix Prompting is associated with a 15% performance decrease on TruthfulQA and no average change on BoolQ. The IP methods achieve the best average TruthfulQA result, reported as a 9% increase over the base models, while SR achieves the best average BoolQ result, reported as a 1% increase. These findings concern the tested model and benchmark settings; they do not establish that prompting preserves performance on all downstream tasks.

  7. Knowl 7 — Ablations show limited benefit from extra refinement steps and support flexible implication generation

    empirical result

    The ablation experiments examine prefix wording, SR iteration count, implication-generator size, and fixed versus prompt-specific implications. Across four models, four instruction-prefix variants average 64.12 ICAT, 0.11 Regard, and 0.045 toxicity, while four role-prefix variants average 65.18 ICAT, 0.11 Regard, and 0.043 toxicity. The authors summarize the role variants as having 1.7% higher ICAT and 4.5% lower toxicity than the instruction variants.

    Increasing SR beyond one revision yields little consistent additional benefit: results for k=2,3,4k=2,3,4 remain in ranges comparable to k=1k=1, leading the authors to regard one refinement step as sufficient to obtain the observed gain over PP. For implication generation, TinyLlama 1.1B and Llama-2 13B produce results close to using Mistral itself as the implication generator when revising GPT-J, MPT-Instruct, and Mistral outputs. The authors also find that prompt-specific implications generally work better overall than one generic fixed implication, though the fixed implication has a higher ICAT than the tested IP variants for GPT-J.

  8. Knowl 8 — Comparison with internal or training-based debiasing is strongest on GPT-2 small

    empirical result

    The paper compares prompting with SelfDebias, SentenceDebias, CAFIE, and counterfactual data augmentation (CDA) methods. Because of compute constraints, this comparison is conducted on GPT-2 small (125M); CDA training-based methods could not be run on GPT-J. On GPT-2 small, the base ICAT is 73.63, while role-based SR with k=2k=2 reaches 81.25, above the reported ICAT values for the compared methods, including 75.96 for CAFIE and 75.99 for SentenceDebias with race. On GPT-J, the highest reported prompting ICAT is 72.51 for role-based SR with k=1k=1, compared with 74.72 for CAFIE. Thus the comparison supports the potential of prompting, especially on GPT-2 small, but does not show that prompting beats every comparator in every tested setting.

  9. Knowl 9 — Bias reductions do not show a uniform language-modeling penalty

    empirical result

    The authors find no consistent relationship between lower bias and worse language modeling across the four evaluated models. Averaged across models, PP is reported to increase StereoSet LM score by 0.61% relative to the base models, SR to decrease it by 0.46%, and IP to decrease it by 0.09%. The model-level trends differ: GPT-J and MPT-Instruct show LM-score decreases as stereotype scores improve, whereas Mistral and Llama-2 can have higher LM scores under multi-step prompting than at baseline. These results indicate that the tested bias reductions do not require a uniform loss in the measured language-modeling ability.

  10. Knowl 10 — Prompt-based debiasing has limits in scope and interpretation

    limitation

    The authors caution that the System 2 framing is a prompt-design analogy, not evidence that language models think or reason as humans do; their outputs may instead reflect textual patterns associated with the prompts. The approach can simplify complex social biases into prompt instructions, depends on models responding consistently across architectures and updates, and may fail when the model’s learned representations do not contain information useful for the task. The evaluation is also constrained by compute: it does not cover larger models such as 70B Llama-2 variants or Mixtral, and it omits other advanced prompting approaches including Tree-of-Thought, Self-Consistency, and Directional Stimulus Prompting.

Coverage note — The complete per-attribute StereoSet matrices and qualitative input-output examples are omitted because they provide supporting detail rather than distinct general findings; aggregate benchmark outcomes and methodological ablations are retained.

References

  1. 1.Pragyan Banerjee, Abhinav Java, Surgan Jandial, Simra Shahid, Shaz Furniturewala, Balaji Krishnamurthy, and Sumit Bhatia. 2023. All should be equal in the eyes of language models: Counterfactually aware fair text generation.
  2. 2.Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the dangers of stochastic parrots: Can language models be too big?. In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency, pages 610–623.
  3. 3.Conrad Borchers, Dalia Gala, Benjamin Gilburt, Eduard Oravkin, Wilfried Bounsi, Yuki M Asano, and Hannah Kirk. 2022. Looking for a handsome carpenter! debiasing GPT-3 job advertisements. In Proceedings of the 4th Workshop on Gender Bias in Natural Language Processing (GeBNLP), pages 212–224, Seattle, Washington. Association for Computational Linguistics.
  4. 4.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020a. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc.
  5. 5.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020b. Language models are few-shot learners.
  6. 6.Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019. Boolq: Exploring the surprising difficulty of natural yes/no questions.
  7. 7.Isabel O Gallegos, Ryan A Rossi, Joe Barrow, Md Mehrab Tanjim, Sungchul Kim, Franck Dernoncourt, Tong Yu, Ruiyi Zhang, and Nesreen K Ahmed. 2023. Bias and fairness in large language models: A survey. arXiv preprint arXiv:2309.00770.
  8. 8.Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A Smith. 2020. Realtoxicityprompts: Evaluating neural toxic degeneration in language models. arXiv preprint arXiv:2009.11462.
  9. 9.Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2023. Mistral 7b.
  10. 10.Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne Lachaux, Pierre Stock, Sandeep Subramanian, Sophia Yang, Szymon Antoniak, Teven Le Scao, Théophile Gervet, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2024. Mixtral of experts.
  11. 11.Daniel Kahneman. 2011. Thinking, fast and slow. macmillan.
  12. 12.Daniel Kahneman and Amos Tversky. 2013. Prospect theory: An analysis of decision under risk. In Handbook of the fundamentals of financial decision making: Part I, pages 99–127. World Scientific.
  13. 13.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. Advances in neural information processing systems, 35:22199–22213.
  14. 14.Aobo Kong, Shiwan Zhao, Hao Chen, Qicheng Li, Yong Qin, Ruiqi Sun, and Xin Zhou. 2023. Better zero-shot reasoning with role-play prompting.
  15. 15.Yingji Li, Mengnan Du, Rui Song, Xin Wang, and Ying Wang. 2023a. A survey on fairness in large language models. arXiv preprint arXiv:2308.10149.
  16. 16.Zekun Li, Baolin Peng, Pengcheng He, Michel Galley, Jianfeng Gao, and Xifeng Yan. 2023b. Guiding large language models via directional stimulus prompting.
  17. 17.Paul Pu Liang, Irene Mengze Li, Emily Zheng, Yao Chong Lim, Ruslan Salakhutdinov, and Louis-Philippe Morency. 2020. Towards debiasing sentence representations. arXiv preprint arXiv:2007.08100.
  18. 18.Paul Pu Liang, Chiyu Wu, Louis-Philippe Morency, and Ruslan Salakhutdinov. 2021. Towards understanding and mitigating social biases in language models. In International Conference on Machine Learning, pages 6565–6576. PMLR.
  19. 19.Stephanie Lin, Jacob Hilton, and Owain Evans. 2022. Truthfulqa: Measuring how models mimic human falsehoods.
  20. 20.Huan Ma, Changqing Zhang, Yatao Bian, Lemao Liu, Zhirui Zhang, Peilin Zhao, Shu Zhang, Huazhu Fu, Qinghua Hu, and Bingzhe Wu. 2023. Fairness-guided few-shot prompting for large language models. arXiv preprint arXiv:2303.13217.
  21. 21.Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Sean Welleck, Bodhisattwa Prasad Majumder, Shashank Gupta, Amir Yazdanbakhsh, and Peter Clark. 2023. Self-refine: Iterative refinement with self-feedback.
  22. 22.Moin Nadeem, Anna Bethke, and Siva Reddy. 2020. Stereoset: Measuring stereotypical bias in pretrained language models. arXiv preprint arXiv:2004.09456.
  23. 23.Timo Schick, Sahana Udupa, and Hinrich Schütze. 2021. Self-diagnosis and self-debiasing: A proposal for reducing corpus-based bias in nlp. Transactions of the Association for Computational Linguistics, 9:1408–1424.
  24. 24.Emily Sheng, Kai-Wei Chang, Premkumar Natarajan, and Nanyun Peng. 2019. The woman worked as a babysitter: On biases in language generation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3407–3412, Hong Kong, China. Association for Computational Linguistics.
  25. 25.Chenglei Si, Zhe Gan, Zhengyuan Yang, Shuohang Wang, Jianfeng Wang, Jordan Boyd-Graber, and Lijuan Wang. 2023. Prompting gpt-3 to be reliable.
  26. 26.Lichao Sun, Yue Huang, Haoran Wang, Siyuan Wu, Qihui Zhang, Chujie Gao, Yixin Huang, Wenhan Lyu, Yixuan Zhang, Xiner Li, Zhengliang Liu, Yixin Liu, Yijue Wang, Zhikun Zhang, Bhavya Kailkhura, Caiming Xiong, Chaowei Xiao, Chunyuan Li, Eric Xing, Furong Huang, Hao Liu, Heng Ji, Hongyi Wang, Huan Zhang, Huaxiu Yao, Manolis Kellis, Marinka Zitnik, Meng Jiang, Mohit Bansal, James Zou, Jian Pei, Jian Liu, Jianfeng Gao, Jiawei Han, Jieyu Zhao, Jiliang Tang, Jindong Wang, John Mitchell, Kai Shu, Kaidi Xu, Kai-Wei Chang, Lifang He, Lifu Huang, Michael Backes, Neil Zhenqiang Gong, Philip S. Yu, Pin-Yu Chen, Quanquan Gu, Ran Xu, Rex Ying, Shuiwang Ji, Suman Jana, Tianlong Chen, Tianming Liu, Tianyi Zhou, William Wang, Xiang Li, Xiangliang Zhang, Xiao Wang, Xing Xie, Xun Chen, Xuyu Wang, Yan Liu, Yanfang Ye, Yinzhi Cao, Yong Chen, and Yue Zhao. 2024. Trustllm: Trustworthiness in large language models.
  27. 27.MosaicML NLP Team. 2023. Introducing mpt-7b: A new standard for open-source, commercially usable llms. Accessed: 2023-05-05.
  28. 28.Vishesh Thakur. 2023. Unveiling gender bias in terms of profession across llms: Analyzing and addressing sociological implications. arXiv preprint arXiv:2307.09162.
  29. 29.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023. Llama 2: Open foundation and fine-tuned chat models.
  30. 30.Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart Shieber. 2020. Investigating gender bias in language models using causal mediation analysis. Advances in neural information processing systems, 33:12388–12401.
  31. 31.Ben Wang and Aran Komatsuzaki. 2021. GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model. https://github.com/kingoflolz/mesh-transformer-jax.
  32. 32.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023. Self-consistency improves chain of thought reasoning in language models.
  33. 33.Kellie Webster, Xuezhi Wang, Ian Tenney, Alex Beutel, Emily Pitler, Ellie Pavlick, Jilin Chen, Ed Chi, and Slav Petrov. 2020. Measuring and reducing gendered correlations in pre-trained models. arXiv preprint arXiv:2010.06032.
  34. 34.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed Chi, Quoc V Le, and Denny Zhou. 2022. Chain-of-thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems, volume 35, pages 24824–24837. Curran Associates, Inc.
  35. 35.Zhongbin Xie and Thomas Lukasiewicz. 2023. An empirical analysis of parameter-efficient methods for debiasing pre-trained language models.
  36. 36.Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik Narasimhan. 2023. Tree of thoughts: Deliberate problem solving with large language models.
  37. 37.Peiyuan Zhang, Guangtao Zeng, Tianduo Wang, and Wei Lu. 2024. Tinyllama: An open-source small language model.
  38. 38.Ran Zmigrod, Sabrina J Mielke, Hanna Wallach, and Ryan Cotterell. 2019. Counterfactual data augmentation for mitigating gender stereotypes in languages with rich morphology. arXiv preprint arXiv:1906.04571.

Citation

MLA
Furniturewala, S., et al. ““Thinking” Fair and Slow: On the Efficacy of Structured Prompts for Debiasing Language Models”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024, pp. 213–27, https://doi.org/10.18653/v1/2024.emnlp-main.13.
APA
Furniturewala, S., Jandial, S., Java, A., Banerjee, P., Shahid, S., Bhatia, S., & Jaidka, K. (2024). “Thinking” Fair and Slow: On the Efficacy of Structured Prompts for Debiasing Language Models. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 213–227. https://doi.org/10.18653/v1/2024.emnlp-main.13
Chicago
Furniturewala, S., S. Jandial, A. Java, et al. 2024. ““Thinking” Fair and Slow: On the Efficacy of Structured Prompts for Debiasing Language Models”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 213–27. https://doi.org/10.18653/v1/2024.emnlp-main.13.
Harvard
Furniturewala, S. et al. (2024) ““Thinking” Fair and Slow: On the Efficacy of Structured Prompts for Debiasing Language Models”, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 213–227. Available at: https://doi.org/10.18653/v1/2024.emnlp-main.13.
Vancouver
1. Furniturewala S, Jandial S, Java A, Banerjee P, Shahid S, Bhatia S, Jaidka K (2024) “Thinking” Fair and Slow: On the Efficacy of Structured Prompts for Debiasing Language Models. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 213–227

BibTeX

@inproceedings{furniturewala-etal-2024-thinking,
    title = "``Thinking'' Fair and Slow: On the Efficacy of Structured Prompts for Debiasing Language Models",
    author = "Furniturewala, Shaz  and
      Jandial, Surgan  and
      Java, Abhinav  and
      Banerjee, Pragyan  and
      Shahid, Simra  and
      Bhatia, Sumit  and
      Jaidka, Kokil",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.13/",
    doi = "10.18653/v1/2024.emnlp-main.13",
    pages = "213--227"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/