"Thinking" Fair and Slow: On the Efficacy of Structured Prompts for Debiasing Language Models
Shaz FurniturewalaSurgan JandialAbhinav JavaPragyan BanerjeeSimra ShahidSumit BhatiaKokil Jaidka
Presents a System 2-inspired prompting framework that enables end users to significantly reduce social biases in closed-source and open-source language models without retraining weights or accessing internal probability distributions.
Large language models frequently inherit and perpetuate societal stereotypes present in their training data. While traditional bias-mitigation methods rely on model retraining, internal parameter adjustments, or access to output probabilities, many state-of-the-art models are deployed as closed systems accessible only through interfaces that restrict internal access. Even when open-source models are available, curating data and fine-tuning models is computationally expensive and risks degrading general task performance. Consequently, end-users requiring fairer text generation need black-box mitigation techniques that work exclusively through prompt design.
The article evaluates whether structured prompting strategies can mitigate bias in closed and open language models without internal modifications or loss in general performance. Specifically, it organizes and tests prompting methods across three structured categories inspired by human cognitive decision-making processes: direct prefix instructions, iterative self-refinement, and multi-step implication prompting that explicitly reasons about underlying stereotypes.
To assess these strategies, the authors evaluated four prominent open-access language models (GPT-J 6B, Mistral 7B, MPT-Instruct 7B, and Llama-2 13B) across established benchmarks measuring stereotypical bias across gender, race, religion, and profession, social regard differences, and completion toxicity. The evaluation compared single-step prefixing, multi-iteration self-refinement, and dynamic implication prompting against baseline models and existing internal modification techniques. Downstream question-answering benchmarks were also evaluated to verify whether general model capability was preserved.
The findings show that structured, reasoning-based prompting significantly reduces bias while matching or exceeding the performance of internal debiasing techniques. Implication Prompting proved the most effective across all benchmarks, reducing stereotype scores by an average of about 4% and bias on social perception metrics by roughly 27% compared to other prompting methods, while decreasing output toxicity by approximately 7%. Iterative Self-Refinement delivered strong improvements over single-step prefixes, though running more than one refinement step yielded diminishing returns, improving stereotype metrics by less than 0.5%. Across all methods, assigning an unbiased persona or role outperformed simple negative instructions, and smaller language models could effectively generate the required implications without degrading overall debiasing quality.
These results demonstrate that end-users can successfully eliminate bias in third-party and proprietary language models without costly fine-tuning or specialized infrastructure. Because these prompting pipelines preserve underlying language modeling capabilities—showing parity on factual and general question-answering tasks—organizations can adopt them without risking operational performance. When implementing prompt-based fairness controls, teams should prioritize dynamically generated implications or single-iteration role-based refinements rather than basic instruction prefixes or compute-intensive multiple refinement loops.
Decision-makers should note certain limitations: the empirical tests were bounded by resource constraints to models up to 13 billion parameters, and they depend on the model's pre-existing internal data containing sufficient context to navigate toward fairer outputs. Furthermore, structured prompts direct stochastic search rather than instill genuine moral reasoning. Nevertheless, the high statistical confidence across diverse benchmarks indicates that structured prompt frameworks are a practical and robust tool for end-user bias mitigation.
- Paper: An Empirical Survey of the Effectiveness of Debiasing Techniques for Pre-trained Language Models, Nicholas Meade et al. (2022). Its comparison of established debiasing methods, including Self-Debias, provides the mitigation baseline against which this paper evaluates prompt-only alternatives.
- Paper: StereoSet: Measuring stereotypical bias in pretrained language models, Moin Nadeem et al. (2020). This paper introduces StereoSet, one of the stereotype-bias benchmarks used to assess the source’s prompting methods.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). Its foundational account of chain-of-thought prompting clarifies the stepwise reasoning approach that informs the source’s structured implication prompts.
No sufficiently relevant recommendations were found.
