A Survey on Bias and Fairness in Machine Learning
Ninareh MehrabiFred MorstatterNripsuta SaxenaKristina LermanAram Galstyan
Categorizes the sources of bias throughout the machine learning pipeline and establishes a structured taxonomy of mathematical fairness definitions and mitigation strategies across real-world artificial intelligence applications.
- Paper: Fairness through awareness, Cynthia Dwork et al. (2012). It introduces foundational definitions of individual and group fairness that form the bedrock of the survey's fairness taxonomy.
- Paper: Equality of Opportunity in Supervised Learning, Moritz Hardt et al. (2016). It formalizes key fairness criteria like equalized odds and equal opportunity in supervised learning, which are centrally categorized in the survey.
- Paper: Inherent Trade-Offs in the Fair Determination of Risk Scores, Jon Kleinberg et al. (2017). It proves the fundamental mathematical trade-offs and impossibilities among core fairness metrics that the survey discusses.
- Paper: Certifying and Removing Disparate Impact, Michael Feldman et al. (2014). It provides the mathematical framework for certifying and removing disparate impact via data pre-processing, a primary mitigation strategy reviewed in the survey.
- Paper: Learning Fair Representations, Richard Zemel et al. (2013). It establishes representation learning techniques to mitigate bias before downstream classification, providing a core pre-processing reference for the survey.
- Paper: Counterfactual Fairness, Matt J. Kusner et al. (2017). It develops counterfactual fairness using causal modeling, offering a vital theoretical perspective reviewed within the survey's fairness taxonomy.
- Paper: Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification, Joy Buolamwini et al. (2018). It provides seminal empirical evidence of intersectional bias in commercial vision systems, illustrating the real-world harms analyzed in the survey.
- Paper: Semantics derived automatically from language corpora contain human-like biases, Aylin Caliskan et al. (2016). It demonstrates how human-like biases are automatically embedded in NLP models and embeddings, establishing foundational evidence for the survey's NLP domain analysis.
- Paper: A Reductions Approach to Fair Classification, Alekh Agarwal et al. (2018). It provides the reductions-based in-processing framework for constrained fair classification that is categorized among standard algorithmic mitigation approaches.
- Paper: Big Data's Disparate Impact, Solon Barocas et al. (2016). It outlines the technical mechanisms through which data mining can introduce legal disparate impact, motivating the survey's taxonomy of bias sources.
- Paper: AI Fairness 360: An extensible toolkit for detecting and mitigating algorithmic bias, Rachel Bellamy et al. (2019). It translates the survey's theoretical taxonomy of metrics and mitigation methods into an open-source, production-ready software toolkit.
- Paper: Fairness and Abstraction in Sociotechnical Systems, Andrew D. Selbst et al. (2019). It critically examines the limitations of purely mathematical fairness definitions by analyzing the sociotechnical traps encountered when deploying fair ML systems.
- Paper: Model Cards for Model Reporting, Margaret Mitchell et al. (2019). It establishes a standardized reporting framework to document subgroup performance disparities and ethical considerations identified in the survey.
- Paper: Datasheets for datasets, Timnit Gebru et al. (2021). It introduces a structured documentation methodology to audit datasets for the specific sources of bias cataloged in the survey.
- Paper: Fairness in Recommendation Ranking through Pairwise Comparisons, Alex Beutel et al. (2019). It extends group fairness principles to large-scale recommender systems using pairwise ranking comparison techniques.
- Paper: Explainable Artificial Intelligence (XAI): Concepts, Taxonomies, Opportunities and Challenges toward Responsible AI, Alejandro Barredo Arrieta et al. (2020). It situates algorithmic fairness within the broader paradigm of explainable and responsible artificial intelligence.
- Paper: Ethical and social risks of harm from Language Models, Laura Weidinger et al. (2022). It builds on fairness taxonomies to construct a dedicated framework of ethical, discriminatory, and social risks specific to large language models.
- Paper: On the Opportunities and Risks of Foundation Models, Rishi Bommasani et al. (2021). It generalizes concerns about algorithmic bias and fairness to the emerging societal risks and systemic failures posed by foundation models.
