Taxonomizing and Measuring Representational Harms: A Look at Image Tagging
Jared KatzmanAngelina WangMorgan Klaus ScheuermanSu Lin BlodgettKristen LairdHanna M. WallachSolon Barocas
Categorizes computational fairness metrics and representational harms in image tagging systems to demonstrate that standard measurement approaches fail to uniquely map to specific harms and that mitigating one harm can inadvertently worsen another.
Computer vision technologies, specifically image tagging systems, are widely deployed in public-facing applications such as generating alternative text for accessibility, indexing search engines, and organizing photos. Because image tags communicate salience directly to human users, problematic outputs can shape broader societal beliefs, attitudes, and cultural understandings. Most algorithmic fairness research has historically focused on allocational harms—decisions that unfairly deny people resources or opportunities, such as in lending or employment. In contrast, image tagging primarily produces representational harms by reinforcing harmful social hierarchies, yet existing technical evaluations often conflate these issues under broad terms like bias, unfairness, or discrimination without clear normative grounding.
The article establishes a taxonomy that categorizes quantitative fairness evaluations for image tagging and conceptualizes the specific representational harms these systems create. Through this taxonomy, the analysis demonstrates how quantitative metrics map onto distinct representational harms and evaluates the inherent trade-offs involved in mitigating them.
The researchers conducted an extensive conceptual and analytical literature review of quantitative measurement methods used across computer vision. They structured these evaluation techniques into five distinct categories: incidence-based (detecting inherently problematic tags or pairings), distribution-based (measuring demographic skew in training sets and outputs), performance-based (comparing tagging accuracy across demographic groups), perturbation-based (testing output consistency under controlled image variations), and internals-based (analyzing model feature representations and gradient-based saliency maps). The article then defined four core types of representational harms and mapped the five measurement categories against each harm type.
The analysis yielded several critical findings. First, representational harms in image tagging manifest in four primary ways: reifying fluid social groups into rigid biological categories, perpetuating stereotypes by reinforcing occupational or behavioral associations, demeaning groups through slurs or derogatory context, and erasing groups by omitting tags for cultural artifacts or marginalized identities. Second, algorithmic categorization frequently denies individuals the ability to self-identify, depriving them of autonomy and compounding group-level harms. Third, there is no direct one-to-one relationship between measurement techniques and harm types; every computational metric category can evaluate any of the four harms, but each captures only a partial facet of the issue. Finally, mitigation strategies often exist in direct technical tension: removing demographic tags prevents reification but risks erasing marginalized identities, while equalizing tag distributions across groups can inadvertently suppress genuine cultural distinctions.
These findings indicate that technical teams cannot rely on a single fairness metric as a definitive measure of system safety or equity. For decision-makers and compliance leaders, relying on off-the-shelf bias benchmarks without explicit ethical objectives creates legal and reputational risks. Because the application of an algorithm determines what constitutes harm, a universal technical fix is impossible; system design must account for the specific social context and downstream user needs.
To manage these challenges, organizations developing or procuring image tagging systems should clearly specify the exact representational harms they intend to measure and mitigate. Technical teams must deploy multi-faceted evaluation suites combining behavioral testing, distribution audits, and internal explainability tools rather than depending on single summary scores. When designing mitigations, teams must evaluate context-specific trade-offs—for example, prioritizing erasure avoidance in accessibility applications while emphasizing stereotype prevention in broad consumer search platforms.
The article focuses exclusively on quantitative measurement methods within image tagging and captioning, excluding qualitative auditing approaches. While confidence in the conceptual taxonomy and its demonstrated mitigation trade-offs is high, practitioners must exercise caution. Because image tags and human visual culture are infinitely variable, pre-defined measurement tests cannot anticipate every possible harm, necessitating ongoing human review and context-aware governance.
- Paper: Language (Technology) is Power: A Critical Survey of “Bias” in NLP, Su Lin Blodgett et al. (2020). This paper establishes the foundational taxonomy of representational versus allocational harms and critiques normative mismatches in NLP bias measurements, which directly motivates the source paper's taxonomy of representational harms in vision systems.
- Paper: Fairness and Abstraction in Sociotechnical Systems, Andrew D. Selbst et al. (2019). This work introduces the formalism and framing traps when reducing sociotechnical harms to mathematical metrics, underpinning the source paper's analysis of why computational fairness metrics fail to map one-to-one to representational harms.
- Paper: Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification, Joy Buolamwini et al. (2018). This seminal study documents severe intersectional representational and classification disparities in commercial computer vision models, providing the core empirical motivation for evaluating fairness in visual tagging.
- Paper: A Survey on Bias and Fairness in Machine Learning, Ninareh Mehrabi et al. (2019). This survey organizes the theoretical definitions of fairness and sources of machine learning bias across modalities, providing foundational background for the computational measurement categories analyzed in the source.
- Paper: Inherent Trade-Offs in the Fair Determination of Risk Scores, Jon Kleinberg et al. (2017). This paper establishes mathematical trade-offs between competing fairness definitions, establishing a theoretical precedent for the source's finding that mitigating different representational harms creates inherent mutual tensions.
- Paper: Discovering and Mitigating Visual Biases Through Keyword Explanation, Younghyun Kim et al. (2024). This work directly operationalizes the discovery and mitigation of visual and semantic biases by mapping mispredictions to descriptive keywords in vision-language frameworks.
- Paper: OpenBias: Open-Set Bias Detection in Text-to-Image Generative Models, Moreno D'Incà et al. (2024). This paper extends the measurement of representational harms from discriminative image tagging to open-set bias discovery and quantification in text-to-image generative models.
- Paper: Large Language Models are Geographically Biased, Rohin Manvi et al. (2024). This study applies fine-grained representational harm and bias auditing to the geographic domain across foundation models.
