Taxonomizing and Measuring Representational Harms: A Look at Image Tagging

Jared KatzmanAngelina WangMorgan Klaus ScheuermanSu Lin BlodgettKristen LairdHanna M. WallachSolon Barocas

article2023AAAI55 citations

Categorizes computational fairness metrics and representational harms in image tagging systems to demonstrate that standard measurement approaches fail to uniquely map to specific harms and that mitigating one harm can inadvertently worsen another.

Listen

Computer vision technologies, specifically image tagging systems, are widely deployed in public-facing applications such as generating alternative text for accessibility, indexing search engines, and organizing photos. Because image tags communicate salience directly to human users, problematic outputs can shape broader societal beliefs, attitudes, and cultural understandings. Most algorithmic fairness research has historically focused on allocational harms—decisions that unfairly deny people resources or opportunities, such as in lending or employment. In contrast, image tagging primarily produces representational harms by reinforcing harmful social hierarchies, yet existing technical evaluations often conflate these issues under broad terms like bias, unfairness, or discrimination without clear normative grounding.

The article establishes a taxonomy that categorizes quantitative fairness evaluations for image tagging and conceptualizes the specific representational harms these systems create. Through this taxonomy, the analysis demonstrates how quantitative metrics map onto distinct representational harms and evaluates the inherent trade-offs involved in mitigating them.

The researchers conducted an extensive conceptual and analytical literature review of quantitative measurement methods used across computer vision. They structured these evaluation techniques into five distinct categories: incidence-based (detecting inherently problematic tags or pairings), distribution-based (measuring demographic skew in training sets and outputs), performance-based (comparing tagging accuracy across demographic groups), perturbation-based (testing output consistency under controlled image variations), and internals-based (analyzing model feature representations and gradient-based saliency maps). The article then defined four core types of representational harms and mapped the five measurement categories against each harm type.

The analysis yielded several critical findings. First, representational harms in image tagging manifest in four primary ways: reifying fluid social groups into rigid biological categories, perpetuating stereotypes by reinforcing occupational or behavioral associations, demeaning groups through slurs or derogatory context, and erasing groups by omitting tags for cultural artifacts or marginalized identities. Second, algorithmic categorization frequently denies individuals the ability to self-identify, depriving them of autonomy and compounding group-level harms. Third, there is no direct one-to-one relationship between measurement techniques and harm types; every computational metric category can evaluate any of the four harms, but each captures only a partial facet of the issue. Finally, mitigation strategies often exist in direct technical tension: removing demographic tags prevents reification but risks erasing marginalized identities, while equalizing tag distributions across groups can inadvertently suppress genuine cultural distinctions.

These findings indicate that technical teams cannot rely on a single fairness metric as a definitive measure of system safety or equity. For decision-makers and compliance leaders, relying on off-the-shelf bias benchmarks without explicit ethical objectives creates legal and reputational risks. Because the application of an algorithm determines what constitutes harm, a universal technical fix is impossible; system design must account for the specific social context and downstream user needs.

To manage these challenges, organizations developing or procuring image tagging systems should clearly specify the exact representational harms they intend to measure and mitigate. Technical teams must deploy multi-faceted evaluation suites combining behavioral testing, distribution audits, and internal explainability tools rather than depending on single summary scores. When designing mitigations, teams must evaluate context-specific trade-offs—for example, prioritizing erasure avoidance in accessibility applications while emphasizing stereotype prevention in broad consumer search platforms.

The article focuses exclusively on quantitative measurement methods within image tagging and captioning, excluding qualitative auditing approaches. While confidence in the conceptual taxonomy and its demonstrated mitigation trade-offs is high, practitioners must exercise caution. Because image tags and human visual culture are infinitely variable, pre-defined measurement tests cannot anticipate every possible harm, necessitating ongoing human review and context-aware governance.

arXiv: 2305.01776
  • Paper: Language (Technology) is Power: A Critical Survey of “Bias” in NLP, Su Lin Blodgett et al. (2020). This paper establishes the foundational taxonomy of representational versus allocational harms and critiques normative mismatches in NLP bias measurements, which directly motivates the source paper's taxonomy of representational harms in vision systems.
  • Paper: Fairness and Abstraction in Sociotechnical Systems, Andrew D. Selbst et al. (2019). This work introduces the formalism and framing traps when reducing sociotechnical harms to mathematical metrics, underpinning the source paper's analysis of why computational fairness metrics fail to map one-to-one to representational harms.
  • Paper: Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification, Joy Buolamwini et al. (2018). This seminal study documents severe intersectional representational and classification disparities in commercial computer vision models, providing the core empirical motivation for evaluating fairness in visual tagging.
  • Paper: A Survey on Bias and Fairness in Machine Learning, Ninareh Mehrabi et al. (2019). This survey organizes the theoretical definitions of fairness and sources of machine learning bias across modalities, providing foundational background for the computational measurement categories analyzed in the source.
  • Paper: Inherent Trade-Offs in the Fair Determination of Risk Scores, Jon Kleinberg et al. (2017). This paper establishes mathematical trade-offs between competing fairness definitions, establishing a theoretical precedent for the source's finding that mitigating different representational harms creates inherent mutual tensions.
Cover for Taxonomizing and Measuring Representational Harms: A Look at Image Tagging

Abstract

In this paper, we examine computational approaches for measuring the “fairness” of image tagging systems, finding that they cluster into five distinct categories, each with its own analytic foundation. We also identify a range of normative concerns that are often collapsed under the terms “unfairness,” “bias,” or even “discrimination” when discussing problematic cases of image tagging. Specifically, we identify four types of representational harms that can be caused by image tagging systems, providing concrete examples of each. We then consider how different computational measurement approaches map to each of these types, demonstrating that there is not a one-to-one mapping. Our findings emphasize that no single measurement approach will be definitive and that it is not possible to infer from the use of a particular measurement approach which type of harm was intended to be measured. Lastly, equipped with this more granular understanding of the types of representational harms that can be caused by image tagging systems, we show that attempts to mitigate some of these types of harms may be in tension with one another.

Table of Contents

  • Introduction
  • Background
  • Image Tagging
  • Representational Harms
  • Computational Measurement Approaches
  • Representational Harms
  • Four Types of Representational Harms
  • Denying People the Ability to Self-Identify
  • Mapping Harms to Measurement Approaches
  • Stereotyping Social Groups
  • Implications of these Mappings
  • Tensions When Mitigating Harms
  • Conclusion
  • References

Knowls

  1. Knowl 1 — Five analytic categories of quantitative image-tagging measurements

    model/method

    The paper organizes quantitative approaches for assessing image-tagging systems into five categories, based on their analytic foundations. Image tagging applies tags to salient aspects of an image—including objects, actions, facial expressions, and social groups—and differs from object recognition, which aims to identify all depicted objects and emphasizes completeness rather than salience.

    • Incidence-based approaches test whether pre-identified problematic tags, image types, or tag–image pairings occur in training data, system inputs, or system outputs. They assume that certain tags, images, or pairings are inherently unacceptable, but require prior knowledge to identify them and are unlikely to enumerate every problematic case.
    • Distribution-based approaches compare distributions of images, social groups, attributes, activities, or tags across datasets, system inputs, and system outputs. They can test representation against a reference distribution, compare tags assigned to different social groups, or test whether a system amplifies differences already present in its training data.
    • Performance-based approaches compare system performance across social groups or image types. They may evaluate group-membership tags, such as gender or race, or evaluate tags for any depicted content, such as everyday objects.
    • Perturbation-based approaches vary an image’s properties—using altered appearances, counterfactual images, or semantically similar image pairs—and test whether the system’s tags change. They are used to identify spurious visual correlations and to test invariance to aspects that should not affect tagging.
    • Internals-based approaches inspect learned representations or the internal basis for predictions. Saliency maps can show which image regions influence a tag, while representation analyses can test whether a system encodes social-group concepts even when it is not intended to apply social-group tags.

    The paper’s taxonomy is explicitly about quantitative measurement; qualitative measurement approaches are left for future work.

  2. Knowl 2 — Representational harms as reproduction of social hierarchies

    definition

    The paper frames the principal harms of image tagging as representational harms: harms that affect the social standing of groups by contributing to harmful social hierarchies, such as casting some groups as inferior or confining them to subordinate positions. These harms concern groups rather than particular individuals. They differ from allocational harms, which arise when people are unfairly denied opportunities or resources such as employment, housing, or credit.

    Image-tagging outputs are often consumed directly by people—for example, as alt text, search labels, or photo-grouping tags—so they influence what people perceive as salient and how they understand social groups. The paper identifies four forms of representational harm in image tagging: reifying social groups, stereotyping social groups, demeaning social groups, and erasing social groups. The taxonomy is not claimed to be exhaustive because the large space of possible image contents, tags, and tag–image relationships may produce additional forms of hierarchy-reproducing harm.

  3. Knowl 3 — Reifying social groups through visual tags

    definition

    Reifying social groups means treating socially constructed and contested groups as if they were natural, fixed, objectively measurable categories with stable boundaries. Image-tagging systems are especially prone to this harm when they infer pre-defined social-group tags from visual appearance, thereby suggesting that group membership is visually observable and that the system’s chosen group boundaries are immutable.

    Binary gender classification illustrates the problem: such a system presupposes both that gender can be determined from appearance and that only two gender identities exist. Even when the system applies its tags consistently, it can reinforce the belief that these contested social categories and their relationships to appearance are natural facts.

  4. Knowl 4 — Stereotyping social groups through tag associations

    definition

    Stereotyping social groups occurs when image-tagging systems perpetuate over-generalized beliefs that connect particular groups with particular appearances, activities, occupations, or properties in ways that reproduce harmful hierarchies.

    This can happen when a system is intended to apply group-membership tags: associating snowboarding or another visual activity with men makes images containing that activity more likely to receive a male tag. It can also happen when the system is intended to apply non-group tags: associating visual characteristics common among women, such as long hair, with nursing can cause female doctors to be tagged as nurses. Systems may stereotype even when individual tags are correct—for example, by applying occupation tags more often to images of men and appearance-related tags more often to images of women, thereby treating professional activity as more salient for men and appearance as more salient for women.

  5. Knowl 5 — Demeaning social groups through tags and unequal evaluation

    definition

    Demeaning social groups means suggesting that members of a group are less worthy of respect or esteem than members of other groups. The harm can arise from explicit racial epithets, from otherwise ordinary tags whose historical or social context makes them degrading, or from tags that trivialize attributes, artifacts, or activities associated with a group.

    For example, applying the tag gorilla to an image of Black people is demeaning because of the history of comparing and dehumanizing Black people through that analogy. Applying costume to a religious group’s wedding clothing can trivialize the group’s traditions. Evaluative tags can also demean through unequal application: applying beautiful less often to images of one racial group than to images of another casts the first group as less worthy of esteem and reinforces restrictive standards of beauty.

  6. Knowl 6 — Erasing social groups through failure of recognition

    definition

    Erasing social groups occurs when an image-tagging system fails to recognize a group, or fails to recognize attributes, artifacts, activities, or experiences bound up with that group’s identity. Such failures imply that the group is not worthy of recognition and can contribute to its marginalization.

    Examples include failing to apply person to images of people wearing hijabs, failing to provide or apply menorah when a menorah is depicted while providing tags for Christian artifacts, and describing women suffragists only as people, walking, and street. The last case records visible objects and actions but omits the significance of gender and the historical injustice represented by the march, thereby denying how group membership shapes lived experiences of oppression and resistance.

  7. Knowl 7 — Denying self-identification as an additional normative concern

    definition

    Image-tagging systems can deny people the ability to self-identify when they impose social-group labels based on appearance without the affected person’s awareness, participation, or consent. This is an individual harm because it directly affects particular people and can deprive them of autonomy over a consequential aspect of their identity.

    Patterns of imposed labeling can also produce representational and allocational harms. They reify the idea that social groups are visually determinable; use visual characteristics to stereotype; demean people through misgendering or unequal recognition of groups; erase groups absent from the available label set, such as non-binary people; and increase the visibility of labeled people in ways that may expose them to violence, surveillance, employment loss, or other harms.

  8. Knowl 8 — Every measurement category can be used to investigate every harm type

    theoretical result

    The paper’s mapping analysis shows that the five measurement categories do not correspond one-to-one with the four representational harms. Each category can be designed to investigate reification, stereotyping, demeaning, or erasure; the following examples illustrate possible uses rather than claims that every listed behavior occurs in every system.

    • Incidence-based measurement: reification can be examined by checking whether available tags divide people into different races; stereotyping by checking for tags such as boyish and girly; demeaning by checking whether racial epithets are available; and erasure by checking whether tags exist for the Bible but not the Quran or Torah.
    • Distribution-based measurement: reification can be examined when gender-tag distributions differ for people with different hair lengths; stereotyping when profession tags are applied more often to images of men than to images of people of other genders; demeaning when beautiful is applied more often to images of white people than to images of other racial groups; and erasure when foods from the Global North receive specific names while foods from the Global South are labeled only food.
    • Performance-based measurement: reification can be examined when a system appears to apply gender tags perfectly despite the impossibility of reliably determining gender from appearance; stereotyping when doctor is applied more accurately to images of men than to images of people of other genders; demeaning when human is recognized more accurately in images of white people than in images of other racial groups; and erasure when objects are recognized more accurately in images from the Global North than in images from the Global South.
    • Perturbation-based measurement: reification can be examined when changing a person’s appearance changes the assigned gender tag; stereotyping when standardized headshots of politicians receive different tags according to gender; demeaning when adding a cheongsam causes an image to receive the tag costume; and erasure when adding an assistive device causes the system to stop applying person.
    • Internals-based measurement: reification can be examined when a system’s feature representation encodes gender even though the system is not intended to apply gender tags; stereotyping when a saliency map shows attention to a truck while the system incorrectly applies man to an image depicting people of other genders; demeaning when a saliency map shows attention to facial features while applying janitor to an image depicting a person of color; and erasure when a saliency map shows that wheelchair users are ignored while the system applies people to an image of a group.

    Thus, the choice of measurement category alone does not identify the normative harm being investigated.

  9. Knowl 9 — Measurement results require explicit normative interpretation

    theoretical result

    No single computational measurement approach is definitive for representational harms. Different approaches reveal different facets of the same harm: incidence tests may expose an explicitly problematic label, distribution comparisons may expose group-level patterns, performance comparisons may expose unequal correctness, perturbations may expose causal sensitivity to image features, and internals-based analyses may expose the system’s representational or visual basis for a prediction.

    Consequently, a measurement study should state explicitly which representational harm it is investigating and interpret its measurements in light of what the selected approach can and cannot reveal. Without that reasoning, the use of a particular fairness metric or measurement category does not establish whether the intended target was reification, stereotyping, demeaning, erasure, or another harm.

  10. Knowl 10 — Mitigating one representational harm can intensify another

    theoretical result

    Mitigation strategies for image-tagging harms can conflict because recognition and non-recognition have different normative consequences.

    • Removing tags that indicate social-group membership can reduce reification and may reduce visibility-based vulnerability, but it can also erase groups by making them unrecognizable.
    • Adding and correctly applying group-membership tags can mitigate erasure and sometimes provide recognition of injustice or privilege, but it can reify group boundaries, impose identities, or increase vulnerability through visibility.
    • Removing tags for group-associated attributes, artifacts, or activities can prevent demeaning misclassification, yet removing those tags can itself erase the groups whose identities are expressed through them.
    • Filtering images considered inherently problematic may require another vision system to identify them and can erase communities when the filter incorrectly treats images associated with those communities as objectionable.
    • Making tag distributions similar across groups may reduce stereotypical differences, but forcing similarity can deny the role that group membership plays in people’s lived experiences and thereby create erasure.

    The appropriate trade-off depends partly on the deployment context. A social-media system might prioritize removing tags that stereotype or demean, whereas an image-description system for blind or low-vision users might judge the consequences of erasure to be worse than those of some incorrect tagging.

Coverage note — No substantial contributed material was omitted; the paper’s cited prior work and its deferred survey of qualitative measurement approaches were excluded because they are background or explicitly outside the paper’s quantitative scope.

References

  1. 1.Algorithmic Justice League. 2021. Drag Vs AI Workshop.
  2. 2.Alvi, M.; Zisserman, A.; and Nellaker, C. 2018. Turning a Blind Eye: Explicit Removal of Biases and Variation from Deep Neural Network Embeddings. In Workshop on Bias Estimation in Face Analytics at ECCV 2018.
  3. 3.Baker, D.; Hanna, A.; and Denton, E. 2020. Algorithmically Encoded Identities: Reframing Human Classification. In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, FAT* ’20, 681. New York, NY, USA: Association for Computing Machinery. ISBN 9781450369367.
  4. 4.Bard, J. 2020. Developing a Legal Framework for Regulating Emotion AI. In University of Florida Levin College of Law Research Paper.
  5. 5.Barlas, P.; Kyriakou, K.; Guest, O.; Kleanthous, S.; and Otterbacher, J. 2021. To ”See” is to Stereotype: Image Tagging Algorithms, Gender Recognition, and the Accuracy-Fairness Trade-off. Proceedings of the ACM on Human-Computer Interaction, 4(CSCW3): 1–31.
  6. 6.Barlas, P.; Kyriakou, K.; Kleanthous, S.; and Otterbacher, J. 2019a. Social B(eye)as: Human and Machine Descriptions of People Images. In Proceedings of the International AAAI Conference on Web and Social Media.
  7. 7.Barlas, P.; Kyriakou, K.; Kleanthous, S.; and Otterbacher, J. 2019b. What Makes an Image Tagger Fair? Proprietary Auto-tagging and Interpretations on People Images. In Proceedings of the 27th ACM Conference On User Modelling, Adaptation And Personalization (UMAP).
  8. 8.Barocas, S.; Crawford, K.; Shapiro, A.; and Wallach, H. 2017. The Problem With Bias: Allocative Versus Representational Harms in Machine Learning. In Proceedings of SIGCIS. Philadelphia, PA.
  9. 9.Benjamin, R. 2019. Race After Technology: Abolitionist Tools for the New Jim Code. John Wiley & Sons.
  10. 10.Bennett, C. L.; Gleason, C.; Scheuerman, M. K.; Bigham, J. P.; Guo, A.; and To, A. 2021. “It’s Complicated”: Negotiating Accessibility and (Mis) Representation in Image Descriptions of Race, Gender, and Disability. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems, 1–19.
  11. 11.Bhargava, S.; and Forsyth, D. 2019. Exposing and correcting the gender bias in image captioning datasets and models. arXiv preprint arXiv:1912.00578.
  12. 12.Blodgett, S. L.; Barocas, S.; Daume, H., III; and Wallach, H. 2020. Language (Technology) is Power: A Critical Survey of “Bias” in NLP. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL), 5454–5476.
  13. 13.Bronstein, C. 2020. Pornography, Trans Visibility, and the Demise of Tumblr. In Transgender Studies Quarterly.
  14. 14.Buolamwini, J.; and Gebru, T. 2018. Gender shades: Intersectional accuracy disparities in commercial gender classification. In Proceedings of the Conference on Fairness, Accountability, and Transparency. New York City, NY.
  15. 15.Crawford, K. 2017. The Trouble with Bias. Conference on Neural Information Processing Systems Keynote.
  16. 16.Crawford, K.; and Paglen, T. 2019. Excavating AI: the politics of images in machine learning training sets. Excavating AI.
  17. 17.Das, A.; Dantcheva, A.; and Bremond, F. 2018. Mitigating Bias in Gender, Age and Ethnicity Classification: a Multi-Task Convolution Neural Network Approach. In Workshop at ECCV 2018.
  18. 18.Denton, E.; Hutchinson, B.; Mitchell, M.; Gebru, T.; and Zaldivar, A. 2021. Image Counterfactual Sensitivity Analysis for Detecting Unintended Bias. arXiv preprint arXiv:1906.06439.
  19. 19.DeVries, T.; Misra, I.; Wang, C.; and van der Maaten, L. 2019. Does Object Recognition Work for Everyone? arXiv:1906.02659.
  20. 20.Hanley, M.; Barocas, S.; Levy, K.; Azenkot, S.; and Nissenbaum, H. 2021. Computer vision and conflicting values: Describing people with automated alt text. In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society, 543–554.
  21. 21.Hanna, A.; Denton, E.; Smart, A.; and Smith-Loud, J. 2020. Towards a Critical Race Methodology in Algorithmic Fairness. In Proceedings of the Conference on Fairness, Accountability, and Transparency, 501–512.
  22. 22.Hendricks, L. A.; Burns, K.; Saenko, K.; Darrell, T.; and Rohrbach, A. 2018. Women also snowboard: Overcoming bias in captioning models. In Proceedings of the European Conference on Computer Vision (ECCV), 771–787.
  23. 23.Hofmann, M.; Kasnitz, D.; Mankoff, J.; and Bennett, C. L. 2020. Living Disability Theory: Reflections on Access, Research, and Design. In The 22nd International ACM SIGACCESS Conference on Computers and Accessibility (ASSETS).
  24. 24.Hu, Z.; and Strout, J. 2018. Exploring Stereotypes and Biased Data with the Crowd. In arXiv:1801.03261.
  25. 25.Jacobs, A. Z.; and Wallach, H. 2021. Measurement and fairness. In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency, 375–385.
  26. 26.Jia, S.; Lansdall-Welfare, T.; and Cristianini, N. 2018. Right for the Right Reason: Training Agnostic Networks. In Lecture Notes in Computer Science.
  27. 27.Kay, M.; Matuszek, C.; and Munson, S. A. 2015a. Unequal Representation and Gender Stereotypes in Image Search Results for Occupations. In Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems, CHI ’15, 3819–3828. New York, NY, USA: Association for Computing Machinery. ISBN 9781450331456.
  28. 28.Kay, M.; Matuszek, C.; and Munson, S. A. 2015b. Unequal Representation and Gender Stereotypes in Image Search Results for Occupations. In Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems, 3819–3828.
  29. 29.Keyes, O. 2018. The misgendering machines: Trans/HCI implications of automatic gender recognition. Proceedings of the ACM on human-computer interaction, 2(CSCW): 1–22.
  30. 30.Khiyari, H. E.; and Wechsler, H. 2016. Face Verification Subject to Varying (Age, Ethnicity, and Gender)Demographics Using Deep Learning. Journal of biometrics & biostatistics, 7: 1–5.
  31. 31.Klare, B. F.; Burge, M. J.; Klontz, J. C.; Vorder Bruegge, R. W.; and Jain, A. K. 2012. Face Recognition Performance: Role of Demographic Information. IEEE Transactions on Information Forensics and Security, 7(6): 1789–1801.
  32. 32.Kyriakou, K.; Barlas, P.; Kleanthous, S.; and Otterbacher, J. 2019. Fairness in Proprietary Image Tagging Algorithms: A Cross-Platform Audit on People Images. In Proceedings of the International AAAI Conference on Web and Social Media.
  33. 33.Karkk̈ainen, K.; and Joo, J. 2021. FairFace: Face Attribute Dataset for Balanced Race, Gender, and Age. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision.
  34. 34.Levin, S. 2016. A beauty contest was judged by AI and the robots didn’t like dark skin. The Guardian.
  35. 35.McDuff, D.; Ma, S.; Song, Y.; and Kapoor, A. 2019. Characterizing Bias in Classifiers using Generative Models. In Conference on Neural Information Processing Systems (NeurIPS).
  36. 36.Miceli, M.; Schuessler, M.; and Yang, T. 2020. Between Subjectivity and Imposition: Power Dynamics in Data Annotation for Computer Vision. In ACM Conference on Computer Supported Cooperative Work (CSCW).
  37. 37.Muthukumar, V.; Pedapati, T.; Ratha, N.; Sattigeri, P.; Wu, C.-W.; Kingsbury, B.; Kumar, A.; Thomas, S.; Mojsilovic, A.; and Varshney, K. R. 2018. Understanding Unequal Gender Classification Accuracy from Face Images. In arXiv:1812.00099.
  38. 38.Noble, S. U. 2018. Algorithms of Oppression: How Search Engines Reinforce Racism. NYU Press.
  39. 39.Offert, F.; and Bell, P. 2020. Perceptual bias and technical metapictures: critical machine vision as a humanities challenge. In AI & Society.
  40. 40.Otterbacher, J. 2018. Social Cues, Social Biases: Stereotypes in Annotations on People Images. In Proceedings of the AAAI Conference on Human Computation and Crowdsourcing.
  41. 41.Otterbacher, J.; Barlas, P.; Kleanthous, S.; and Kyriakou, K. 2019. How Do We Talk about Other People? Group (Un)Fairness in Natural Language Image Descriptions. In Proceedings of the AAAI Conference on Human Computation and Crowdsourcing.
  42. 42.Pilipets, E.; and Paasonen, S. 2020. Nipples, memes, and algorithmic failure: NSFW critique of Tumblr censorship. In New Media & Society.
  43. 43.Prabhu, V. U.; and Birhane, A. 2020. Large image datasets: A pyrrhic win for computer vision? arXiv:2006.16923.
  44. 44.Rhue, L. 2018. Racial Influence on Automated Perceptions of Emotions. SSRN Electronic Journal.
  45. 45.Roberts, C. V. 2018. Quantifying the Extent to which Popular Pre-Trained Convolutional Neural Networks Implicitly Learn High-Level Protected Attributes. In Princeton Masters Thesis.
  46. 46.Rodden, K. 2017. Is that a boy or a girl? Exploring a neural network’s construction of gender. Medium.
  47. 47.Scheuerman, M. K.; Paul, J. M.; and Brubaker, J. R. 2019. How Computers See Gender: An Evaluation of Gender Classification in Commercial Facial Analysis Services. Proc. ACM Hum.-Comput. Interact., 3(CSCW).
  48. 48.Scheuerman, M. K.; Wade, K.; Lustig, C.; and Brubaker, J. R. 2020. How We’ve Taught Algorithms to See Identity: Constructing Race and Gender in Image Databases for Facial Analysis. Proc. ACM Hum.-Comput. Interact., 4(CSCW1).
  49. 49.Schwemmer, C.; Knight, C.; Bello-Pardo, E. D.; Oklobdzija, S.; Schoonvelde, M.; and Lockhart, J. W. 2020. Diagnosing gender bias in image recognition systems. Socius, 6: 2378023120967171.
  50. 50.Sedenberg, E.; and Chuang, J. 2017. Smile for the Camera: Privacy and Policy Implications of Emotion AI. arXiv:1709.00396.
  51. 51.Selbst, A.; and Barocas, S. 2023. Unfair Artificial Intelligence: How FTC Intervention Can Overcome the Limitations of Discrimination Law. University of Pennsylvania Law Review, 171.
  52. 52.Shankar, S.; Halpern, Y.; Breck, E.; Atwood, J.; Wilson, J.; and Sculley, D. 2017a. No Classification without Representation: Assessing Geodiversity Issues in Open Data Sets for the Developing World. arXiv:1711.08536.
  53. 53.Shankar, S.; Halpern, Y.; Breck, E.; Atwood, J.; Wilson, J.; and Sculley, D. 2017b. No classification without representation: Assessing geodiversity issues in open data sets for the developing world. In Proceedings of the NIPS 2017 Workshop on Machine Learning for the Developing World.
  54. 54.Simonite, T. 2018. When It Comes to Gorillas, Google Photos Remains Blind. Wired, January.
  55. 55.Singh, K. K.; Mahajan, D.; Grauman, K.; Lee, Y. J.; Feiszli, M.; and Ghadiyaram, D. 2020. Don’t Judge an Object by Its Context: Learning to Overcome Contextual Bias. arXiv:2001.03152.
  56. 56.Song, C.; and Shmatikov, V. 2020. Overlearning Reveals Sensitive Attributes. arXiv:1905.11742.
  57. 57.Steed, R.; and Caliskan, A. 2021. Image Representations Learned With Unsupervised Pre-Training Contain Human-like Biases. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’21, 701–713. New York, NY, USA: Association for Computing Machinery. ISBN 9781450383097.
  58. 58.Stock, P.; and Cisse, M. 2018. ConvNets and ImageNet Beyond Accuracy: Understanding Mistakes and Uncovering Biases. arXiv:1711.11443.
  59. 59.Wang, A.; Barocas, S.; Laird, K.; and Wallach, H. 2022. Measuring Representational Harms in Image Captioning. ACM Conference on Fairness, Accountability, and Transparency (FAccT).
  60. 60.Wang, A.; Liu, A.; Zhang, R.; Kleiman, A.; Kim, L.; Zhao, D.; Shirai, I.; Narayanan, A.; and Russakovsky, O. 2021. REVISE: A Tool for Measuring and Mitigating Bias in Visual Datasets. arXiv:2004.07999.
  61. 61.Wang, A.; and Russakovsky, O. 2021. Directional Bias Amplification. arXiv:2102.12594.
  62. 62.Wang, T.; Zhao, J.; Yatskar, M.; Chang, K.-W.; and Ordonez, V. 2019. Balanced Datasets Are Not Enough: Estimating and Mitigating Gender Bias in Deep Image Representations. arXiv:1811.08489.
  63. 63.Wang, Z.; Qinami, K.; Karakozis, I. C.; Genova, K.; Nair, P.; Hata, K.; and Russakovsky, O. 2020. Towards Fairness in Visual Recognition: Effective Strategies for Bias Mitigation. arXiv:1911.11834.
  64. 64.Wilson, B.; Hoffman, J.; and Morgenstern, J. 2019. Predictive Inequity in Object Detection. arXiv:1902.11097.
  65. 65.Wu, S.; Wieland, J.; Farivar, O.; and Schiller, J. 2017. Automatic Alt-text: Computer-generated Image Descriptions for Blind Users on a Social Network Service. ACM Conference on Computer-Supported Cooperative Work And Social Computing (CSCW).
  66. 66.Yang, K.; Qinami, K.; Fei-Fei, L.; Deng, J.; and Russakovsky, O. 2020. Towards fairer datasets. Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency.
  67. 67.Zhao, J.; Wang, T.; Yatskar, M.; Ordonez, V.; and Chang, K.-W. 2017. Men also like shopping: Reducing gender bias amplification using corpus-level constraints. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), 2979–2989.

Citation

MLA
Katzman, J., et al. “Taxonomizing and Measuring Representational Harms: A Look at Image Tagging”. Thirty-Seventh AAAI Conference on Artificial Intelligence (AAAI 2023), 2023, http://arxiv.org/abs/2305.01776v1.
APA
Katzman, J., Wang, A., Scheuerman, M., Blodgett, S. L., Laird, K., Wallach, H., & Barocas, S. (2023). Taxonomizing and Measuring Representational Harms: A Look at Image Tagging. Thirty-Seventh AAAI Conference on Artificial Intelligence (AAAI 2023). http://arxiv.org/abs/2305.01776v1
Chicago
Katzman, J., A. Wang, M. Scheuerman, et al. 2023. “Taxonomizing and Measuring Representational Harms: A Look at Image Tagging”. Thirty-Seventh AAAI Conference on Artificial Intelligence (AAAI 2023). http://arxiv.org/abs/2305.01776v1.
Harvard
Katzman, J. et al. (2023) “Taxonomizing and Measuring Representational Harms: A Look at Image Tagging”, Thirty-Seventh AAAI Conference on Artificial Intelligence (AAAI 2023) [Preprint]. Available at: http://arxiv.org/abs/2305.01776v1.
Vancouver
1. Katzman J, Wang A, Scheuerman M, Blodgett SL, Laird K, Wallach H, Barocas S (2023) Taxonomizing and Measuring Representational Harms: A Look at Image Tagging. Thirty-Seventh AAAI Conference on Artificial Intelligence (AAAI 2023)

BibTeX

@article{katzman2023taxonomizing,
  title = {Taxonomizing and Measuring Representational Harms: A Look at Image Tagging},
  author = {Katzman, Jared and Wang, Angelina and Scheuerman, Morgan and Blodgett, Su Lin and Laird, Kristen and Wallach, Hanna and Barocas, Solon},
  year = {2023},
  journal = {Thirty-Seventh AAAI Conference on Artificial Intelligence (AAAI 2023)},
  url = {http://arxiv.org/abs/2305.01776v1},
  eprint = {2305.01776}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF