Built independently by an author, for readers. Read the story and support ChapterPal

keyword

representational harms

Representational harms refer to negative impacts produced by computational and artificial intelligence systems that degrade the social standing, visibility, or perception of specific individuals or demographic groups. Unlike allocative harms, which involve the unfair distribution or denial of tangible resources and opportunities such as employment, loans, or medical care, representational harms operate at a cultural and symbolic level. They occur when algorithms produce outputs that reinforce harmful stereotypes, generate demeaning or toxic associations, misclassify individuals based on sensitive attributes, or minimize and erase marginalized communities from public view. Because these harms influence how people perceive themselves and others, they can perpetuate systemic prejudice and long-term societal inequality across applications such as natural language processing, computer vision, and information retrieval.

3 items

Magic, Madness, Heaven, Sin: LLM Output Diversity is Everything, Everywhere, All at Once

Magic, Madness, Heaven, Sin: LLM Output Diversity is Everything, Everywhere, All at Once

Harnoor Dhingra

OrganizationsMicrosoft

Why you should read this

Proposes a unified evaluation framework for large language model output diversity across four normative contexts, exposing critical trade-offs where optimizing for safety or factuality undermines demographic representation and creative utility.

Research on Large Language Models (LLMs) studies output variation across generation, reasoning, alignment, and representational analysis, often under the umbrella of "diversity." Yet the terminology remains fragmented, largely because the normative objectives underlying tasks are rarely made explicit. We introduce the Magic, Madness, Heaven, Sin framework, which models output variation along a homogeneity-heterogeneity axis, where valuation is determined by the task and its normative objective. We organize tasks into four normative contexts: epistemic (factuality), interactional (user utility), societal (representation), and safety (robustness). For each, we examine the failure modes and vocabulary such as hallucination, mode collapse, bias, and erasure through which variation is studied. We apply the framework to analyze all pairwise cross-contextual interactions, revealing that optimizing for one objective, such as improving safety, can inadvertently harm demographic representation or creative diversity. We argue for context-aware evaluation of output variation, reframing it as a property shaped by task objectives rather than a model's intrinsic trait.

Added

2026-09-30

Taxonomizing and Measuring Representational Harms: A Look at Image Tagging

Taxonomizing and Measuring Representational Harms: A Look at Image Tagging

Jared Katzman, Angelina Wang, Morgan Klaus Scheuerman, Su Lin Blodgett, Kristen Laird, Hanna M. Wallach, Solon Barocas

OrganizationsMicrosoftPrinceton UniversityUniversity of Colorado BoulderUniversity of Michigan

Why you should read this

Categorizes computational fairness metrics and representational harms in image tagging systems to demonstrate that standard measurement approaches fail to uniquely map to specific harms and that mitigating one harm can inadvertently worsen another.

In this paper, we examine computational approaches for measuring the “fairness” of image tagging systems, finding that they cluster into five distinct categories, each with its own analytic foundation. We also identify a range of normative concerns that are often collapsed under the terms “unfairness,” “bias,” or even “discrimination” when discussing problematic cases of image tagging. Specifically, we identify four types of representational harms that can be caused by image tagging systems, providing concrete examples of each. We then consider how different computational measurement approaches map to each of these types, demonstrating that there is not a one-to-one mapping. Our findings emphasize that no single measurement approach will be definitive and that it is not possible to infer from the use of a particular measurement approach which type of harm was intended to be measured. Lastly, equipped with this more granular understanding of the types of representational harms that can be caused by image tagging systems, we show that attempts to mitigate some of these types of harms may be in tension with one another.

Added

2026-09-26

Ethical and social risks of harm from Language Models

Ethical and social risks of harm from Language Models

Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, Zac Kenton, Sasha Brown, Will Hawkins, Tom Stepleton, Courtney Biles, Abeba Birhane, Julia Haas, Laura Rimell, Lisa Anne Hendricks, William Isaac, Sean Legassick, Geoffrey Irving, Iason Gabriel

OrganizationsCalifornia Institute of TechnologyGoogleUniversity College DublinUniversity of Toronto

Why you should read this

Systematically organizes the scattered landscape of language model risks into a structured taxonomy that makes potential harms identifiable and actionable for developers, researchers, and policymakers working to build safer AI systems.

This paper aims to help structure the risk landscape associated with large-scale Language Models (LMs). In order to foster advances in responsible innovation, an in-depth understanding of the potential risks posed by these models is needed. A wide range of established and anticipated risks are analysed in detail, drawing on multidisciplinary expertise and literature from computer science, linguistics, and social sciences. We outline six specific risk areas: I. Discrimination, Exclusion and Toxicity, II. Information Hazards, III. Misinformation Harms, V. Malicious Uses, V. Human-Computer Interaction Harms, VI. Automation, Access, and Environmental Harms. The first area concerns the perpetuation of stereotypes, unfair discrimination, exclusionary norms, toxic language, and lower performance by social group for LMs. The second focuses on risks from private data leaks or LMs correctly inferring sensitive information. The third addresses risks arising from poor, false or misleading information including in sensitive domains, and knock-on risks such as the erosion of trust in shared information. The fourth considers risks from actors who try to use LMs to cause harm. The fifth focuses on risks specific to LLMs used to underpin conversational agents that interact with human users, including unsafe use, manipulation or deception. The sixth discusses the risk of environmental harm, job automation, and other challenges that may have a disparate effect on different social groups or communities. In total, we review 21 risks in-depth. We discuss the points of origin of different risks and point to potential mitigation approaches. Lastly, we discuss organisational responsibilities in implementing mitigations, and the role of collaboration and participation. We highlight directions for further research, particularly on expanding the toolkit for assessing and evaluating the outlined risks in LMs.

Added

2026-02-21