Magic, Madness, Heaven, Sin: LLM Output Diversity is Everything, Everywhere, All at Once

Harnoor Dhingra

article2026arXiv0 citations

Proposes a unified evaluation framework for large language model output diversity across four normative contexts, exposing critical trade-offs where optimizing for safety or factuality undermines demographic representation and creative utility.

Listen

Large language models are deployed across a wide range of applications, from medical question answering and creative writing to high-stakes compliance and automated customer support. However, evaluating model output variation has remained fractured across siloed domains such as fairness, safety, alignment, and natural language generation, leading to fragmented terminology where output variation is alternately praised or penalized without a shared conceptual baseline. The article addresses this fragmentation by establishing a unified lens to evaluate when output variation is beneficial versus when it represents system failure.

The main objective of the article is to introduce the Magic, Madness, Heaven, Sin framework, which demonstrates that large language model output variation is not an intrinsic model trait but a context-dependent property evaluated along a continuous axis of homogeneity and heterogeneity. The article evaluates how task-specific normative goals determine whether variation is rewarded or penalized, and systematically maps the structural trade-offs that occur when optimizing models across competing objectives.

To construct this framework, the article synthesizes findings across contemporary machine learning literature, spanning research in alignment, representation, natural language generation, and safety benchmarking. It organizes tasks into four distinct normative contexts: epistemic, interactional, societal, and safety. Using this conceptual taxonomy, the article conducts a systematic pairwise analysis across all six cross-contextual interactions to identify how interventions targeted at one objective affect the others.

The analysis reveals several core findings. First, output valuation directly depends on context: heterogeneity is rewarded as creative utility in interactional contexts ("Magic") but penalized as hallucination in epistemic settings ("Madness"), while homogeneity is rewarded as robustness in safety contexts ("Heaven") but penalized as erasure or stereotyping in societal contexts ("Sin"). Second, the article demonstrates that standard safety and epistemic alignment pipelines systematically drive output convergence, which directly induces mode collapse, suppresses creative diversity, and homogenizes demographic and cultural viewpoints. Third, models exhibit severe cultural and demographic skews, frequently defaulting to Western, English-speaking societal norms and rendering marginalized identities invisible. Fourth, user personalization creates structural trade-offs, where optimizing for individual preferences risks confining users to demographic stereotypes or ideological filter bubbles. Finally, the analysis shows that single real-world queries regularly activate multiple competing objectives simultaneously, requiring models to satisfy opposing demands for consistency and variation within the same response.

These findings indicate that treating diversity or consistency as universally positive attributes leads to flawed system designs, as optimizing solely for safety or factuality will inherently compromise creative range and equitable demographic representation. For enterprise leaders and system developers, these trade-offs directly impact compliance, brand consistency, user engagement, and fairness risks. Rather than pursuing one-size-fits-all model alignment, developers must implement context-aware deployment strategies that evaluate output distributions relative to specific task objectives.

Decision-makers should transition from monolithic alignment toward context-specific control mechanisms and modular guardrails. When deploying systems in high-stakes fields like finance or law, teams should prioritize deterministic compliance and factuality; conversely, creative and exploratory deployments require mechanisms that preserve output entropy and representation. Organizations should also evaluate multi-objective queries to balance factual boundaries with broad option generation. Because the article presents a conceptual framework and literature synthesis rather than a technical control mechanism, future work must focus on building and piloting dynamic steering architectures that can resolve these cross-contextual trade-offs in real time.

arXiv: 2604.01504
Cover for Magic, Madness, Heaven, Sin: LLM Output Diversity is Everything, Everywhere, All at Once

Abstract

Research on Large Language Models (LLMs) studies output variation across generation, reasoning, alignment, and representational analysis, often under the umbrella of "diversity." Yet the terminology remains fragmented, largely because the normative objectives underlying tasks are rarely made explicit. We introduce the Magic, Madness, Heaven, Sin framework, which models output variation along a homogeneity-heterogeneity axis, where valuation is determined by the task and its normative objective. We organize tasks into four normative contexts: epistemic (factuality), interactional (user utility), societal (representation), and safety (robustness). For each, we examine the failure modes and vocabulary such as hallucination, mode collapse, bias, and erasure through which variation is studied. We apply the framework to analyze all pairwise cross-contextual interactions, revealing that optimizing for one objective, such as improving safety, can inadvertently harm demographic representation or creative diversity. We argue for context-aware evaluation of output variation, reframing it as a property shaped by task objectives rather than a model's intrinsic trait.

Table of Contents

  • 1 Introduction
  • 2 Epistemic Context
  • 2.1 Task: Fact-based QA
  • 2.2 Reasoning Tasks
  • 3 Interactional Context
  • 3.1 Task: Creative Writing and Brainstorming
  • 3.2 Task: Open-Ended Dialogue
  • 3.3 Task: Open-Ended QA
  • 4 Societal Context
  • 4.1 Demographic Representation
  • 4.2 Cultural Representation
  • 4.3 Values and Politics
  • 5 Safety Context
  • 5.1 Refusal and Safe Completion
  • 5.2 Compliance
  • 6 Discussion
  • 6.1 Applications of framework
  • 6.2 Contemporary Frameworks and Scope
  • 7 Conclusion
  • References
  • A Disambiguating Diversity in Post-Training

Knowls

  1. Knowl 1 — Output variation is valued relative to a task’s normative objective

    definition

    The Magic, Madness, Heaven, Sin framework treats language-model output variation as a continuum from homogeneity (convergent, consistent outputs) to heterogeneity (divergent, varied outputs). Variation is not inherently good or bad: its valuation depends on the task’s normative objective. The framework distinguishes four dominant contexts: epistemic tasks seek factuality and penalize heterogeneity as hallucination or error (Madness); interactional tasks seek utility, engagement, and novelty and often reward heterogeneity as creativity (Magic); societal tasks seek fair representation and penalize homogeneity as erasure or stereotyping (Sin); and safety tasks seek robustness and reward homogeneity as consistent compliance (Heaven). These are normative categories, not numerical scores or exhaustive task types.

  2. Knowl 2 — Apply the framework by identifying objectives and their levels of analysis

    model/method

    To evaluate output variation in a task, identify its active normative objectives, then determine separately for each objective whether homogeneity or heterogeneity is desirable. Make that judgment at the relevant level of analysis: the same task can value convergence in one aspect of a response and variation in another. For example, a request for help managing severe chronic pain activates epistemic, safety, and interactional objectives. Medical facts and avoidance of dangerous advice call for consistency, while user utility can call for a varied set of treatment options. The framework therefore calls for satisfying competing objectives together, rather than assigning one overall value to the response’s diversity.

  3. Knowl 3 — Epistemic tasks penalize semantic divergence, not every form of variation

    model/method

    In factual question answering and reasoning-intensive tasks, the dominant objective is correctness and reliability. When a question has one correct answer or a small, well-defined answer set, divergence from that answer is an epistemic failure, commonly described as hallucination. In reasoning tasks, correctness can depend on the intermediate logical trajectory as well as the final result, so divergent reasoning paths that lead to inconsistent or incorrect solutions are a concern. However, surface variation can be acceptable: lexical, structural, or syntactic differences need not matter when the semantic answer or functional result remains correct, as with code variants that execute to the same result. Miscalibrated confidence is another risk; when knowledge is insufficient, uncertainty expression or abstention is preferable to an incorrect, confident answer.

  4. Knowl 4 — Creative generation can fail through convergence at several levels

    model/method

    Creative writing, brainstorming, and ideation often require heterogeneity: users seek novel, surprising, meaningfully distinct outputs, so predictable repetition is a form of homogenization. The paper’s synthesis describes convergence within a model across repeated samples (intra-model mode collapse) and across independently trained models (inter-model similarity). Variation in wording alone may not provide substantive novelty: generated stories can differ lexically while reusing similar plot elements. Homogenization is also reported in human–AI collaboration, where using language models as creativity tools can make different users’ ideas more similar and aligned-model co-writing can yield less semantic diversity than writing with base models.

  5. Knowl 5 — Interactional tasks require variation to be conditioned on users and answer structure

    model/method

    Interactional tasks do not have a single uniform preference for variation. In multi-turn dialogue, different users should be able to receive different conversational trajectories reflecting their intent and preferences, while a single user’s conversation should remain consistent with that user’s established preferences. In open-ended question answering, heterogeneity can represent coverage of multiple supported perspectives or interpretations; that coverage may appear across separately sampled answers or within one answer that presents several viewpoints. Ambiguous queries may require multiple plausible interpretations for completeness. Recommendations are an exception to a simple preference for divergence: personalization favors options suited to a user, but excessive narrowing can create filter-bubble effects, so useful recommendations must balance personalized consistency with novel options.

  6. Knowl 6 — Societal representational harm arises from output homogeneity

    model/method

    In the societal context, the objective is fair representation, and convergence on dominant defaults can erase variation in human populations. The paper describes two demographic mechanisms: erasure, in which underspecified prompts lead models to represent majority groups disproportionately or leave marginalized identities invisible; and stereotyping, in which represented groups are confined to reductive identity–role associations. Both reduce the complexity of represented lives. The same concern extends to culture and values: models are reported to perform unevenly across regions and languages, to give culturally different answers depending on query language, and to reflect Anglocentric or WEIRD (Western, Educated, Industrialized, Rich, Democratic) moral and political defaults. Under this objective, greater representational coverage and heterogeneity are desirable.

  7. Knowl 7 — Safety values consistent behavior through refusal and compliance

    model/method

    In the safety context, the objective is robustness: prescribed behavioral constraints should hold consistently across prompts, including adversarial paraphrases, role-play, translation, jailbreaks, and prompt injection. Homogeneity is therefore valued. Safety includes refusal, which excludes policy-prohibited outputs, and compliance, which requires adherence to external professional, legal, ethical, or other standards. Safe completion is a less rigid alternative to refusal: it avoids actionable harmful details while still trying to address the user’s underlying intent helpfully. In commercial and high-stakes settings, reproducibility, auditability, and consistent answers can also be treated as safety or compliance requirements.

  8. Knowl 8 — Pairwise cross-context analysis exposes six distinct alignment tensions

    model/method

    The paper’s system-level analysis groups the six pairs of normative contexts into three kinds of tension. These are structural conflicts illustrated through prior findings, not estimates of a single causal effect.

    • Opposing preferred behaviors: Safety and interactional objectives pull in opposite directions because safety favors convergence while creative interaction often rewards divergence; the paper notes reports associating alignment and post-training with creative homogenization. Societal and epistemic objectives can also conflict: factual consistency may suppress culturally relevant variation, while omitting demographic attributes to avoid stereotypes can degrade medical diagnostic accuracy when those attributes are clinically relevant.
    • Same behavior, different valuations: Safety rewards homogeneity as robust alignment, whereas societal objectives may penalize it as representational erasure. The paper cites reports that alignment can compress demographic or cultural diversity and that safety reward models can penalize non-standard dialects such as African American English. Interactional and epistemic objectives can both concern heterogeneity but value it oppositely: creative novelty may be useful interactionally but unreliable epistemically. Sycophancy can cause a model to abandon correct answers in dialogue, while suppressing uncertainty to improve factuality may reduce creative diversity.
    • Same broad valuation, different objectives: Societal and interactional tasks can both favor heterogeneity, but societal heterogeneity means coverage across groups, whereas interactional heterogeneity means novelty or user-specific adaptation. Personalization can promote within-user homogeneity and may reinforce stereotypes. Safety and epistemic tasks can both favor homogeneity, respectively for robust behavior and factual correctness; optimizing these objectives together can leave societal and interactional goals disadvantaged.
  9. Knowl 9 — Post-training effects depend on what kind of diversity is measured

    data/table

    The paper’s comparison of prior post-training studies shows that “diversity” does not name a single outcome, and reported effects vary by task, diversity dimension, and whether variation is measured across or within responses. In summarization, RLHF is reported to improve out-of-distribution generalization at the expense of output diversity. In five natural-language-generation tasks, instruction tuning is reported to increase lexical diversity while reducing syntactic and semantic diversity relative to base models. In open-ended code generation, preference-tuned models are reported to have lower lexical and syntactic diversity but greater quality-controlled effective semantic diversity than base or supervised-fine-tuned models, largely because they produce more valid outputs. For open-ended question answering, alignment is reported to reduce diversity across responses while increasing pluralism within a response by combining perspectives. Other cited results report reduced coverage of valid outputs after post-training and reduced conceptual diversity across simulated personas after RLHF or RLAIF. These contrasts support evaluating variation with task-appropriate dimensions rather than treating a single diversity measure as decisive.

  10. Knowl 10 — The framework is normative and does not prescribe how to resolve conflicts

    limitation

    The four contexts are illustrative rather than exhaustive. The framework classifies how variation is valued but abstracts away from mechanisms that produce or change it, including sampling strategies, training stages, and prompting techniques. It is also agnostic to the specific metric and level of analysis, leaving the disentangling of those levels and measures outside its scope. Although the framework makes conflicts between objectives explicit, it does not determine how to resolve them or provide context-aware control mechanisms.

Coverage note — The paper’s broader reference-by-reference discussion of task vocabularies and individual diversity metrics is not reproduced beyond the distinctions needed to explain the framework; it synthesizes prior literature rather than adding a separate method or result.

References

  1. 1.Mohammad Abdollahi, Khandaker Rifah Tasnia, Soumit Kanti Saha, Jinqiu Yang, Song Wang, and Hadi Hemmati. Demystifying errors in llm reasoning traces: An empirical study of code execution simulation, 2025. URL https://arxiv.org/abs/2512.00215.
  2. 2.Muhammad Farid Adilazuarda, Sagnik Mukherjee, Pradhyumna Lavania, Siddhant Shivdutt Singh, Alham Fikri Aji, Jacki O’Neill, Ashutosh Modi, and Monojit Choudhury. Towards measuring and modeling “culture” in LLMs: A survey. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (eds.), Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp. 15763–15784, Miami, Florida, USA, November 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.emnlp-main.882. URL https://aclanthology.org/2024.emnlp-main.882/.
  3. 3.Favour Y. Aghaebe, Elizabeth A Williams, Tanefa Apekey, and Nafise Sadat Moosavi. LLMs do not see age: Assessing demographic bias in automated systematic review synthesis. In Kentaro Inui, Sakriani Sakti, Haofen Wang, Derek F. Wong, Pushpak Bhattacharyya, Biplab Banerjee, Asif Ekbal, Tanmoy Chakraborty, and Dhirendra Pratap Singh (eds.), Proceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics, pp. 1815–1833, Mumbai, India, December 2025. The Asian Federation of Natural Language Processing and The Association for Computational Linguistics. ISBN 979-8-89176-298-5. doi: 10.18653/v1/2025.ijcnlp-long.98. URL https://aclanthology.org/2025.ijcnlp-long.98/.
  4. 4.Dalia Ali, Dora Zhao, Allison Koenecke, and Orestis Papakyriakopoulos. Operationalizing pluralistic values in large language model alignment reveals trade-offs in safety, inclusivity, and model behavior, 2025. URL https://arxiv.org/abs/2511.14476.
  5. 5.Mohammad Aliannejadi, Julia Kiseleva, Aleksandr Chuklin, Jeff Dalton, and Mikhail Burtsev. Building and evaluating open-domain dialogue corpora with clarifying questions. In Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih (eds.), Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 4473–4484, Online and Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.emnlp-main.367. URL https://aclanthology.org/2021.emnlp-main.367/.
  6. 6.Badr AlKhamissi, Muhammad ElNokrashy, Mai Alkhamissi, and Mona Diab. Investigating cultural alignment of large language models. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar (eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 12404–12422, Bangkok, Thailand, August 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.acl-long.671. URL https://aclanthology.org/2024.acl-long.671/.
  7. 7.Barrett R Anderson, Jash Hemant Shah, and Max Kreminski. Homogenization effects of large language models on human creative ideation. In Proceedings of the 16th Conference on Creativity and Cognition, pp. 413–425, New York, NY, USA, 2024. Association for Computing Machinery. ISBN 9798400704857. doi: 10.1145/3635636.3656204. URL https://doi.org/10.1145/3635636.3656204.
  8. 8.Qazi Mohammad Areeb, Mohammad Nadeem, Shahab Saquib Sohail, Raza Imam, Faiyaz Doctor, Yassine Himeur, Amir Hussain, and Abbes Amira. Filter bubbles in recommender systems: Fact or fallacy – a systematic review, 2023. URL https://arxiv.org/abs/2307.01221.
  9. 9.Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Ben Mann, Nova DasSarma, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Jackson Kernion, Kamal Ndousse, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, and Jared Kaplan. A general language assistant as a laboratory for alignment, 2021. URL https://arxiv.org/abs/2112.00861.
  10. 10.Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Chen, Catherine Olsson, Christopher Olah, Danny Hernandez, Dawn Drain, Deep Ganguli, Dustin Li, Eli Tran-Johnson, Ethan Perez, Jamie Kerr, Jared Mueller, Jeffrey Ladish, Joshua Landau, Kamal Ndousse, Kamile Lukosuite, Liane Lovitt, Michael Sellitto, Nelson Elhage, Nicholas Schiefer, Noemi Mercado, Nova DasSarma, Robert Lasenby, Robin Larson, Sam Ringer, Scott Johnston, Shauna Kravec, Sheer El Showk, Stanislav Fort, Tamera Lanham, Timothy Telleen-Lawton, Tom Conerly, Tom Henighan, Tristan Hume, Samuel R. Bowman, Zac Hatfield-Dodds, Ben Mann, Dario Amodei, Nicholas Joseph, Sam McCandlish, Tom Brown, and Jared Kaplan. Constitutional ai: Harmlessness from ai feedback, 2022. URL https://arxiv.org/abs/2212.08073.
  11. 11.Solon Barocas, Moritz Hardt, and Arvind Narayanan. Fairness and machine learning limitations and opportunities. 2018. URL https://api.semanticscholar.org/CorpusID:113402716.
  12. 12.Noam Benkler, Drisana Mosaphir, Scott Friedman, Andrew Smart, and Sonja Schmer-Galunder. Assessing llms for moral value pluralism, 2023. URL https://arxiv.org/abs/2312.10075.
  13. 13.Su Lin Blodgett. Sociolinguistically Driven Approaches for Just Natural Language Processing. PhD thesis, University of Massachusetts Amherst, 2021. URL https://scholarworks.umass.edu/dissertations 2/2251. PhD Dissertation.
  14. 14.Su Lin Blodgett, Solon Barocas, Hal Daume III, and Hanna Wallach. Language (technology) is power: A critical survey of “bias” in NLP. In Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (eds.), Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 5454–5476, Online, July 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.acl-main.485. URL https://aclanthology.org/2020.acl-main.485/.
  15. 15.Yong Cao, Li Zhou, Seolhwa Lee, Laura Cabello, Min Chen, and Daniel Hershcovich. Assessing cross-cultural alignment between ChatGPT and human societies: An empirical study. In Sunipa Dev, Vinodkumar Prabhakaran, David Ifeoluwa Adelani, Dirk Hovy, and Luciana Benotti (eds.), Proceedings of the First Workshop on Cross-Cultural Considerations in NLP (C3NLP), pp. 53–67, Dubrovnik, Croatia, May 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.c3nlp-1.7. URL https://aclanthology.org/2023.c3nlp-1.7/.
  16. 16.Sihao Chen, Daniel Khashabi, Wenpeng Yin, Chris Callison-Burch, and Dan Roth. Seeing things from a different angle:discovering diverse perspectives about claims. In Jill Burstein, Christy Doran, and Thamar Solorio (eds.), Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 542–557, Minneapolis, Minnesota, June 2019. Association for Computational Linguistics. doi: 10.18653/v1/N19-1053. URL https://aclanthology.org/N19-1053/.
  17. 17.Myra Cheng, Esin Durmus, and Dan Jurafsky. Marked personas: Using natural language prompts to measure stereotypes in language models. In Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (eds.), Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 1504–1532, Toronto, Canada, July 2023a. Association for Computational Linguistics. doi: 10.18653/v1/2023.acl-long.84. URL https://aclanthology.org/2023.acl-long.84/.
  18. 18.Myra Cheng, Sunny Yu, Cinoo Lee, Pranav Khadpe, Lujain Ibrahim, and Dan Jurafsky. Elephant: Measuring and understanding social sycophancy in llms, 2025. URL https://arxiv.org/abs/2505.13995.
  19. 19.Shu-Li Cheng, Shih-Jen Tsai, Ya-Mei Bai, Chih-Hung Ko, Chih-Wei Hsu, Fu-Chi Yang, Chia-Kuang Tsai, Yu-Kang Tu, Szu-Nian Yang, Ping-Tao Tseng, Tien-Wei Hsu, Chih-Sung Liang, and Kuan-Pin Su. Comparisons of quality, correctness, and similarity between chatgpt-generated and human-written abstracts for basic research: Cross-sectional study. J Med Internet Res, 25:e51229, Dec 2023b. ISSN 1438-8871. doi: 10.2196/51229. URL https://www.jmir.org/2023/1/e51229.
  20. 20.Paul Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei. Deep reinforcement learning from human preferences, 2023. URL https://arxiv.org/abs/1706.03741.
  21. 21.Kenneth Ward Church and William A. Gale. Poisson mixtures. Natural Language Engineering, 1:163 – 190, 1995. URL https://api.semanticscholar.org/CorpusID:8121803.
  22. 22.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training verifiers to solve math word problems, 2021. URL https://arxiv.org/abs/2110.14168.
  23. 23.Sunipa Dev, Masoud Monajatipoor, Anaelia Ovalle, Arjun Subramonian, Jeff Phillips, and Kai-Wei Chang. Harms of gender exclusivity and challenges in non-binary representation in language technologies. In Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih (eds.), Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 1968–1994, Online and Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.emnlp-main.150. URL https://aclanthology.org/2021.emnlp-main.150/.
  24. 24.Sherwin Dewan, Ismail Shaikh, Connie Shaw, Abhilash Sahoo, Akshita Jha, and Alisha Pradhan. Examining age-bias and stereotypes of aging in llms. In Proceedings of the 27th International ACM SIGACCESS Conference on Computers and Accessibility, ASSETS ’25, New York, NY, USA, 2025. Association for Computing Machinery. ISBN 9798400706769. doi: 10.1145/3663547.3746464. URL https://doi.org/10.1145/3663547.3746464.
  25. 25.Harnoor Dhingra, Preetiha Jayashanker, Sayali Moghe, and Emma Strubell. Queer people are people first: Deconstructing sexual identity stereotypes in large language models, 2023. URL https://arxiv.org/abs/2307.00101.
  26. 26.Yi Dong, Ronghui Mu, Gaojie Jin, Yi Qi, Jinwei Hu, Xingyu Zhao, Jie Meng, Wenjie Ruan, and Xiaowei Huang. Building guardrails for large language models, 2024. URL https://arxiv.org/abs/2402.01822.
  27. 27.Berin Doru, Christoph Maier, Johanna Sophie Busse, Thomas Lucke, Judith Sch önhoff, Elena Enax-Krumova, Steffen Hessler, Maria Berger, and Marianne Tokic. Detecting artificial intelligence–generated versus human-written medical student essays: Semirandomized controlled study. JMIR Med Educ, 11:e62779, Mar 2025. ISSN 2369-3762. doi: 10.2196/62779. URL https://mededu.jmir.org/2025/1/e62779.
  28. 28.Louis Esteve, Marie-Catherine de Marneffe, Nurit Melnik, Agata Savary, and Olha Kanishcheva. A survey of diversity quantification in natural language processing: The why, what, where and how, 2026. URL https://arxiv.org/abs/2507.20858.
  29. 29.Shangbin Feng, Chan Young Park, Yuhan Liu, and Yulia Tsvetkov. From pretraining data to language models to downstream tasks: Tracking the trails of political biases leading to unfair NLP models. In Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (eds.), Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 11737–11762, Toronto, Canada, July 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.acl-long.656. URL https://aclanthology.org/2023.acl-long.656/.
  30. 30.Isabel O. Gallegos, Ryan A. Rossi, Joe Barrow, Md Mehrab Tanjim, Sungchul Kim, Franck Dernoncourt, Tong Yu, Ruiyi Zhang, and Nesreen K. Ahmed. Bias and fairness in large language models: A survey. Computational Linguistics, 50(3):1097–1179, September 2024. doi: 10.1162/coli a 00524. URL https://aclanthology.org/2024.cl-3.8/.
  31. 31.Cuiyun Gao, Guodong Fan, Chun Yong Chong, Shizhan Chen, Chao Liu, David Lo, Zibin Zheng, and Qing Liao. A systematic literature review of code hallucinations in llms: Characterization, mitigation methods, challenges, and future directions for reliable ai, 2025. URL https://arxiv.org/abs/2511.00776.
  32. 32.Yanzhu Guo, Guokan Shang, and Chloe Clavel. Benchmarking linguistic diversity of large language models, 2025. URL https://arxiv.org/abs/2412.10271.
  33. 33.Hilda Hadan, Derrick M. Wang, Reza Hadi Mogavi, Joseph Tu, Leah Zhang-Kennedy, and Lennart E. Nacke. The great ai witch hunt: Reviewers’ perception and (mis)conception of generative ai in research writing. Computers in Human Behavior: Artificial Humans, 2(2):100095, 2024. ISSN 2949-8821. doi: https://doi.org/10.1016/j.chbah.2024.100095. URL https://www.sciencedirect.com/science/article/pii/S2949882124000550.
  34. 34.Saad Hassan, Matt Huenerfauth, and Cecilia Ovesdotter Alm. Unpacking the interdependent systems of discrimination: Ableist bias in NLP systems through an intersectional lens. In Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih (eds.), Findings of the Association for Computational Linguistics: EMNLP 2021, pp. 3116–3123, Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.findings-emnlp.267. URL https://aclanthology.org/2021.findings-emnlp.267/.
  35. 35.Shirley Anugrah Hayati, Minhwa Lee, Dheeraj Rajagopal, and Dongyeop Kang. How far can we extract diverse perspectives from large language models? In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (eds.), Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp. 5336–5366, Miami, Florida, USA, November 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.emnlp-main.306. URL https://aclanthology.org/2024.emnlp-main.306/.
  36. 36.Jiseung Hong, Grace Byun, Seungone Kim, and Kai Shu. Measuring sycophancy of language models in multi-turn dialogues. In Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, and Violet Peng (eds.), Findings of the Association for Computational Linguistics: EMNLP 2025, pp. 2239–2259, Suzhou, China, November 2025. Association for Computational Linguistics. ISBN 979-8-89176-335-7. doi: 10.18653/v1/2025.findings-emnlp.121. URL https://aclanthology.org/2025.findings-emnlp.121/.
  37. 37.Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, and Ting Liu. A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. ACM Transactions on Information Systems, 43(2):1–55, January 2025a. ISSN 1558-2868. doi: 10.1145/3703155. URL http://dx.doi.org/10.1145/3703155.
  38. 38.Yue Huang, Chujie Gao, Siyuan Wu, Haoran Wang, Xiangqi Wang, Yujun Zhou, Yanbo Wang, Jiayi Ye, Jiawen Shi, Qihui Zhang, Yuan Li, Han Bao, Zhaoyi Liu, Tianrui Guan, Dongping Chen, Ruoxi Chen, Kehan Guo, Andy Zou, Bryan Hooi Kuen-Yew, Caiming Xiong, Elias Stengel-Eskin, Hongyang Zhang, Hongzhi Yin, Huan Zhang, Huaxiu Yao, Jaehong Yoon, Jieyu Zhang, Kai Shu, Kaijie Zhu, Ranjay Krishna, Swabha Swayamdipta, Taiwei Shi, Weijia Shi, Xiang Li, Yiwei Li, Yuexing Hao, Zhihao Jia, Zhize Li, Xiuying Chen, Zhengzhong Tu, Xiyang Hu, Tianyi Zhou, Jieyu Zhao, Lichao Sun, Furong Huang, Or Cohen Sasson, Prasanna Sattigeri, Anka Reuel, Max Lamparth, Yue Zhao, Nouha Dziri, Yu Su, Huan Sun, Heng Ji, Chaowei Xiao, Mohit Bansal, Nitesh V. Chawla, Jian Pei, Jianfeng Gao, Michael Backes, Philip S. Yu, Neil Zhenqiang Gong, Pin-Yu Chen, Bo Li, Dawn Song, and Xiangliang Zhang. On the trustworthiness of generative foundation models: Guideline, assessment, and perspective, 2025b. URL https://arxiv.org/abs/2502.14296.
  39. 39.Zheng Hui, Yijiang River Dong, Ehsan Shareghi, and Nigel Collier. Trident: Benchmarking llm safety in finance, medicine, and law, 2025. URL https://arxiv.org/abs/2507.21134.
  40. 40.Frederick Jelinek, Robert L. Mercer, Lalit R. Bahl, and Janet M. Baker. Perplexity—a measure of the difficulty of speech recognition tasks. Journal of the Acoustical Society of America, 62, 1977. URL https://api.semanticscholar.org/CorpusID:121680873.
  41. 41.Jiabao Ji, Min Li, Priyanshu Kumar, Shiyu Chang, and Saloni Potdar. Deepambigqa: Ambiguous multi-hop questions for benchmarking llm answer completeness, 2025. URL https://arxiv.org/abs/2511.01323.
  42. 42.Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12):1–38, March 2023. ISSN 1557-7341. doi: 10.1145/3571730. URL http://dx.doi.org/10.1145/3571730.
  43. 43.Liwei Jiang, Yuanjun Chai, Margaret Li, Mickel Liu, Raymond Fok, Nouha Dziri, Yulia Tsvetkov, Maarten Sap, Alon Albalak, and Yejin Choi. Artificial hivemind: The open-ended homogeneity of language models (and beyond), 2025. URL https://arxiv.org/abs/2510.22954.
  44. 44.Rebecca L Johnson, Giada Pistilli, Natalia Menedez-González, Leslye Denisse Dias Duran, Enrico Panai, Julija Kalpokiene, and Donald Jay Bertulfo. The ghost in the machine has an american accent: value conflict in gpt-3, 2022. URL https://arxiv.org/abs/2203.07785.
  45. 45.Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, Scott Johnston, Sheer El-Showk, Andy Jones, Nelson Elhage, Tristan Hume, Anna Chen, Yuntao Bai, Sam Bowman, Stanislav Fort, Deep Ganguli, Danny Hernandez, Josh Jacobson, Jackson Kernion, Shauna Kravec, Liane Lovitt, Kamal Ndousse, Catherine Olsson, Sam Ringer, Dario Amodei, Tom Brown, Jack Clark, Nicholas Joseph, Ben Mann, Sam McCandlish, Chris Olah, and Jared Kaplan. Language models (mostly) know what they know, 2022. URL https://arxiv.org/abs/2207.05221.
  46. 46.Anjali Kantharuban, Jeremiah Milbauer, Maarten Sap, Emma Strubell, and Graham Neubig. Stereotype or personalization? user identity biases chatbot recommendations. In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar (eds.), Findings of the Association for Computational Linguistics: ACL 2025, pp. 24418–24436, Vienna, Austria, July 2025. Association for Computational Linguistics. ISBN 979-8-89176-256-5. doi: 10.18653/v1/2025.findings-acl.1254. URL https://aclanthology.org/2025.findings-acl.1254/.
  47. 47.Gautam Siddharth Kashyap, Mark Dras, and Usman Naseem. Too helpful, too harmless, too honest or just right? In Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, and Violet Peng (eds.), Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pp. 29723–29734, Suzhou, China, November 2025. Association for Computational Linguistics. ISBN 979-8-89176-332-6. doi: 10.18653/v1/2025.emnlp-main.1510. URL https://aclanthology.org/2025.emnlp-main.1510/.
  48. 48.Joshua Kelsall, Xingwei Tan, Aislinn Bergin, Jiahong Chen, Maria Waheed, Tom Sorell, Rob Procter, Maria Liakata, Jenny Chim, and Serene Chi. A rapid evidence review of evaluation techniques for large language models in legal use cases: trends, gaps, and recommendations for future research. AI and Society, 2025. ISSN 1435-5655. doi: 10.1007/s00146-025-02741-9. URL https://doi.org/10.1007/s00146-025-02741-9.
  49. 49.Saurabh Khanna and Xinxu Li. Invisible languages of the llm universe, 2025. URL https://arxiv.org/abs/2510.11557.
  50. 50.Yubin Kim, Hyewon Jeong, Shan Chen, Shuyue Stella Li, Mingyu Lu, Kumail Alhamoud, Jimin Mun, Cristina Grau, Minseok Jung, Rodrigo Gameiro, Lizhou Fan, Eugene Park, Tristan Lin, Joonsik Yoon, Wonjin Yoon, Maarten Sap, Yulia Tsvetkov, Paul Liang, Xuhai Xu, Xin Liu, Daniel McDuff, Hyeonhoon Lee, Hae Won Park, Samir Tulebaev, and Cynthia Breazeal. Medical hallucination in foundation models and their impact on healthcare. medRxiv, 2025. doi: 10.1101/2025.02.28.25323115. URL https://www.medrxiv.org/content/early/2025/03/03/2025.02.28.25323115.
  51. 51.Robert Kirk, Ishita Mediratta, Christoforos Nalmpantis, Jelena Luketina, Eric Hambro, Edward Grefenstette, and Roberta Raileanu. Understanding the effects of rlhf on llm generalisation and diversity, 2024. URL https://arxiv.org/abs/2310.06452.
  52. 52.Hadas Kotek, Rikker Dockum, and David Sun. Gender bias and stereotypes in large language models. In Proceedings of The ACM Collective Intelligence Conference, CI ’23, pp. 12–24, New York, NY, USA, 2023. Association for Computing Machinery. ISBN 9798400701139. doi: 10.1145/3582269.3615599. URL https://doi.org/10.1145/3582269.3615599.
  53. 53.Preethi Lahoti, Nicholas Blumm, Xiao Ma, Raghavendra Kotikalapudi, Sahitya Potluri, Qijun Tan, Hansa Srinivasan, Ben Packer, Ahmad Beirami, Alex Beutel, and Jilin Chen. Improving diversity of demographic representation in large language models via collective-critiques and self-voting. In Houda Bouamor, Juan Pino, and Kalika Bali (eds.), Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 10383–10405, Singapore, December 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.emnlp-main.643. URL https://aclanthology.org/2023.emnlp-main.643/.
  54. 54.Thom Lake, Eunsol Choi, and Greg Durrett. From distributional to overton pluralism: Investigating large language model alignment, 2025. URL https://arxiv.org/abs/2406.17692.
  55. 55.Messi H. J. Lee and Soyeon Jeon. Token sampling uncertainty does not explain homogeneity bias in large language models, 2025. URL https://arxiv.org/abs/2501.19337.
  56. 56.Zhehui Liao, Maria Antoniak, Inyoung Cheong, Evie Yu-Yen Cheng, Ai-Heng Lee, Kyle Lo, Joseph Chee Chang, and Amy X. Zhang. Llms as research tools: A large scale survey of researchers’ usage and perceptions, 2024. URL https://arxiv.org/abs/2411.05025.
  57. 57.MingShan Liu and Jialing Fang. Enhancing mathematical reasoning in large language models with self-consistency-based hallucination detection, 2025. URL https://arxiv.org/abs/2504.09440.
  58. 58.Songyang Liu, Chaozhuo Li, Jiameng Qiu, Xi Zhang, Feiran Huang, Litian Zhang, Yiming Hei, and Philip S. Yu. The scales of justitia: A comprehensive survey on safety evaluation of llms, 2025a. URL https://arxiv.org/abs/2506.11094.
  59. 59.Xiaoou Liu, Tiejin Chen, Longchao Da, Chacha Chen, Zhen Lin, and Hua Wei. Uncertainty quantification and confidence calibration in large language models: A survey, 2025b. URL https://arxiv.org/abs/2503.15850.
  60. 60.Yang Liu, Yuanshun Yao, Jean-Francois Ton, Xiaoying Zhang, Ruocheng Guo, Hao Cheng, Yegor Klochkov, Muhammad Faaiz Taufiq, and Hang Li. Trustworthy llms: a survey and guideline for evaluating large language models’ alignment, 2024. URL https://arxiv.org/abs/2308.05374.
  61. 61.Fangrui Lv, Kaixiong Gong, Jian Liang, Xinyu Pang, and Changshui Zhang. Subjective topic meets LLMs: Unleashing comprehensive, reflective and creative thinking through the negation of negation. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (eds.), Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp. 12318–12341, Miami, Florida, USA, November 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.emnlp-main.686. URL https://aclanthology.org/2024.emnlp-main.686/.
  62. 62.Jiaju Ma, Lei Shi, Kenneth Aleksander Robertsen, and Peggy Chi. Ambigchat: Interactive hierarchical clarification for ambiguous open-domain question answering. In Proceedings of the 38th Annual ACM Symposium on User Interface Software and Technology, UIST ’25, New York, NY, USA, 2025. Association for Computing Machinery. ISBN 9798400720376. doi: 10.1145/3746059.3747686. URL https://doi.org/10.1145/3746059.3747686.
  63. 63.Bertalan Mesko and Eric J. Topol. The imperative for regulatory oversight of large language models (or generative ai) in healthcare. npj Digital Medicine, 6(1):120, 2023. ISSN 2398-6352. doi: 10.1038/s41746-023-00873-0. URL https://doi.org/10.1038/s41746-023-00873-0.
  64. 64.Sewon Min, Julian Michael, Hannaneh Hajishirzi, and Luke Zettlemoyer. AmbigQA: Answering ambiguous open-domain questions. In Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu (eds.), Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 5783–5797, Online, November 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.emnlp-main.466. URL https://aclanthology.org/2020.emnlp-main.466/.
  65. 65.Joel Mire, Zubin Trivadi Aysola, Daniel Chechelnitsky, Nicholas Deas, Chrysoula Zerva, and Maarten Sap. Rejected dialects: Biases against African American language in reward models. In Luis Chiruzzo, Alan Ritter, and Lu Wang (eds.), Findings of the Association for Computational Linguistics: NAACL 2025, pp. 7468–7487, Albuquerque, New Mexico, April 2025. Association for Computational Linguistics. ISBN 979-8-89176-195-7. doi: 10.18653/v1/2025.findings-naacl.417. URL https://aclanthology.org/2025.findings-naacl.417/.
  66. 66.Kibum Moon, Adam E. Green, and Kostadin Kushlev. Homogenizing effect of large language models (llms) on creative diversity: An empirical comparison of human and chatgpt writing. Computers in Human Behavior: Artificial Humans, 6:100207, 2025. ISSN 2949-8821. doi: https://doi.org/10.1016/j.chbah.2025.100207. URL https://www.sciencedirect.com/science/article/pii/S294988212500091X.
  67. 67.Sonia Krishna Murthy, Tomer Ullman, and Jennifer Hu. One fish, two fish, but not the whole sea: Alignment reduces language models’ conceptual diversity. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pp. 11241–11258. Association for Computational Linguistics, 2025. doi: 10.18653/v1/2025.naacl-long.561. URL http://dx.doi.org/10.18653/v1/2025.naacl-long.561.
  68. 68.Tarek Naous, Michael J Ryan, Alan Ritter, and Wei Xu. Having beer after prayer? measuring cultural bias in large language models. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar (eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 16366–16393, Bangkok, Thailand, August 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.acl-long.862. URL https://aclanthology.org/2024.acl-long.862/.
  69. 69.Vera Neplenbroek, Arianna Bisazza, and Raquel Fernandez. Reading between the prompts: How stereotypes shape LLM’s implicit personalization. In Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, and Violet Peng (eds.), Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pp. 20367–20400, Suzhou, China, November 2025. Association for Computational Linguistics. ISBN 979-8-89176-332-6. doi: 10.18653/v1/2025.emnlp-main.1029. URL https://aclanthology.org/2025.emnlp-main.1029/.
  70. 70.Xuan-Phi Nguyen, Mahani Aljunied, Shafiq Joty, and Lidong Bing. Democratizing LLMs for low-resource languages by leveraging their English dominant abilities with linguistically-diverse prompts. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar (eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 3501–3516, Bangkok, Thailand, August 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.acl-long.192. URL https://aclanthology.org/2024.acl-long.192/.
  71. 71.Owen O’Neill, Rajitha Ramanayake, Abhishek Mandal, Urja Pawar, Will Flanagan, Houssem Chatbri, and Christopher Martin. A practical taxonomy for finance-specific LLM risk detection and monitoring. In NeurIPS 2025 Workshop: Generative AI in Finance, 2026. URL https://openreview.net/forum?id=n0tbeSkK9i.
  72. 72.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. Training language models to follow instructions with human feedback, 2022. URL https://arxiv.org/abs/2203.02155.
  73. 73.Vishakh Padmakumar and He He. Does writing with language models reduce content diversity?, 2024. URL https://arxiv.org/abs/2309.05196.
  74. 74.Eileen Pan, Anna Seo Gyeong Choi, Maartje Ter Hoeve, Skyler Seto, and Allison Koenecke. Analyzing dialectical biases in LLMs for knowledge and reasoning benchmarks. In Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, and Violet Peng (eds.), Findings of the Association for Computational Linguistics: EMNLP 2025, pp. 20882–20893, Suzhou, China, November 2025a. Association for Computational Linguistics. ISBN 979-8-89176-335-7. doi: 10.18653/v1/2025.findings-emnlp.1139. URL https://aclanthology.org/2025.findings-emnlp.1139/.
  75. 75.Wenbo Pan, Jie Xu, Qiguang Chen, Junhao Dong, Libo Qin, Xinfeng Li, Haining Yu, and Xiaohua Jia. Can llms refuse questions they do not know? measuring knowledge-aware refusal in factual tasks, 2025b. URL https://arxiv.org/abs/2510.01782.
  76. 76.Giulio Pelosio, Devesh Batra, Noemie Bovey, Robert Hankache, Cristovao Iglesias, Greig Cowan, and Raad Khraishi. Obscured but not erased: Evaluating nationality bias in llms via name-based bias benchmarks, 2025. URL https://arxiv.org/abs/2507.16989.
  77. 77.Sriyash Poddar, Yanming Wan, Hamish Ivison, Abhishek Gupta, and Natasha Jaques. Personalizing reinforcement learning from human feedback with variational preference learning. In Proceedings of the 38th International Conference on Neural Information Processing Systems, NIPS ’24, Red Hook, NY, USA, 2024. Curran Associates Inc. ISBN 9798331314385.
  78. 78.Sonal Prabhune, Balaji Padmanabhan, and Kaushik Dutta. Information-consistent language model recommendations through group relative policy optimization, 2025. URL https://arxiv.org/abs/2512.12858.
  79. 79.Rida Qadri, Aida M. Davani, Kevin Robinson, and Vinodkumar Prabhakaran. Risks of cultural erasure in large language models, 2025. URL https://arxiv.org/abs/2501.01056.
  80. 80.Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, and Chelsea Finn. Direct preference optimization: Your language model is secretly a reward model, 2024. URL https://arxiv.org/abs/2305.18290.
  81. 81.Ian Rios-Sialer. Structure-aware diversity pursuit as an ai safety strategy against homogenization, 2026. URL https://arxiv.org/abs/2601.06116.
  82. 82.Corby Rosset, Ho-Lam Chung, Guanghui Qin, Ethan C. Chau, Zhuo Feng, Ahmed Awadallah, Jennifer Neville, and Nikhil Rao. Researchy questions: A dataset of multi-perspective, decompositional questions for llm web agents, 2024. URL https://arxiv.org/abs/2402.17896.
  83. 83.Ivan Sekulic, Mohammad Aliannejadi, and Fabio Crestani. Towards facet-driven generation of clarifying questions for conversational search. In Proceedings of the 2021 ACM SIGIR International Conference on Theory of Information Retrieval, ICTIR ’21, pp. 167–175, New York, NY, USA, 2021. Association for Computing Machinery. ISBN 9781450386111. doi: 10.1145/3471158.3472257. URL https://doi.org/10.1145/3471158.3472257.
  84. 84.Orit Shaer, Angelora Cooper, Osnat Mokryn, Andrew L Kun, and Hagit Ben Shoshan. Ai-augmented brainwriting: Investigating the use of llms in group ideation. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, CHI ’24, New York, NY, USA, 2024. Association for Computing Machinery. ISBN 9798400703300. doi: 10.1145/3613904.3642414. URL https://doi.org/10.1145/3613904.3642414.
  85. 85.Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R. Bowman, Newton Cheng, Esin Durmus, Zac Hatfield-Dodds, Scott R. Johnston, Shauna Kravec, Timothy Maxwell, Sam McCandlish, Kamal Ndousse, Oliver Rausch, Nicholas Schiefer, Da Yan, Miranda Zhang, and Ethan Perez. Towards understanding sycophancy in language models, 2025. URL https://arxiv.org/abs/2310.13548.
  86. 86.Siqi Shen, Lajanugen Logeswaran, Moontae Lee, Honglak Lee, Soujanya Poria, and Rada Mihalcea. Understanding the capabilities and limitations of large language models for cultural commonsense. In Kevin Duh, Helena Gomez, and Steven Bethard (eds.), Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pp. 5668–5680, Mexico City, Mexico, June 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.naacl-long.316. URL https://aclanthology.org/2024.naacl-long.316/.
  87. 87.Emily Sheng, Kai-Wei Chang, Prem Natarajan, and Nanyun Peng. Societal biases in language generation: Progress and challenges. In Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli (eds.), Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pp. 4275–4293, Online, August 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.acl-long.330. URL https://aclanthology.org/2021.acl-long.330/.
  88. 88.Zhengyan Shi, Giuseppe Castellucci, Simone Filice, Saar Kuzi, Elad Kravi, Eugene Agichtein, Oleg Rokhlenko, and Shervin Malmasi. Ambiguity detection and uncertainty calibration for question answering with large language models. In Trista Cao, Anubrata Das, Tharindu Kumarage, Yixin Wan, Satyapriya Krishna, Ninareh Mehrabi, Jwala Dhamala, Anil Ramakrishna, Aram Galystan, Anoop Kumar, Rahul Gupta, and Kai-Wei Chang (eds.), Proceedings of the 5th Workshop on Trustworthy NLP (TrustNLP 2025), pp. 41–55, Albuquerque, New Mexico, May 2025. Association for Computational Linguistics. ISBN 979-8-89176-233-6. doi: 10.18653/v1/2025.trustnlp-main.4. URL https://aclanthology.org/2025.trustnlp-main.4/.
  89. 89.Alexander Shypula, Shuo Li, Botong Zhang, Vishakh Padmakumar, Kayo Yin, and Osbert Bastani. Evaluating the diversity and quality of llm generated content, 2025. URL https://arxiv.org/abs/2504.12522.
  90. 90.Peiyang Song, Pengrui Han, and Noah Goodman. Large language model reasoning failures, 2026. URL https://arxiv.org/abs/2602.06176.
  91. 91.Taylor Sorensen, Jared Moore, Jillian Fisher, Mitchell Gordon, Niloofar Mireshghallah, Christopher Michael Rytting, Andre Ye, Liwei Jiang, Ximing Lu, Nouha Dziri, Tim Althoff, and Yejin Choi. A roadmap to pluralistic alignment, 2024. URL https://arxiv.org/abs/2402.05070.
  92. 92.Taylor Sorensen, Benjamin Newman, Jared Moore, Chan Park, Jillian Fisher, Niloofar Mireshghallah, Liwei Jiang, and Yejin Choi. Spectrum tuning: Post-training for distributional coverage and in-context steerability, 2025. URL https://arxiv.org/abs/2510.06084.
  93. 93.Peiqi Sui. Llms exhibit significantly lower uncertainty in creative writing than professional writers, 2026. URL https://arxiv.org/abs/2602.16162.
  94. 94.Yan Tao, Olga Viberg, Ryan S Baker, and Rene F Kizilcec. Cultural bias and cultural alignment of large language models. PNAS Nexus, 3(9), September 2024. ISSN 2752-6542. doi: 10.1093/pnasnexus/pgae346. URL http://dx.doi.org/10.1093/pnasnexus/pgae346.
  95. 95.Edward Tian. Identifying GPT: First principles for generative AI detection, 2023. URL http://arks.princeton.edu/ark:/88435/dsp0100000330z.
  96. 96.Jonathan Uesato, Nate Kushman, Ramana Kumar, Francis Song, Noah Siegel, Lisa Wang, Antonia Creswell, Geoffrey Irving, and Irina Higgins. Solving math word problems with process- and outcome-based feedback, 2022. URL https://arxiv.org/abs/2211.14275.
  97. 97.Yanming Wan, Jiaxing Wu, Marwa Abdulhai, Lior Shani, and Natasha Jaques. Enhancing personalized multi-turn dialogue with curiosity reward, 2025. URL https://arxiv.org/abs/2504.03206.
  98. 98.Cunxiang Wang, Xiaoze Liu, Yuanhao Yue, Xiangru Tang, Tianhang Zhang, Cheng Jiayang, Yunzhi Yao, Wenyang Gao, Xuming Hu, Zehan Qi, Yidong Wang, Linyi Yang, Jindong Wang, Xing Xie, Zheng Zhang, and Yue Zhang. Survey on factuality in large language models: Knowledge, retrieval and domain-specificity, 2023. URL https://arxiv.org/abs/2310.07521.
  99. 99.Danqing Wang, Kevin Yang, Hanlin Zhu, Xiaomeng Yang, Andrew Cohen, Lei Li, and Yuandong Tian. Learning personalized alignment for evaluating open-ended text generation. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (eds.), Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp. 13274–13292, Miami, Florida, USA, November 2024a. Association for Computational Linguistics. doi: 10.18653/v1/2024.emnlp-main.737. URL https://aclanthology.org/2024.emnlp-main.737/.
  100. 100.Jiashuo Wang, Kaitao Song, Chunpu Xu, Changhe Song, Yang Xiao, Dongsheng Li, Lili Qiu, and Wenjie Li. Enhancing user engagement in socially-driven dialogue through interactive llm alignments, 2025. URL https://arxiv.org/abs/2506.21497.
  101. 101.Yuxia Wang, Haonan Li, Xudong Han, Preslav Nakov, and Timothy Baldwin. Do-not-answer: Evaluating safeguards in LLMs. In Yvette Graham and Matthew Purver (eds.), Findings of the Association for Computational Linguistics: EACL 2024, pp. 896–911, St. Julian’s, Malta, March 2024b. Association for Computational Linguistics. doi: 10.18653/v1/2024.findings-eacl.61. URL https://aclanthology.org/2024.findings-eacl.61/.
  102. 102.Jakub Wdowicz. Not a mirror, a caricature: How llms reproduce cultural identity? AI and Ethics, 6(1):48, 2025. doi: 10.1007/s43681-025-00898-z. URL https://doi.org/10.1007/s43681-025-00898-z.
  103. 103.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou. Chain-of-thought prompting elicits reasoning in large language models, 2023. URL https://arxiv.org/abs/2201.11903.
  104. 104.Weijia Xu, Nebojsa Jojic, Sudha Rao, Chris Brockett, and Bill Dolan. Echoes in ai: Quantifying lack of plot diversity in llm outputs. Proceedings of the National Academy of Sciences, 122(35):e2504966122, 2025. doi: 10.1073/pnas.2504966122. URL https://www.pnas.org/doi/abs/10.1073/pnas.2504966122.
  105. 105.Yifan Yang, Xiaoyu Liu, Qiao Jin, Furong Huang, and Zhiyong Lu. Unmasking and quantifying racial bias of large language models in medical report generation. Communications Medicine, 4(1):176, 2024. ISSN 2730-664X. doi: 10.1038/s43856-024-00601-z. URL https://doi.org/10.1038/s43856-024-00601-z.
  106. 106.Zhangyue Yin, Qiushi Sun, Qipeng Guo, Jiawen Wu, Xipeng Qiu, and Xuanjing Huang. Do large language models know what they don’t know?, 2023. URL https://arxiv.org/abs/2305.18153.
  107. 107.Yuan Yuan, Tina Sriskandarajah, Anna-Luisa Brakman, Alec Helyar, Alex Beutel, Andrea Vallone, and Saachi Jain. From hard refusals to safe-completions: Toward output-centric safety training, 2025. URL https://arxiv.org/abs/2508.09224.
  108. 108.Sarfaroz Yunusov, Kaige Chen, Kazi Nishat Anwar, and Ali Emami. Personality matters: User traits predict LLM preferences in multi-turn collaborative tasks. In Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, and Violet Peng (eds.), Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pp. 1359–1372, Suzhou, China, November 2025. Association for Computational Linguistics. ISBN 979-8-89176-332-6. doi: 10.18653/v1/2025.emnlp-main.71. URL https://aclanthology.org/2025.emnlp-main.71/.
  109. 109.Jiayi Zhang, Simon Yu, Derek Chong, Anthony Sicilia, Michael R. Tomz, Christopher D. Manning, and Weiyan Shi. Verbalized sampling: How to mitigate mode collapse and unlock llm diversity, 2025. URL https://arxiv.org/abs/2510.01171.
  110. 110.Ke Zhou, Marios Constantinides, and Daniele Quercia. Should llms be weird? exploring weirdness and human rights in large language models, 2025. URL https://arxiv.org/abs/2508.19269.

Citation

MLA
Dhingra, H. “Magic, Madness, Heaven, Sin: LLM Output Diversity Is Everything, Everywhere, All at Once”. arXiv, 2026, http://arxiv.org/abs/2604.01504v1.
APA
Dhingra, H. (2026). Magic, Madness, Heaven, Sin: LLM Output Diversity is Everything, Everywhere, All at Once. arXiv. http://arxiv.org/abs/2604.01504v1
Chicago
Dhingra, H. 2026. “Magic, Madness, Heaven, Sin: LLM Output Diversity Is Everything, Everywhere, All at Once”. arXiv. http://arxiv.org/abs/2604.01504v1.
Harvard
Dhingra, H. (2026) “Magic, Madness, Heaven, Sin: LLM Output Diversity is Everything, Everywhere, All at Once”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2604.01504v1.
Vancouver
1. Dhingra H (2026) Magic, Madness, Heaven, Sin: LLM Output Diversity is Everything, Everywhere, All at Once. arXiv

BibTeX

@article{dhingra2026magic,
  title = {Magic, Madness, Heaven, Sin: LLM Output Diversity is Everything, Everywhere, All at Once},
  author = {Dhingra, Harnoor},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2604.01504v1},
  eprint = {2604.01504}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/