Built independently by an author, for readers. Read the story and support ChapterPal

keyword

negation bias

Negation bias refers to a reporting imbalance in language corpora where negative commonsense facts are rarely stated explicitly compared to affirmative assertions. Because everyday communication naturally assumes obvious non-occurrences and negative truths rather than documenting what things are not or what does not happen, text datasets overwhelmingly favor positive statements over explicit negations. In natural language processing, this imbalance causes models trained on general text to underrepresent negative knowledge, resulting in difficulty generating, recognizing, or reasoning about commonsense negative facts despite their prevalence in the real world.

1 item

Say What You Mean! Large Language Models Speak Too Positively about Negative Commonsense Knowledge

Say What You Mean! Large Language Models Speak Too Positively about Negative Commonsense Knowledge

Jiangjie Chen, Wei Shi, Ziquan Fu, Sijie Cheng, Lei Li, Yanghua Xiao

OrganizationsBrain Technologies, Inc.Fudan-Aishu Cognitive Intelligence Joint Research CenterFudan UniversitySystem Inc.University of California, Santa Barbara

Why you should read this

Reveals a fundamental belief conflict in large language models where they correctly answer yes-or-no questions about negative commonsense facts yet fail to generate text incorporating that same negative knowledge due to pre-training reporting biases.

Large language models (LLMs) have been widely studied for their ability to store and utilize positive knowledge. However, negative knowledge, such as “lions don’t live in the ocean”, is also ubiquitous in the world but rarely mentioned explicitly in the text. What do LLMs know about negative knowledge? This work examines the ability of LLMs to negative commonsense knowledge. We design a constrained keywords-to-sentence generation task (CG) and a Boolean question-answering task (QA) to probe LLMs. Our experiments reveal that LLMs frequently fail to generate valid sentences grounded in negative commonsense knowledge, yet they can correctly answer polar yes-or-no questions. We term this phenomenon the belief conflict of LLMs. Our further analysis shows that statistical shortcuts and negation reporting bias from language modeling pre-training cause this conflict.

Added

2026-09-26