Say What You Mean! Large Language Models Speak Too Positively about Negative Commonsense Knowledge
Jiangjie ChenWei ShiZiquan FuSijie ChengLei LiYanghua Xiao
Reveals a fundamental belief conflict in large language models where they correctly answer yes-or-no questions about negative commonsense facts yet fail to generate text incorporating that same negative knowledge due to pre-training reporting biases.
Large language models have shown strong capabilities in storing and utilizing positive factual knowledge, but human reasoning relies equally on negative commonsense knowledge—understanding what is false, impossible, or non-existent (such as knowing that lions do not live in the ocean). Because negative knowledge is rarely stated explicitly in natural text, these models may struggle to acquire and use it properly. This lack of reliability poses a significant risk as generative artificial intelligence systems are increasingly deployed in decision-making and automated content generation workflows.
The article investigates whether large language models genuinely acquire implicit negative commonsense knowledge and whether their generated text faithfully reflects their internal knowledge. To test this, the authors constructed a balanced dataset of 4,000 relational commonsense triples and evaluated a range of prominent models—including variants of Flan-T5, GPT-3, Codex, InstructGPT, and ChatGPT—across two tuning-free tasks: a Boolean Question Answering task to probe recognition, and a keyword-to-sentence Constrained Generation task to evaluate generative expression.
The investigation revealed a pronounced behavioral mismatch termed the "belief conflict." While models consistently achieved high accuracy (typically 80% to 85%) when answering direct yes-or-no questions about negative facts, their ability to generate truthful, negated sentences from keywords dropped sharply, often falling below 30% to 50% on negative cases without extensive tuning. Further analysis demonstrated that models rely heavily on statistical word co-occurrences in pre-training corpora; concepts that frequently appear together (such as "worm" and "bird") routinely triggered models to assert false positive relationships (such as "worms eat birds"). Models fine-tuned with reinforcement learning from human feedback, such as newer InstructGPT and ChatGPT versions, performed substantially better at generating negative knowledge than older, raw base models.
These findings indicate that high accuracy on standardized question-answering benchmarks masks a serious tendency for models to hallucinate false facts during free-form text generation. This creates operational and reputational risks in high-stakes applications like automated reasoning, compliance, and explanation generation, where an organization requires faithful and logically sound outputs. Relying solely on standard probing methods gives a misleading picture of model reliability.
To mitigate this risk, practitioners should not deploy standard large language models for open generation of negative commonsense claims without safeguards. Teams should incorporate explicit reasoning techniques, such as chain-of-thought deductive prompting and fact comparison, or increase the proportion of negative examples in prompting contexts, both of which significantly improved negative generation accuracy during testing. Decision-makers should also favor models aligned with human feedback for tasks requiring negative reasoning.
Readers should note that the evaluation was confined to relational commonsense triples from ConceptNet and evaluated primarily through automated negation detection that, while agreeing with human judgment 95% of the time, may overlook subtle phrasing or complex edge cases. Because real-world commonsense knowledge contains exceptions and extends into social, temporal, and spatial domains, further research and domain-specific validation remain necessary before deploying these systems in mission-critical environments.
- Paper: Beyond Accuracy: Behavioral Testing of NLP Models with CheckList, Marco Túlio Ribeiro et al. (2020). Introduces systematic behavioral testing for NLP capabilities, establishing the foundational methodology and demonstrating widespread failure modes in basic negation handling.
- Paper: TruthfulQA: Measuring How Models Mimic Human Falsehoods, Stephanie C. Lin et al. (2022). Establishes how autoregressive language models reliably mimic statistical falsehoods and misconceptions from pre-training corpora, providing essential context for why models struggle with negative commonsense facts.
- Paper: Language Models (Mostly) Know What They Know, Saurav Kadavath et al. (2022). Demonstrates the divergence between internal model probability calibration on recognition tasks and free-form generation, directly underpinning the source's exploration of 'belief conflict'.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). Provides the foundational chain-of-thought prompting methodology that the source adapts as an explicit reasoning mitigation strategy against hallucinated positive assertions.
- Paper: CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge, Alon Talmor et al. (2019). Presents the ConceptNet relational commonsense framework and baseline evaluation setups that inform the dataset construction in the source paper.
- Paper: Right for the Wrong Reasons: Diagnosing Syntactic Heuristics in Natural Language Inference, R. Thomas McCoy et al. (2019). Demonstrates that neural language models frequently rely on shallow statistical co-occurrence and syntactic heuristics rather than genuine semantic reasoning.
- Paper: Consistency Analysis of ChatGPT, Myeongjun Jang et al. (2023). Extends the study of negation and logical inconsistencies in newer instruction-tuned models like ChatGPT across formal reasoning dimensions.
- Paper: Fine-Tuning Language Models for Factuality, Katherine Tian et al. (2024). Applies targeted preference-tuning and reinforcement learning techniques to systematically reduce hallucinations and improve factual reliability during open-ended text generation.
- Paper: Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting, Miles Turpin et al. (2023). Investigates how step-by-step rationales and chain-of-thought explanations can remain unfaithful to internal model drivers, expanding upon the source's analysis of generative versus recognition mismatches.
- Paper: A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions, Lei Huang et al. (2023). Synthesizes broad taxonomic principles, detection methods, and mitigation strategies for factuality hallucinations in large language models.
- Paper: LM vs LM: Detecting Factual Errors via Cross Examination, Roi Cohen et al. (2023). Develops zero-shot conversational cross-examination protocols to detect internal factual inconsistencies and hallucinations exposed by generative tasks.
- Paper: Position: TrustLLM: Trustworthiness in Large Language Models, Yue Huang et al. (2024). Provides a comprehensive multi-dimensional benchmark evaluating truthfulness, safety, and trustworthiness across commercial and open-weight models.
