Say What You Mean! Large Language Models Speak Too Positively about Negative Commonsense Knowledge

Jiangjie ChenWei ShiZiquan FuSijie ChengLei LiYanghua Xiao

article2023ACL62 citations

Reveals a fundamental belief conflict in large language models where they correctly answer yes-or-no questions about negative commonsense facts yet fail to generate text incorporating that same negative knowledge due to pre-training reporting biases.

Listen

Large language models have shown strong capabilities in storing and utilizing positive factual knowledge, but human reasoning relies equally on negative commonsense knowledge—understanding what is false, impossible, or non-existent (such as knowing that lions do not live in the ocean). Because negative knowledge is rarely stated explicitly in natural text, these models may struggle to acquire and use it properly. This lack of reliability poses a significant risk as generative artificial intelligence systems are increasingly deployed in decision-making and automated content generation workflows.

The article investigates whether large language models genuinely acquire implicit negative commonsense knowledge and whether their generated text faithfully reflects their internal knowledge. To test this, the authors constructed a balanced dataset of 4,000 relational commonsense triples and evaluated a range of prominent models—including variants of Flan-T5, GPT-3, Codex, InstructGPT, and ChatGPT—across two tuning-free tasks: a Boolean Question Answering task to probe recognition, and a keyword-to-sentence Constrained Generation task to evaluate generative expression.

The investigation revealed a pronounced behavioral mismatch termed the "belief conflict." While models consistently achieved high accuracy (typically 80% to 85%) when answering direct yes-or-no questions about negative facts, their ability to generate truthful, negated sentences from keywords dropped sharply, often falling below 30% to 50% on negative cases without extensive tuning. Further analysis demonstrated that models rely heavily on statistical word co-occurrences in pre-training corpora; concepts that frequently appear together (such as "worm" and "bird") routinely triggered models to assert false positive relationships (such as "worms eat birds"). Models fine-tuned with reinforcement learning from human feedback, such as newer InstructGPT and ChatGPT versions, performed substantially better at generating negative knowledge than older, raw base models.

These findings indicate that high accuracy on standardized question-answering benchmarks masks a serious tendency for models to hallucinate false facts during free-form text generation. This creates operational and reputational risks in high-stakes applications like automated reasoning, compliance, and explanation generation, where an organization requires faithful and logically sound outputs. Relying solely on standard probing methods gives a misleading picture of model reliability.

To mitigate this risk, practitioners should not deploy standard large language models for open generation of negative commonsense claims without safeguards. Teams should incorporate explicit reasoning techniques, such as chain-of-thought deductive prompting and fact comparison, or increase the proportion of negative examples in prompting contexts, both of which significantly improved negative generation accuracy during testing. Decision-makers should also favor models aligned with human feedback for tasks requiring negative reasoning.

Readers should note that the evaluation was confined to relational commonsense triples from ConceptNet and evaluated primarily through automated negation detection that, while agreeing with human judgment 95% of the time, may overlook subtle phrasing or complex edge cases. Because real-world commonsense knowledge contains exceptions and extends into social, temporal, and spatial domains, further research and domain-specific validation remain necessary before deploying these systems in mission-critical environments.

arXiv: 2305.05976
Cover for Say What You Mean! Large Language Models Speak Too Positively about Negative Commonsense Knowledge

Abstract

Large language models (LLMs) have been widely studied for their ability to store and utilize positive knowledge. However, negative knowledge, such as “lions don’t live in the ocean”, is also ubiquitous in the world but rarely mentioned explicitly in the text. What do LLMs know about negative knowledge? This work examines the ability of LLMs to negative commonsense knowledge. We design a constrained keywords-to-sentence generation task (CG) and a Boolean question-answering task (QA) to probe LLMs. Our experiments reveal that LLMs frequently fail to generate valid sentences grounded in negative commonsense knowledge, yet they can correctly answer polar yes-or-no questions. We term this phenomenon the belief conflict of LLMs. Our further analysis shows that statistical shortcuts and negation reporting bias from language modeling pre-training cause this conflict.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Probing Protocol
  • 3.1 The CSK-PN Dataset
  • 3.2 Probing Task Formulation
  • 3.3 Evaluation Metrics
  • 4 Do LLMs have negative commonsense knowledge?
  • 4.1 Probing LLMs with In-Context Learning
  • 4.2 The Belief Conflict
  • 5 Analysis on the Belief Conflict
  • 5.1 Could keywords as task input hinder the manifestation of LLMs' belief?
  • 5.2 Will the keyword co-occurrence within corpus affect LLMs' generation?
  • 5.3 How does the balance of positive and negative examples affect negation bias?
  • 5.4 Do Chain-of-Thought help generate texts with negative commonsense knowledge?
  • 6 Closing Remarks
  • Limitations
  • Ethical Statement
  • Acknowledgement
  • References
  • A Demonstrations for In-Context Learning
  • A.1 Manually-written Examples for In-Context Learning
  • A.2 Example Prompts for the Probing Tasks
  • B Additional Results
  • B.1 Sensitivity to Temperature Tuning
  • B.2 Abnormal Results of GPT-3 (davinci)
  • B.3 Results of Different Relation Types
  • B.4 Do LLMs hold concerns about exceptions for commonsense knowledge?
  • B.5 Case Study
  • ACL 2023 Responsible NLP Checklist

Knowls

  1. Knowl 1 — Belief conflict between knowing and generating negative knowledge

    empirical result

    The paper identifies a belief conflict in large language models (LLMs): a model can assign the correct truth value to a negative commonsense proposition when directly asked a yes-or-no question, yet fail to generate a sentence that expresses the same negative fact. The conflict is therefore between the model’s apparent belief, measured by Boolean question answering, and its action, measured by constrained sentence generation. Across LLMs with different architectures, sizes, and instruction-tuning histories, positive commonsense facts are usually verbalized reliably, whereas negative facts are often verbalized as affirmative or otherwise unfaithful statements.

  2. Knowl 2 — CSK-PN dataset for positive and negative commonsense relations

    definition

    The CSK-PN probing dataset contains 4,000 relational commonsense triples of the form ⟨s,r,o⟩\langle s,r,o\rangle, where ss and oo are everyday concepts and rr is a relation. The proposition represented by a triple is r(s,o)r(s,o). A triple is treated as a negative fact when the proposition is false according to commonsense knowledge, written ¬r(s,o)\neg r(s,o); otherwise it is treated as a positive fact. The dataset is derived from challenging ConceptNet triples and excludes cases with pronouns, negation, or adjectives as subjects or objects. Positive and negative subsets are balanced at 2,000 triples each and cover six positive/negative relation pairs: Desires/NotDesires (73/1,084), IsA/NotIsA (813/108), HasA/NotHasA (139/69), HasProperty/NotHasProperty (405/190), CapableOf/NotCapableOf (533/537), and MadeOf/NotMadeOf (37/12).

  3. Knowl 3 — Paired probing protocol for LLM beliefs and generation

    model/method

    The paper probes each commonsense triple with two tasks. In Boolean Question Answering (QA), a triple is converted into a natural yes-or-no question; for a negative relation, the explicit NOT prefix is removed so that the model must answer the underlying proposition, such as “Do lions live in the ocean?” In Constrained Sentence Generation (CG), the model receives the subject, relation, and object as keywords and must produce a short factual sentence; the NOT prefix is again removed, so a negative triple requires the model to insert an appropriate negation cue, such as “Lions do not live in the ocean.” Keyword paraphrasing is allowed, making CG a soft-constrained generation task.

    For QA, TP is accuracy on positive triples, TN is accuracy on negative triples, and overall accuracy is Acc=(TP+TN)/2\mathrm{Acc}=(\mathrm{TP}+\mathrm{TN})/2 because the two subsets are equally sized. For CG, the same quantities measure whether generated sentences correctly express positive and negative facts. CG outputs are evaluated using keyword matching followed by a RoBERTa negation-cue classifier; the detector was correct in 97% of 200 manually checked cases, and its agreement with human commonsense judgments was 95% on 100 cases. For fewer than 1% of QA outputs that contain neither “yes” nor “no,” the conditional probabilities of those two tokens are compared.

  4. Knowl 4 — In-context experimental design across LLM families

    experimental setup

    The probing procedure uses in-context learning without parameter updates. The authors manually construct 32 demonstrations: 16 positive and 16 negative examples. Unless otherwise specified, a prompt randomly samples kk examples with equal numbers of positive and negative demonstrations, using either 2-shot or 10-shot settings. Models are evaluated with greedy decoding at temperature 0.00.0.

    The evaluated systems include Flan-T5-3B, Flan-T5-11B, GPT-3 davinci (175B), Codex002, InstructGPT-curie001 (6.7B), InstructGPT001, InstructGPT002, InstructGPT003, and ChatGPT. InstructGPT002 is used as the default model for most analyses. The experiments therefore compare encoder-decoder and autoregressive architectures, base and instruction-tuned models, and models with different code-training or human-feedback histories.

  5. Knowl 5 — Main comparison of QA and constrained generation

    data/table

    The main comparison shows that QA performance is generally balanced between positive and negative triples, while CG performance is often highly imbalanced: models generate positive facts accurately but fail to insert negation for negative facts. The consistency score measures agreement between QA and CG behavior. All values are percentages.

    Could not parse LaTeX table

    The strongest overall behavior comes from InstructGPT003 and ChatGPT, but even these models retain a gap between positive and negative CG accuracy. Increasing the number of demonstrations from 0 to 32 has little effect on QA, improves both positive and negative CG performance, and does not eliminate the CG gap between TP and TN.

  6. Knowl 6 — Question formats expose a statistical shortcut

    empirical result

    Changing the input format shows that the CG failure is not primarily caused by the model’s inability to interpret keywords. With InstructGPT002 and 10-shot prompting, replacing natural-language questions with keyword triples in QA reduces performance by only about 4 percentage points, while preserving relatively balanced positive and negative accuracy. Conversely, replacing keyword triples with natural-language questions in CG raises performance to approximately the QA level.

    The question-to-sentence format nevertheless introduces a shortcut: in zero-shot prompting, 84.58% of InstructGPT002 generations and 80.28% of InstructGPT003 generations begin with “yes” or “no,” effectively turning sentence generation into QA followed by a templated answer. After adding 10 demonstrations, this rate falls below 4%. These results support keyword-to-sentence CG as the more revealing test of whether an LLM can verbalize negative commonsense knowledge rather than merely answer a polarity question.

  7. Knowl 7 — Entity co-occurrence predicts failures on negative facts

    empirical result

    The paper tests whether subject-object co-occurrence in pretraining corpora creates a shortcut for CG. For a triple with subject words ss and object words oo, its co-occurrence score is defined as

    CoOccur(s,o)=∑i,jcooccur(wi,wj)lslo,wi∈s,  wj∈o,\mathrm{CoOccur}(s,o)=\frac{\sum_{i,j}\mathrm{cooccur}(w_i,w_j)}{l_s l_o},\qquad w_i\in s,\;w_j\in o,

    where cooccur(wi,wj)\mathrm{cooccur}(w_i,w_j) is the number of corpus co-occurrences of subject word wiw_i and object word wjw_j, and lsl_s and lol_o are the numbers of words in the subject and object after stopwords are discarded. The score is computed from OMCS and a Wikipedia corpus.

    For Codex002, InstructGPT002, and InstructGPT003, negative-triple TN accuracy drops in the most frequent co-occurrence bucket, with scores above 1,000, whereas positive-triple TP accuracy is best in that same bucket. The authors interpret this pattern as evidence that frequent subject-object associations encourage affirmative completions even when the target relation is negative. For example, because “worm” and “bird” commonly occur together, a model may generate “Worms can eat birds” for a triple labeled as negative.

  8. Knowl 8 — More negative demonstrations mitigate reporting bias

    empirical result

    The authors test whether an imbalance between positive and negative demonstrations causes a negation-reporting bias. Let E+E^+ be the set of positive demonstrations, E−E^- the set of negative demonstrations, kk the total number of demonstrations, and η=∣E−∣/k\eta=|E^-|/k the negative-demonstration ratio. With k=10k=10, InstructGPT002’s QA performance remains largely stable as η\eta changes, except when all demonstrations are positive or all are negative. CG TN accuracy increases as η\eta grows, with the clearest improvement when negative demonstrations constitute nearly all examples; TP falls as positive demonstrations are removed.

    A second experiment holds ∣E+∣=5|E^+|=5 fixed and increases ∣E−∣|E^-| from 5 to 15. In this setting, CG TN continues to improve without sacrificing TP. The results suggest that additional negative examples in training or prompting can reduce the conflict caused by negation-reporting bias, although supplying sufficiently many negative examples may not be practical in every deployment.

  9. Knowl 9 — Chain-of-thought reasoning improves negative-fact generation

    model/method

    The paper evaluates two chain-of-thought (CoT) prompting strategies for CG. Deductive prompting adds intermediate commonsense reasoning: negative propositions are supported using a modus-tollens pattern (“If PP then QQ; not QQ; therefore not PP”), while positive propositions use a modus-ponens pattern (“If PP then QQ; PP; therefore QQ”). Fact-comparison prompting adds a related counterpart fact, such as pairing “lions do not live in the ocean” with “lions live on land.” Demonstrations are manually written with these intermediate rationales.

    Both strategies raise TN accuracy, indicating that explicit reasoning helps models derive and verbalize negative knowledge. The improvement can reduce TP because the baseline models tend to overpredict positive statements. Fact comparison generally preserves TP better than deduction, plausibly because rigid deduction is more sensitive to exceptions in commonsense generalizations.

    Could not parse LaTeX table

    The table reports percentages on CG. For both models and both shot counts, each CoT strategy improves TN over the corresponding no-CoT condition, though the TP/TN trade-off differs by model and prompting strategy.

  10. Knowl 10 — Scope limitations of the negative-knowledge probe

    limitation

    The conclusions are restricted to relational commonsense knowledge between everyday concepts, primarily drawn from ConceptNet-style knowledge bases. The study focuses on commonsense statements that are generally accepted and does not systematically model exceptions, even though facts such as “birds can fly” may have counterexamples. The keyword-to-sentence task is deliberately simple and controlled so that negation can be evaluated reliably; the paper does not establish whether the same belief conflict appears in more complex generation, explanation, argumentation, or reasoning tasks.

    Other forms of negative knowledge, including social, temporal, and spatial knowledge, are not evaluated. The underlying triples may also contain noise, offensive content, or culturally variable judgments. Consequently, the reported negation detector and the observed conflict should not be interpreted as universal measures of all negative knowledge or all LLM generation behavior.

Coverage note — The relation-specific breakdown, temperature-sensitivity analysis, anomalous GPT-3 pattern, detailed case study, and exception-adverb ablation were omitted because they are secondary diagnostics rather than load-bearing contributions among the top ten knowls.

References

  1. 1.Shourya Aggarwal, Divyanshu Mandowara, Vishwajeet Agrawal, Dinesh Khandelwal, Parag Singla, and Dinesh Garg. 2021. Explanations for CommonsenseQA: New Dataset and Models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 3050–3065, Online. Association for Computational Linguistics.
  2. 2.Emily Allaway, Jena D Hwang, Chandra Bhagavatula, Kathleen McKeown, Doug Downey, and Yejin Choi. 2022. Penguins don’t fly: Reasoning about generics through instantiations and exceptions. arXiv preprint arXiv:2205.11658.
  3. 3.Hiba Arnaout, Simon Razniewski, Gerhard Weikum, and Jeff Z. Pan. 2021. Negative knowledge for open-world wikidata. In Companion Proceedings of the Web Conference 2021, WWW ’21, page 544–551, New York, NY, USA. Association for Computing Machinery.
  4. 4.Hiba Arnaout, Simon Razniewski, Gerhard Weikum, and Jeff Z. Pan. 2022. Uncommonsense: Informative negative knowledge about everyday concepts. In Proceedings of the 31st ACM International Conference on Information & Knowledge Management, CIKM ’22, page 37–46, New York, NY, USA. Association for Computing Machinery.
  5. 5.Sören Auer, Christian Bizer, Georgi Kobilarov, Jens Lehmann, Richard Cyganiak, and Zachary Ives. 2007. Dbpedia: A nucleus for a web of open data. In The semantic web, pages 722–735. Springer.
  6. 6.Stephen Barker and Mark Jago. 2012. Being positive about negative facts. Philosophy and Phenomenological research, pages 117–138.
  7. 7.Susanne Bobzien. 2020. Ancient Logic. In Edward N. Zalta, editor, The Stanford Encyclopedia of Philosophy, Summer 2020 edition. Metaphysics Research Lab, Stanford University.
  8. 8.Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al. 2021. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258.
  9. 9.Ruben Branco, António Branco, João António Rodrigues, and João Ricardo Silva. 2021. Shortcutted commonsense: Data spuriousness in deep learning of commonsense reasoning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 1504–1521, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  10. 10.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared J Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc.
  11. 11.Oana-Maria Camburu, Tim Rocktäschel, Thomas Lukasiewicz, and Phil Blunsom. 2018. e-snli: Natural language inference with natural language explanations. In Advances in Neural Information Processing Systems, volume 31. Curran Associates, Inc.
  12. 12.Boxi Cao, Hongyu Lin, Xianpei Han, Le Sun, Lingyong Yan, Meng Liao, Tong Xue, and Jin Xu. 2021. Knowledgeable or educated guess? revisiting language models as knowledge bases. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1860–1874, Online. Association for Computational Linguistics.
  13. 13.Jiangjie Chen, Rui Xu, Ziquan Fu, Wei Shi, Zhongqiao Li, Xinbo Zhang, Changzhi Sun, Lei Li, Yanghua Xiao, and Hao Zhou. 2022. E-KAR: A benchmark for rationalizing natural language analogical reasoning. In Findings of the Association for Computational Linguistics: ACL 2022, pages 3941–3955, Dublin, Ireland. Association for Computational Linguistics.
  14. 14.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Alexander Slama, Greg Brockman, et al. 2021. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374.
  15. 15.Sijie Cheng, Zhiyong Wu, Jiangjie Chen, Zhixing Li, Yang Liu, and Lingpeng Kong. 2022. Unsupervised explanation generation via correct instantiations. arXiv preprint arXiv:2211.11160.
  16. 16.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.
  17. 17.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. 2022. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416.
  18. 18.Peter Clark, Oyvind Tafjord, and Kyle Richardson. 2020. Transformers as soft reasoners over language. In Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence, IJCAI-20, pages 3882–3890. International Joint Conferences on Artificial Intelligence Organization. Main track.
  19. 19.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  20. 20.Nam Do and Ellie Pavlick. 2021. Are rotten apples edible? challenging commonsense inference ability with exceptions. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 2061–2073, Online. Association for Computational Linguistics.
  21. 21.Allyson Ettinger. 2020. What BERT is not: Lessons from a new suite of psycholinguistic diagnostics for language models. Transactions of the Association for Computational Linguistics, 8:34–48.
  22. 22.Melvin Fitting. 2006. Intensional logic.
  23. 23.Tianyu Gao, Xingcheng Yao, and Danqi Chen. 2021. SimCSE: Simple contrastive learning of sentence embeddings. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6894–6910, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  24. 24.Reto Gubelmann and Siegfried Handschuh. 2022. Context matters: A pragmatic study of PLMs’ negation understanding. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 4602–4621, Dublin, Ireland. Association for Computational Linguistics.
  25. 25.Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel Bowman, and Noah A. Smith. 2018. Annotation artifacts in natural language inference data. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), pages 107–112, New Orleans, Louisiana. Association for Computational Linguistics.
  26. 26.Laurence R. Horn and Heinrich Wansing. 2022. Negation. In Edward N. Zalta and Uri Nodelman, editors, The Stanford Encyclopedia of Philosophy, Winter 2022 edition. Metaphysics Research Lab, Stanford University.
  27. 27.Md Mosharaf Hossain, Dhivya Chinnappa, and Eduardo Blanco. 2022. An analysis of negation in natural language understanding corpora. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 716–723, Dublin, Ireland. Association for Computational Linguistics.
  28. 28.Arian Hosseini, Siva Reddy, Dzmitry Bahdanau, R Devon Hjelm, Alessandro Sordoni, and Aaron Courville. 2021. Understanding by understanding not: Modeling negation in language models. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1301–1312, Online. Association for Computational Linguistics.
  29. 29.Joel Jang, Seonghyeon Ye, and Minjoon Seo. 2022. Can large language models truly understand prompts? a case study with negated prompts. arXiv preprint arXiv:2209.12711.
  30. 30.Jaehun Jung, Lianhui Qin, Sean Welleck, Faeze Brahman, Chandra Bhagavatula, Ronan Le Bras, and Yejin Choi. 2022. Maieutic prompting: Logically consistent reasoning with recursive explanations. arXiv preprint arXiv:2205.11822.
  31. 31.Nora Kassner and Hinrich Schütze. 2020. Negated and misprimed probes for pretrained language models: Birds can talk, but cannot fly. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7811–7818, Online. Association for Computational Linguistics.
  32. 32.Nora Kassner, Oyvind Tafjord, Hinrich Schütze, and Peter Clark. 2021. BeliefBank: Adding memory to a pre-trained language model for a systematic notion of belief. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 8849–8861, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  33. 33.Yuxuan Lai, Chen Zhang, Yansong Feng, Quzhe Huang, and Dongyan Zhao. 2021. Why machine reading comprehension models learn shortcuts? In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 989–1002, Online. Association for Computational Linguistics.
  34. 34.Bill Yuchen Lin, Wangchunshu Zhou, Ming Shen, Pei Zhou, Chandra Bhagavatula, Yejin Choi, and Xiang Ren. 2020. CommonGen: A constrained text generation challenge for generative commonsense reasoning. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 1823–1840, Online. Association for Computational Linguistics.
  35. 35.Ye Liu, Yao Wan, Lifang He, Hao Peng, and Philip S. Yu. 2021. Kg-bart: Knowledge graph-augmented bart for generative commonsense reasoning. Proceedings of the AAAI Conference on Artificial Intelligence, 35(7):6418–6425.
  36. 36.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692.
  37. 37.Charles MacDonald. 1965. The role of negation in human knowledge. Laval théologique et philosophique, 21(1):80–114.
  38. 38.Sewon Min, Mike Lewis, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2022a. MetaICL: Learning to learn in context. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2791–2809, Seattle, United States. Association for Computational Linguistics.
  39. 39.Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2022b. Rethinking the role of demonstrations: What makes in-context learning work? arXiv preprint arXiv:2202.12837.
  40. 40.Marvin Minsky. 1997. Negative expertise.
  41. 41.George Molnar. 2000. Truthmakers for negative truths. Australasian Journal of philosophy, 78(1):72–86.
  42. 42.OpenAI. 2022. Chatgpt.
  43. 43.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Gray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems.
  44. 44.Fabio Petroni, Patrick Lewis, Aleksandra Piktus, Tim Rocktäschel, Yuxiang Wu, Alexander H. Miller, and Sebastian Riedel. 2020. How context affects language models’ factual predictions. In Automated Knowledge Base Construction.
  45. 45.Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019. Language models as knowledge bases? In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2463–2473, Hong Kong, China. Association for Computational Linguistics.
  46. 46.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. OpenAI Blog, 1(8):9.
  47. 47.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67.
  48. 48.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. SQuAD: 100,000+ questions for machine comprehension of text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 2383–2392, Austin, Texas. Association for Computational Linguistics.
  49. 49.Abhilasha Ravichander, Matt Gardner, and Ana Marasovic. 2022. Condaqa: A contrastive reading comprehension dataset for reasoning about negation. arXiv preprint arXiv:2211.00295.
  50. 50.Ehud Reiter. 2019. Natural language generation challenges for explainable ai. arXiv preprint arXiv:1911.08794.
  51. 51.Kyle Richardson, Ronen Tamari, Oren Sultan, Reut Tsarfaty, Dafna Shahaf, and Ashish Sabharwal. 2022. Breakpoint transformers for modeling and tracking intermediate beliefs. arXiv preprint arXiv:2211.07950.
  52. 52.Ohad Rubin, Jonathan Herzig, and Jonathan Berant. 2022. Learning to retrieve prompts for in-context learning. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2655–2671, Seattle, United States. Association for Computational Linguistics.
  53. 53.Tara Safavi, Jing Zhu, and Danai Koutra. 2021. NegatER: Unsupervised Discovery of Negatives in Commonsense Knowledge Bases. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 5633–5646, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  54. 54.Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilic, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al. 2022. Bloom: A 176b-parameter open-access multilingual language model. arXiv preprint arXiv:2211.05100.
  55. 55.Push Singh, Thomas Lin, Erik T. Mueller, Grace Lim, Travell Perkins, and Wan Li Zhu. 2002. Open mind common sense: Knowledge acquisition from the general public. In On the Move to Meaningful Internet Systems 2002: CoopIS, DOA, and ODBASE, pages 1223–1237, Berlin, Heidelberg. Springer Berlin Heidelberg.
  56. 56.Robyn Speer, Joshua Chin, and Catherine Havasi. 2017. Conceptnet 5.5: An open multilingual graph of general knowledge. Proceedings of the AAAI Conference on Artificial Intelligence, 31(1).
  57. 57.J.L. Speranza and Laurence R. Horn. 2010. A brief history of negation. Journal of Applied Logic, 8(3):277–301.
  58. 58.Theodore R Sumers, Robert D Hawkins, Mark K Ho, and Thomas L Griffiths. 2021. Extending rational models of communication from beliefs to actions. arXiv preprint arXiv:2105.11950.
  59. 59.Oyvind Tafjord, Bhavana Dalvi, and Peter Clark. 2021. ProofWriter: Generating implications, proofs, and abductive statements over natural language. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 3621–3634, Online. Association for Computational Linguistics.
  60. 60.Oyvind Tafjord, Bhavana Dalvi Mishra, and Peter Clark. 2022. Entailer: Answering questions with faithful and truthful chains of reasoning. arXiv preprint arXiv:2210.12217.
  61. 61.Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019. CommonsenseQA: A question answering challenge targeting commonsense knowledge. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4149–4158, Minneapolis, Minnesota. Association for Computational Linguistics.
  62. 62.Bing Tian, Yixin Cao, Yong Zhang, and Chunxiao Xing. 2022. Debiasing nlu models via causal intervention and counterfactual reasoning.
  63. 63.Denny Vrandecic and Markus Krötzsch. 2014. Wikidata: a free collaborative knowledgebase. Communications of the ACM, 57(10):78–85.
  64. 64.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, and Denny Zhou. 2022. Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171.
  65. 65.Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus. 2022a. Emergent abilities of large language models. Transactions on Machine Learning Research. Survey Certification.
  66. 66.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed H. Chi, Quoc V Le, and Denny Zhou. 2022b. Chain of thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems.
  67. 67.Sean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan, Kyunghyun Cho, and Jason Weston. 2020. Neural text generation with unlikelihood training. In International Conference on Learning Representations.
  68. 68.Wenhao Yu, Chenguang Zhu, Zaitang Li, Zhiting Hu, Qingyun Wang, Heng Ji, and Meng Jiang. 2022. A survey of knowledge-enhanced text generation. ACM Computing Surveys (CSUR).

Citation

MLA
Chen, J., et al. “Say What You Mean! Large Language Models Speak Too Positively About Negative Commonsense Knowledge”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 9890–908, https://doi.org/10.18653/v1/2023.acl-long.550.
APA
Chen, J., Shi, W., Fu, Z., Cheng, S., Li, L., & Xiao, Y. (2023). Say What You Mean! Large Language Models Speak Too Positively about Negative Commonsense Knowledge. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 9890–9908. https://doi.org/10.18653/v1/2023.acl-long.550
Chicago
Chen, J., W. Shi, Z. Fu, S. Cheng, L. Li, and Y. Xiao. 2023. “Say What You Mean! Large Language Models Speak Too Positively About Negative Commonsense Knowledge”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 9890–9908. https://doi.org/10.18653/v1/2023.acl-long.550.
Harvard
Chen, J. et al. (2023) “Say What You Mean! Large Language Models Speak Too Positively about Negative Commonsense Knowledge”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 9890–9908. Available at: https://doi.org/10.18653/v1/2023.acl-long.550.
Vancouver
1. Chen J, Shi W, Fu Z, Cheng S, Li L, Xiao Y (2023) Say What You Mean! Large Language Models Speak Too Positively about Negative Commonsense Knowledge. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 9890–9908

BibTeX

@inproceedings{chen-etal-2023-say,
    title = "Say What You Mean! Large Language Models Speak Too Positively about Negative Commonsense Knowledge",
    author = "Chen, Jiangjie  and
      Shi, Wei  and
      Fu, Ziquan  and
      Cheng, Sijie  and
      Li, Lei  and
      Xiao, Yanghua",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.550/",
    doi = "10.18653/v1/2023.acl-long.550",
    pages = "9890--9908"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/