A Systematic Investigation of Commonsense Knowledge in Large Language Models

Xiang Lorraine LiAdhiguna KuncoroJordan HoffmannCyprien de Masson d'AutumePhil BlunsomAida Nematzadeh

article2022EMNLP80 citations

Reveals that scaling model parameters and few-shot prompting fail to close the gap to human-level commonsense reasoning in large language models once superficial cues and evaluation artifacts are strictly controlled.

Listen

Artificial intelligence systems increasingly rely on large pre-trained language models as foundational components for everyday applications. However, these models must understand basic commonsense facts—such as physical constraints and social norms—to operate safely and reliably. While recent systems appear to perform well on standard reasoning tests without specialized training, it remains unclear whether they genuinely understand everyday concepts or merely exploit superficial statistical cues and test formatting quirks.

The article systematically evaluates the commonsense reasoning capabilities of large pre-trained language models in zero-shot and few-shot settings, where models receive no task-specific fine-tuning. The researchers examine whether apparent performance gains reflect genuine reasoning, whether scaling model size or providing brief examples can close the gap with human competence, and how much arbitrary evaluation choices influence published results.

The researchers assessed language model performance across four standard multiple-choice benchmarks covering physical, temporal, and social reasoning: HellaSwag, PIQA, Social IQa, and WinoGrande. They tested models across six parameter scales, ranging from 44 million to 280 billion parameters, using the autoregressive Transformer model Gopher. To separate true reasoning from superficial pattern matching, the analysis introduced an "answer-only" baseline that measured how often models select the correct choice without ever seeing the question. The team further evaluated the effects of providing up to 64 demonstration examples, augmenting prompts with external knowledge bases, and varying technical evaluation settings such as scoring functions and prompt text formats.

The investigation produced several key findings. First, existing benchmarks suffer from substantial answer-only bias; on tests such as HellaSwag and PIQA, the answer-only baseline exceeded random guessing by 32% and 23% respectively, showing that models often pick correct answers using superficial artifacts rather than contextual reasoning. Second, while larger models achieve higher raw accuracy, scaling alone is insufficient; linear projections indicate that reaching human-level performance on three of the four benchmarks would require models ranging from 100 trillion to over 10 to the 18th power parameters, which is computationally infeasible. Third, providing few-shot demonstration examples offered minimal help, improving accuracy by less than 2% on most tasks and failing to close the gap with specialized models. Fourth, minor evaluation design choices—such as the mathematical scoring function and sentence framing—caused performance swings of up to 19.7% on identical models without any change in actual commonsense capability. Finally, retrieving external knowledge base entries provided no meaningful accuracy benefit once base evaluation settings were properly optimized.

These findings demonstrate that current language models lack robust, human-level commonsense reasoning out of the box. High benchmark scores frequently mask reliance on dataset flaws and sensitive prompt formatting rather than true understanding. Consequently, deploying unassisted language models in high-stakes environments carries operational and safety risks. Furthermore, relying purely on increasing model size or few-shot prompting will not resolve these reasoning deficiencies, creating diminishing returns for massive computational investments.

Organizations developing or deploying artificial intelligence should avoid relying on raw scaling alone to achieve commonsense reasoning. Instead, researchers and practitioners should invest in alternative architectures, such as explicit task supervision, multimodal grounding, and physical embodiment. For model evaluation, practitioners must rigorously benchmark systems against answer-only controls, use scoring metrics that account for answer priors such as pointwise mutual information, and report performance variances across multiple prompt and score configurations to ensure robustness.

These conclusions are bounded by certain experimental limits. The evaluations focused entirely on multiple-choice formats rather than open-ended text generation, examined models trained strictly on text rather than multimodal inputs, and relied on the Gopher model family. Nonetheless, because large language models share core architectures and training paradigms across the industry, confidence in the central findings remains high: current text-only models cannot reliably master common sense through unsupervised scale alone.

Cover for A Systematic Investigation of Commonsense Knowledge in Large Language Models

Abstract

Language models (LMs) trained on large amounts of data (e.g., Brown et al., 2020; Patwary et al., 2021) have shown impressive performance on many NLP tasks under the zero-shot and few-shot setup. Here we aim to better understand the extent to which such models learn commonsense knowledge — a critical component of many NLP applications. We conduct a systematic and rigorous zero-shot and few-shot commonsense evaluation of large pre-trained LMs, where we: (i) carefully control for the LMs’ ability to exploit potential surface cues and annotation artefacts, and (ii) account for variations in performance that arise from factors that are not related to commonsense knowledge. Our findings highlight the limitations of pre-trained LMs in acquiring commonsense knowledge without task-specific supervision; furthermore, using larger models or few-shot evaluation are insufficient to achieve human-level commonsense performance.

Table of Contents

  • Abstract
  • 1 Introduction
  • 2 Experimental Setting
  • 2.1 Commonsense Benchmarks
  • 2.2 Pre-trained Language Model
  • 2.3 Baselines
  • 3 Zero-shot Performance
  • 3.1 Answer-only bias
  • 3.2 Does Increasing Model Size Help?
  • 4 Few-shot Performance
  • 5 Robustness of Reported Results
  • 5.1 Do These Design Choices Matter?
  • 6 Related Work
  • 7 Conclusion
  • Ethical Considerations
  • Limitations
  • Acknowledgments
  • References
  • A Appendix Structure
  • B Scaling Behavior
  • C Cross-entropy vs answer length for all datasets
  • D Commonsense Knowledge Bases
  • E Examples

Knowls

  1. Knowl 1 — Context-conditioned and answer-only evaluation

    model/method

    For a multiple-choice question, the evaluation selects the candidate answer with the highest language-model score. With prompt xx, candidate answer yy, and model parameters θ\theta, the default score is the mean conditional token log probability, sθ(y∣x)=1L∑i=1Llog⁡pθ(yi∣x,y<i)s_\theta(y\mid x)=\frac{1}{L}\sum_{i=1}^{L}\log p_\theta(y_i\mid x,y_{<i}), where yiy_i is the iith token of the answer, LL is its token length, and y<iy_{<i} denotes its preceding answer tokens. The prediction is y^=arg⁡max⁡y∈Y(x)sθ(y∣x)\hat{y}=\arg\max_{y\in\mathcal{Y}(x)}s_\theta(y\mid x), where Y(x)\mathcal{Y}(x) is the set of candidate answers for prompt xx.

    The study also scores answers without conditioning on the question, using the same answer-only language-model score sθ(y)s_\theta(y). This Answer-only baseline measures how well a model can select the labeled answer from its prior likelihood alone, potentially exploiting answer-surface cues or annotation artifacts rather than reasoning over the question. A Random baseline chooses uniformly among the candidates, giving expected accuracy 1/K1/K for a question with KK choices. For WinoGrande, the answer-only score is computed on text beginning with the candidate substituted for the pronoun.

  2. Knowl 2 — Zero-shot performance against baselines and supervised state of the art

    empirical result

    The 280-billion-parameter Gopher model was evaluated zero-shot on four multiple-choice commonsense benchmarks. Its accuracy exceeded random choice on every benchmark, but remained below the reported state-of-the-art (SOTA) results, which were achieved by the 11-billion-parameter UNICORN model, pretrained on six commonsense datasets.

    • HellaSwag: Random 25.00%, Answer-only 57.03%, Gopher zero-shot 79.14%, SOTA 93.85%.
    • PIQA: Random 50.00%, Answer-only 73.18%, Gopher zero-shot 80.47%, SOTA 90.13%.
    • Social IQa: Random 33.33%, Answer-only 36.34%, Gopher zero-shot 50.15%, SOTA 83.15%.
    • WinoGrande: Random 50.00%, Answer-only 50.83%, Gopher zero-shot 71.11%, SOTA 91.28%.

    Thus, strong zero-shot accuracy did not establish human-level or SOTA commonsense performance: the zero-shot-to-SOTA gaps exceeded 10 percentage points on all four benchmarks and were especially large on Social IQa and WinoGrande.

  3. Knowl 3 — Answer-only bias varies sharply across benchmarks

    empirical result

    The Answer-only baseline substantially exceeded random choice on HellaSwag and PIQA, demonstrating that a model could often select the benchmark’s labeled answer without using its question or context. The Answer-only-minus-Random accuracy differences were 32.03 percentage points for HellaSwag, 23.18 for PIQA, 3.01 for Social IQa, and 0.83 for WinoGrande.

    This bias makes raw zero-shot accuracy a less direct measure of commonsense reasoning on HellaSwag and PIQA: some measured success can come from answer priors or dataset cues. By comparison, the near-random Answer-only results on Social IQa and WinoGrande make their zero-shot accuracy more indicative of using the supplied context, although they do not by themselves prove that the model reasons correctly.

  4. Knowl 4 — Scaling improves context use, but extrapolated human-level sizes are enormous

    empirical result

    Across models from 44 million to 280 billion parameters, zero-shot accuracy increased with model size on all four benchmarks. However, answer-only accuracy also rose with size, particularly on HellaSwag and PIQA. The zero-shot-minus-Answer-only accuracy margins across model sizes 44M, 117M, 417M, 1.4B, 7.1B, and 280B were:

    • HellaSwag: 2.20, 4.34, 8.48, 13.42, 19.02, and 22.11 percentage points.
    • PIQA: 2.39, 3.37, 5.01, 6.04, 7.40, and 7.29 points.
    • Social IQa: 6.50, 7.52, 9.62, 11.05, 11.16, and 13.82 points.
    • WinoGrande: 2.76, 1.19, 2.37, 8.36, 12.15, and 20.28 points.

    The authors fit accuracy as a linear function of log parameter count and extrapolated to benchmark-specific human-performance thresholds: 95% for HellaSwag and PIQA, 84% for Social IQa, and 94% for WinoGrande. The estimates were 1.4 trillion parameters for HellaSwag, 102 trillion for PIQA, more than 2,000 trillion for WinoGrande, and more than 101810^{18} parameters for Social IQa. These are model-based extrapolations from the measured scaling trend, not experimentally demonstrated capabilities.

  5. Knowl 5 — Few-shot demonstrations help mainly on Social IQa

    empirical result

    Gopher was evaluated with 1, 10, or 64 question-and-correct-answer demonstrations sampled from each benchmark’s training split and placed before the evaluated item. Few-shot scores were averaged across 5–10 runs with different sampled demonstrations. The model’s accuracy percentages were:

    • HellaSwag: zero-shot 79.1; 1-shot 77.8; 10-shot 79.2; 64-shot 79.3.
    • PIQA: zero-shot 80.5; 1-shot 79.3; 10-shot 81.4; 64-shot 81.5.
    • Social IQa: zero-shot 50.2; 1-shot 50.2; 10-shot 55.3; 64-shot 57.5.
    • WinoGrande: zero-shot 71.1; 1-shot 69.2; 10-shot 71.4; 64-shot 74.6.

    A single demonstration sometimes reduced accuracy. HellaSwag and PIQA gained less than two points even with 64 demonstrations, whereas Social IQa gained 7.3 points. The authors suggest that demonstrations help Social IQa in part because its original question format is less natural for a language model. Few-shot prompting did not close the gap to SOTA or human performance on any benchmark.

  6. Knowl 6 — Evaluation choices cause large accuracy swings

    empirical result

    In a robustness study using a 7-billion-parameter model, the researchers varied scoring function, prompt format, question-to-sentence conversion, and whether to score the answer conditional on the question or the concatenated question–answer text. They varied options within one category at a time while holding other categories fixed; the reported best and worst settings were independently selected for each benchmark. The resulting best-minus-worst accuracy differences were 19.7 percentage points for HellaSwag (70.5% versus 50.8%), 16.2 for PIQA (78.7% versus 62.5%), 4.6 for Social IQa (48.5% versus 43.9%), and 2.3 for WinoGrande (62.0% versus 59.7%). Because not all combinations were tested, these differences are a lower bound on the variation across all possible combinations of the tested choices.

    Scoring function had the largest effect. The tested alternatives were mean token log probability (cross-entropy scoring), summed sequence log probability, and pointwise mutual information, log⁡[p(y∣x)/p(y)]\log[p(y\mid x)/p(y)], where xx is the prompt and yy is a candidate answer. Mean token log probability produced the highest accuracy on most benchmarks, while summed sequence log probability was slightly better on WinoGrande. Converting questions to natural sentences mattered most for Social IQa. Scoring only the answer conditioned on the question generally worked better than scoring the concatenated question and answer, except on WinoGrande, which has no question. The results show that benchmark accuracy can depend strongly on evaluation choices unrelated to the commonsense content being tested.

  7. Knowl 7 — Benchmark and model coverage

    experimental setup

    The study used the validation splits, because the test splits were not public, of four multiple-choice benchmarks: HellaSwag (10,042 questions; four options; temporal and physical commonsense), WinoGrande (1,267; two options; social and physical commonsense), Social IQa (1,954; three options; social commonsense), and PIQA (1,838; two options; physical commonsense). The tasks cover story-ending selection, pronoun coreference, social-interaction question answering, and choosing a solution to an everyday physical task, respectively.

    The primary model was Gopher, an autoregressive Transformer with 280 billion parameters, pretrained on more than two trillion MassiveText tokens. The evaluation also used Gopher-family models of 44M, 117M, 417M, 1.4B, and 7.1B parameters for scaling analyses. The pretraining authors removed documents with substantial overlap with the evaluation sets. The study evaluated pretrained models without commonsense-specific fine-tuning; few-shot demonstrations were supplied in context rather than used to update model parameters.

  8. Knowl 8 — Retrieved commonsense triplets do not robustly improve Social IQa

    empirical result

    The study tested whether appending pre-extracted commonsense knowledge-base triplets at evaluation time improves Social IQa accuracy. For each answer candidate, the approach scored the candidate together with retrieved triplets conditional on the question and used the best-scoring triplet; the triplets were generated from COMET or drawn from ATOMIC or ConceptNet. The reported zero-shot accuracies were:

    • 44M model: baseline 42.3%; with COMET 42.9%; with ATOMIC 42.3%; with ConceptNet 40.6%.
    • 117M: 43.6%; 44.0%; 43.6%; 42.2%, respectively.
    • 400M: 46.3%; 46.8%; 44.7%; 44.1%.
    • 1.3B: 47.0%; 46.8%; 46.4%; 44.7%.
    • 7B: 48.5%; 48.6%; 47.5%; 46.1%.

    Across the five model sizes and three knowledge resources, augmentation yielded no substantial, consistent improvement and sometimes reduced accuracy. The authors note that this contrasts with earlier reported gains and emphasize that a knowledge-augmented method should be compared with a strong baseline whose evaluation choices have been tuned.

  9. Knowl 9 — Mean token log probability is associated with answer-length bias

    limitation

    The authors found that mean token log probability, the default answer score, tended to favor longer answer choices to varying degrees on PIQA, Social IQa, and WinoGrande. They attribute this tendency to later answer tokens having more preceding context and therefore being easier for the model to predict; adding such tokens can make an answer’s average log probability more favorable. With summed sequence log probability, the tendency can reverse, with shorter sequences often receiving higher scores. The authors report no correlation between answer length and correctness in these evaluations, so they found that this bias did not change the results they reported.

  10. Knowl 10 — Evaluation scope limits the conclusions

    limitation

    The experiments covered multiple-choice evaluation rather than generative commonsense tasks, and used only the Gopher model family and variants. All evaluated models were trained on language alone; the study therefore does not establish how multimodal training or grounding would affect commonsense performance. The findings characterize these benchmarks and models under the tested settings, rather than directly measuring commonsense competence in open-ended generation or across all language-model families.

Coverage note — Individual qualitative examples of questions that all models answered correctly or incorrectly were omitted because they are illustrative cases rather than generalizable findings.

References

  1. 1.Lisa Bauer and Mohit Bansal. 2021. Identify, align, and integrate: Matching knowledge graphs to commonsense reasoning tasks. EACL.
  2. 2.Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the dangers of stochastic parrots: Can language models be too big? . In Proc. of FAccT.
  3. 3.Chandra Bhagavatula, Ronan Le Bras, Chaitanya Malaviya, Keisuke Sakaguchi, Ari Holtzman, Hannah Rashkin, Doug Downey, Wen tau Yih, and Yejin Choi. 2020. Abductive commonsense reasoning. In International Conference on Learning Representations.
  4. 4.Yonatan Bisk, Ari Holtzman, Jesse Thomason, Jacob Andreas, Yoshua Bengio, Joyce Chai, Mirella Lapata, Angeliki Lazaridou, Jonathan May, Aleksandr Nisnevich, Nicolas Pinto, and Joseph Turian. 2020a. Experience grounds language. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP).
  5. 5.Yonatan Bisk, Rowan Zellers, Jianfeng Gao, Yejin Choi, et al. 2020b. Piqa: Reasoning about physical commonsense in natural language. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 7432–7439.
  6. 6.Su Lin Blodgett, Lisa Green, and Brendan O’Connor. 2016. Demographic dialectal variation in social media: A case study of African-American English. In Proc. of EMNLP.
  7. 7.Rishi Bommasani, Drew A. Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S. Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, Erik Brynjolfsson, S. Buch, D. Card, Rodrigo Castellon, Niladri S. Chatterji, Annie Chen, Kathleen Creel, Jared Davis, Dora Demszky, Chris Donahue, Moussa Doumbouya, Esin Durmus, Stefano Ermon, John Etchemendy, Kawin Ethayarajh, Li Fei-Fei, Chelsea Finn, Trevor Gale, Lauren E. Gillespie, Karan Goel, Noah D. Goodman, Shelby Grossman, Neel Guha, Tatsunori Hashimoto, Peter Henderson, John Hewitt, Daniel E. Ho, Jenny Hong, Kyle Hsu, Jing Huang, Thomas F. Icard, Saahil Jain, Dan Jurafsky, Pratyusha Kalluri, Siddharth Karamcheti, Geoff Keeling, Fereshte Khani, O. Khattab, Pang Wei Koh, Mark S. Krass, Ranjay Krishna, Rohith Kuditipudi, Ananya Kumar, Faisal Ladhak, Mina Lee, Tony Lee, Jure Leskovec, Isabelle Levent, Xiang Lisa Li, Xuechen Li, Tengyu Ma, Ali Malik, Christopher D. Manning, Suvir P. Mirchandani, Eric Mitchell, Zanele Munyikwa, Suraj Nair, Avanika Narayan, Deepak Narayanan, Ben Newman, Allen Nie, Juan Carlos Niebles, Hamed Nilforoshan, J. F. Nyarko, Giray Ogut, Laurel Orr, Isabel Papadimitriou, Joon Sung Park, Chris Piech, Eva Portelance, Christopher Potts, Aditi Raghunathan, Robert Reich, Hongyu Ren, Frieda Rong, Yusuf H. Roohani, Camilo Ruiz, Jackson K. Ryan, Christopher R’e, Dorsa Sadigh, Shiori Sagawa, Keshav Santhanam, Andy Shih, Krishna Parasuram Srinivasan, Alex Tamkin, Rohan Taori, Armin W. Thomas, Florian Tramèr, Rose E. Wang, William Wang, Bohan Wu, Jiajun Wu, Yuhuai Wu, Sang Michael Xie, Michihiro Yasunaga, Jiaxuan You, Matei A. Zaharia, Michael Zhang, Tianyi Zhang, Xikun Zhang, Yuhui Zhang, Lucia Zheng, Kaitlyn Zhou, and Percy Liang. 2021. On the opportunities and risks of foundation models. ArXiv, abs/2108.07258.
  8. 8.Michael Boratko, Xiang Lorraine Li, Rajarshi Das, Tim O’Gorman, Dan Le, and Andrew McCallum. 2020. Protoqa: A question answering dataset for prototypical common-sense reasoning. EMNLP 2020.
  9. 9.Antoine Bosselut, Ronan Le Bras, and Yejin Choi. 2021. Dynamic neuro-symbolic knowledge graph construction for zero-shot commonsense question answering. In Proceedings of the 35th AAAI Conference on Artificial Intelligence (AAAI).
  10. 10.Antoine Bosselut, Hannah Rashkin, Maarten Sap, Chaitanya Malaviya, Asli Çelikyilmaz, and Yejin Choi. 2019. Comet: Commonsense transformers for automatic knowledge graph construction. In ACL.
  11. 11.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc.
  12. 12.Yang Trista Cao and Hal Daumé III. 2021. Toward gender-inclusive coreference resolution: An analysis of gender and bias throughout the machine learning lifecycle*. Computational Linguistics.
  13. 13.Elizabeth Clark, Tal August, Sofia Serrano, Nikita Haduong, Suchin Gururangan, and Noah A. Smith. 2021. All that’s ‘human’ is not gold: Evaluating human evaluation of generated text. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 7282–7296, Online. Association for Computational Linguistics.
  14. 14.Joe Davison, Joshua Feldman, and Alexander M Rush. 2019. Commonsense knowledge mining from pretrained models. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 1173–1178.
  15. 15.Yanai Elazar, Hongming Zhang, Yoav Goldberg, and Dan Roth. 2021. Back to square one: Bias detection, training and commonsense disentanglement in the winograd schema. arXiv preprint arXiv:2104.08161.
  16. 16.John H Flavell. 2004. Theory-of-mind development: Retrospect and prospect. Merrill-Palmer Quarterly (1982-), pages 274–290.
  17. 17.Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith. 2020. RealToxicityPrompts: Evaluating neural toxic degeneration in language models. In Findings of EMNLP.
  18. 18.Jonathan Gordon and Benjamin Van Durme. 2013. Reporting bias and knowledge acquisition. In Proceedings of the 2013 workshop on Automated knowledge base construction, pages 25–30.
  19. 19.David Gunning. 2018. Machine common sense concept paper. arXiv preprint arXiv:1810.07528.
  20. 20.Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles. 2015. Activitynet: A large-scale video benchmark for human activity understanding. In 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 961–970.
  21. 21.Peter Henderson, Riashat Islam, Philip Bachman, Joelle Pineau, Doina Precup, and David Meger. 2018. Deep reinforcement learning that matters. In Proc. of AAAI.
  22. 22.Ari Holtzman, Peter West, Vered Schwartz, Yejin Choi, and Luke Zettlemoyer. 2021. Surface form competition: Why the highest probability answer isn’t always right. arXiv preprint arXiv:2104.08315.
  23. 23.Lifu Huang, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2019. Cosmos qa: Machine reading comprehension with contextual commonsense reasoning. EMNLP, abs/1909.00277.
  24. 24.Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020. Scaling laws for neural language models. CoRR, abs/2001.08361.
  25. 25.Philipp Koehn and Rebecca Knowles. 2017. Six challenges for neural machine translation. In NMT@ACL.
  26. 26.Mahnaz Koupaee and William Yang Wang. 2018. Wikihow: A large scale text summarization dataset. arXiv preprint arXiv:1810.09305.
  27. 27.Hector Levesque, Ernest Davis, and Leora Morgenstern. 2012. The winograd schema challenge. In Thirteenth International Conference on the Principles of Knowledge Representation and Reasoning.
  28. 28.Xiang Lisa Li and Percy Liang. 2021. Prefix-tuning: Optimizing continuous prompts for generation. In Proc. of ACL-IJCNLP.
  29. 29.Bill Yuchen Lin, Haitian Sun, Bhuwan Dhingra, Manzil Zaheer, Xiang Ren, and William W Cohen. 2021. Differentiable open-ended commonsense reasoning. NAACL.
  30. 30.Bill Yuchen Lin, Wangchunshu Zhou, Ming Shen, Pei Zhou, Chandra Bhagavatula, Yejin Choi, and Xiang Ren. 2020. CommonGen: A constrained text generation challenge for generative commonsense reasoning. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 1823–1840, Online. Association for Computational Linguistics.
  31. 31.Chia-Wei Liu, Ryan Lowe, Iulian Serban, Mike Noseworthy, Laurent Charlin, and Joelle Pineau. 2016. How NOT to evaluate your dialogue system: An empirical study of unsupervised evaluation metrics for dialogue response generation. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing.
  32. 32.Hugo Liu and Push Singh. 2004. Commonsense reasoning in and over natural language. In International Conference on Knowledge-Based and Intelligent Information and Engineering Systems, pages 293–306. Springer.
  33. 33.Nicholas Lourie, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2021. Unicorn on rainbow: A universal commonsense reasoning model on a new multitask benchmark. In AAAI.
  34. 34.John McCarthy et al. 1960. Programs with common sense. RLE and MIT computation center.
  35. 35.Gábor Melis, Chris Dyer, and Phil Blunsom. 2018. On the state of the art of evaluation in neural language models. In Proc. of ICLR.
  36. 36.Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2022. Rethinking the role of demonstrations: What makes in-context learning work? arXiv preprint.
  37. 37.Jekaterina Novikova, Ondřej Dušek, Amanda Cercas Curry, and Verena Rieser. 2017. Why we need new evaluation metrics for NLG. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing.
  38. 38.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In ACL.
  39. 39.David A. Patterson, Joseph Gonzalez, Quoc V. Le, Chen Liang, Lluis-Miquel Munguia, Daniel Rothchild, David R. So, Maud Texier, and Jeff Dean. 2021. Carbon emissions and large neural network training. CoRR, abs/2104.10350.
  40. 40.Mostofa Patwary, Mohammad Shoeybi, Patrick LeGresley, Shrimai Prabhumoye, Jared Casper, Vijay Korthikanti, Vartika Singh, Julie Bernauer, Michael Houston, Bryan Catanzaro, Shaden Smith, Brandon Norick, Samyam Rajbhandari, Zhun Liu, George Zerveas, Elton Zhang, Reza Yazdani Aminabadi, Xia Song, Yuxiong He, Jeffrey Zhu, Jennifer Cruzan, Umesh Madan, Luis Vargas, and Saurabh Tiwary. 2021. Using deepspeed and megatron to train megatron-turing nlg 530b, the world’s largest and most powerful generative language model.
  41. 41.Fabio Petroni, Tim Rocktäschel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, Alexander H. Miller, and Sebastian Riedel. 2019. Language models as knowledge bases? In EMNLP.
  42. 42.Adam Poliak, Jason Naradowsky, Aparajita Haldar, Rachel Rudinger, and Benjamin Van Durme. 2018. Hypothesis only baselines for natural language inference. In The Seventh Joint Conference on Lexical and Computational Semantics (*SEM).
  43. 43.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9.
  44. 44.Jack W. Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, H. Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, Eliza Rutherford, Tom Hennigan, Jacob Menick, Albin Cassirer, Richard Powell, George van den Driessche, Lisa Anne Hendricks, Maribeth Rauh, Po-Sen Huang, Amelia Glaese, Johannes Welbl, Sumanth Dathathri, Saffron Huang, Jonathan Uesato, John Mellor, Irina Higgins, Antonia Creswell, Nat McAleese, Amy Wu, Erich Elsen, Siddhant M. Jayakumar, Elena Buchatskaya, David Budden, Esme Sutherland, Karen Simonyan, Michela Paganini, Laurent Sifre, Lena Martens, Xiang Lorraine Li, Adhiguna Kuncoro, Aida Nematzadeh, Elena Gribovskaya, Domenic Donato, Angeliki Lazaridou, Arthur Mensch, Jean-Baptiste Lespiau, Maria Tsimpoukelli, Nikolai Grigorev, Doug Fritz, Thibault Sottiaux, Mantas Pajarskas, Toby Pohlen, Zhitao Gong, Daniel Toyama, Cyprien de Masson d’Autume, Yujia Li, Tayfun Terzi, Vladimir Mikulik, Igor Babuschkin, Aidan Clark, Diego de Las Casas, Aurelia Guy, Chris Jones, James Bradbury, Matthew Johnson, Blake A. Hechtman, Laura Weidinger, Iason Gabriel, William S. Isaac, Edward Lockhart, Simon Osindero, Laura Rimell, Chris Dyer, Oriol Vinyals, Kareem Ayoub, Jeff Stanway, Lorrayne Bennett, Demis Hassabis, Koray Kavukcuoglu, and Geoffrey Irving. 2021. Scaling language models: Methods, analysis & insights from training gopher. CoRR, abs/2112.11446.
  45. 45.Rachel Rudinger, Jason Naradowsky, Brian Leonard, and Benjamin Van Durme. 2018. Gender bias in coreference resolution. In Proc. of NAACL-HLT.
  46. 46.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S Sara Mahdavi, Rapha Gontijo Lopes, et al. 2022. Photorealistic text-to-image diffusion models with deep language understanding. arXiv preprint arXiv:2205.11487.
  47. 47.Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2020. Winogrande: An adversarial winograd schema challenge at scale. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 8732–8740.
  48. 48.Maarten Sap, Ronan Le Bras, Emily Allaway, Chandra Bhagavatula, Nicholas Lourie, Hannah Rashkin, Brendan Roof, Noah A Smith, and Yejin Choi. 2019a. Atomic: An atlas of machine commonsense for if-then reasoning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, pages 3027–3035.
  49. 49.Maarten Sap, Hannah Rashkin, Derek Chen, Ronan LeBras, and Yejin Choi. 2019b. Socialiqa: Commonsense reasoning about social interactions. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing.
  50. 50.Vered Shwartz and Yejin Choi. 2020. Do neural language models overcome reporting bias? In COLING.
  51. 51.Vered Shwartz, Peter West, Ronan Le Bras, Chandra Bhagavatula, , and Yejin Choi. 2020. Unsupervised commonsense question answering with self-talk. In EMNLP.
  52. 52.Felix Stahlberg and Bill Byrne. 2019. On NMT search errors and model errors: Cat got your tongue? In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3354–3360, Hong Kong, China. Association for Computational Linguistics.
  53. 53.Emma Strubell, Ananya Ganesh, and Andrew McCallum. 2019. Energy and policy considerations for deep learning in NLP. In Proc. of ACL.
  54. 54.Paul Trichelair, Ali Emami, Adam Trischler, Kaheer Suleman, and Jackie Chi Kit Cheung. 2019. How reasonable are common-sense reasoning tasks: A case-study on the Winograd schema challenge and SWAG. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3382–3387, Hong Kong, China. Association for Computational Linguistics.
  55. 55.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in neural information processing systems, pages 5998–6008.
  56. 56.Nathaniel Weir, Adam Poliak, and Benjamin Van Durme. 2020. Probing neural language models for human tacit assumptions. arXiv: Computation and Language.
  57. 57.Dani Yogatama, Cyprien de Masson d’Autume, Jerome Connor, Tomas Kocisky, Mike Chrzanowski, Lingpeng Kong, Angeliki Lazaridou, Wang Ling, Lei Yu, Chris Dyer, et al. 2019. Learning and evaluating general linguistic intelligence. arXiv preprint arXiv:1901.11373.
  58. 58.Rowan Zellers, Yonatan Bisk, Roy Schwartz, and Yejin Choi. 2018. Swag: A large-scale adversarial dataset for grounded commonsense inference. EMNLP.
  59. 59.Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019a. Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics.
  60. 60.Rowan Zellers, Ari Holtzman, Hannah Rashkin, Yonatan Bisk, Ali Farhadi, Franziska Roesner, and Yejin Choi. 2019b. Defending against Neural Fake News.
  61. 61.Xuhui Zhou, Yue Zhang, Leyang Cui, and Dandan Huang. 2020. Evaluating commonsense in pretrained language models. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 9733–9740.

Citation

MLA
Li, X. L., et al. “A Systematic Investigation of Commonsense Knowledge in Large Language Models”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 11838–55, https://doi.org/10.18653/V1/2022.EMNLP-MAIN.812.
APA
Li, X. L., Kuncoro, A., Hoffmann, J., de Masson d’Autume, C., Blunsom, P., & Nematzadeh, A. (2022). A Systematic Investigation of Commonsense Knowledge in Large Language Models. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 11838–11855. https://doi.org/10.18653/V1/2022.EMNLP-MAIN.812
Chicago
Li, X. L., A. Kuncoro, J. Hoffmann, C. de Masson d’Autume, P. Blunsom, and A. Nematzadeh. 2022. “A Systematic Investigation of Commonsense Knowledge in Large Language Models”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 11838–55. https://doi.org/10.18653/V1/2022.EMNLP-MAIN.812.
Harvard
Li, X.L. et al. (2022) “A Systematic Investigation of Commonsense Knowledge in Large Language Models”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 11838–11855. Available at: https://doi.org/10.18653/V1/2022.EMNLP-MAIN.812.
Vancouver
1. Li XL, Kuncoro A, Hoffmann J, de Masson d’Autume C, Blunsom P, Nematzadeh A (2022) A Systematic Investigation of Commonsense Knowledge in Large Language Models. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 11838–11855

BibTeX

@inproceedings{Li_2022, title={A Systematic Investigation of Commonsense Knowledge in Large Language Models}, url={http://dx.doi.org/10.18653/V1/2022.EMNLP-MAIN.812}, DOI={10.18653/v1/2022.emnlp-main.812}, booktitle={Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing}, publisher={Association for Computational Linguistics}, author={Li, Xiang Lorraine and Kuncoro, Adhiguna and Hoffmann, Jordan and de Masson d’Autume, Cyprien and Blunsom, Phil and Nematzadeh, Aida}, year={2022}, pages={11838–11855} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/