A Systematic Investigation of Commonsense Knowledge in Large Language Models
Xiang Lorraine LiAdhiguna KuncoroJordan HoffmannCyprien de Masson d'AutumePhil BlunsomAida Nematzadeh
Reveals that scaling model parameters and few-shot prompting fail to close the gap to human-level commonsense reasoning in large language models once superficial cues and evaluation artifacts are strictly controlled.
Artificial intelligence systems increasingly rely on large pre-trained language models as foundational components for everyday applications. However, these models must understand basic commonsense facts—such as physical constraints and social norms—to operate safely and reliably. While recent systems appear to perform well on standard reasoning tests without specialized training, it remains unclear whether they genuinely understand everyday concepts or merely exploit superficial statistical cues and test formatting quirks.
The article systematically evaluates the commonsense reasoning capabilities of large pre-trained language models in zero-shot and few-shot settings, where models receive no task-specific fine-tuning. The researchers examine whether apparent performance gains reflect genuine reasoning, whether scaling model size or providing brief examples can close the gap with human competence, and how much arbitrary evaluation choices influence published results.
The researchers assessed language model performance across four standard multiple-choice benchmarks covering physical, temporal, and social reasoning: HellaSwag, PIQA, Social IQa, and WinoGrande. They tested models across six parameter scales, ranging from 44 million to 280 billion parameters, using the autoregressive Transformer model Gopher. To separate true reasoning from superficial pattern matching, the analysis introduced an "answer-only" baseline that measured how often models select the correct choice without ever seeing the question. The team further evaluated the effects of providing up to 64 demonstration examples, augmenting prompts with external knowledge bases, and varying technical evaluation settings such as scoring functions and prompt text formats.
The investigation produced several key findings. First, existing benchmarks suffer from substantial answer-only bias; on tests such as HellaSwag and PIQA, the answer-only baseline exceeded random guessing by 32% and 23% respectively, showing that models often pick correct answers using superficial artifacts rather than contextual reasoning. Second, while larger models achieve higher raw accuracy, scaling alone is insufficient; linear projections indicate that reaching human-level performance on three of the four benchmarks would require models ranging from 100 trillion to over 10 to the 18th power parameters, which is computationally infeasible. Third, providing few-shot demonstration examples offered minimal help, improving accuracy by less than 2% on most tasks and failing to close the gap with specialized models. Fourth, minor evaluation design choices—such as the mathematical scoring function and sentence framing—caused performance swings of up to 19.7% on identical models without any change in actual commonsense capability. Finally, retrieving external knowledge base entries provided no meaningful accuracy benefit once base evaluation settings were properly optimized.
These findings demonstrate that current language models lack robust, human-level commonsense reasoning out of the box. High benchmark scores frequently mask reliance on dataset flaws and sensitive prompt formatting rather than true understanding. Consequently, deploying unassisted language models in high-stakes environments carries operational and safety risks. Furthermore, relying purely on increasing model size or few-shot prompting will not resolve these reasoning deficiencies, creating diminishing returns for massive computational investments.
Organizations developing or deploying artificial intelligence should avoid relying on raw scaling alone to achieve commonsense reasoning. Instead, researchers and practitioners should invest in alternative architectures, such as explicit task supervision, multimodal grounding, and physical embodiment. For model evaluation, practitioners must rigorously benchmark systems against answer-only controls, use scoring metrics that account for answer priors such as pointwise mutual information, and report performance variances across multiple prompt and score configurations to ensure robustness.
These conclusions are bounded by certain experimental limits. The evaluations focused entirely on multiple-choice formats rather than open-ended text generation, examined models trained strictly on text rather than multimodal inputs, and relied on the Gopher model family. Nonetheless, because large language models share core architectures and training paradigms across the industry, confidence in the central findings remains high: current text-only models cannot reliably master common sense through unsupervised scale alone.
- Paper: Scaling Language Models: Methods, Analysis & Insights from Training Gopher, Jack W. Rae et al. (2021). The source evaluates Gopher across model scales, so the Gopher training and evaluation study provides the model-family and scaling context its experiments build on.
- Paper: Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?, Nishant Balepur et al. (2024). This study develops the source’s answer-only-bias finding into a choices-only analysis, investigating how models can answer multiple-choice benchmarks when the question is removed.
