Discovering Knowledge Deficiencies of Language Models on Massive Knowledge Base
Linxin SongXuwei DingJieyu ZhangTaiwei ShiRyotaro ShimizuRahul GuptaYang LiuJian KangJie-Yu Zhao
Proposes stochastic error ascent, a scalable framework that treats knowledge failure discovery in closed-weight language models as an optimization problem, uncovering up to forty times more factual errors across massive knowledge bases while drastically reducing query costs.
Large language models frequently fail to retain and generate factual knowledge accurately, leading to hallucinations and misinformation in high-stakes environments such as healthcare, law, and scientific research. Exhaustively evaluating these systems against massive knowledge repositories is computationally and financially prohibitive, especially for proprietary models accessible only via programming interfaces. The article addresses this challenge by introducing Stochastic Error Ascent, an automated framework designed to systematically uncover knowledge deficiencies in language models under strict query and budgetary constraints.
To identify systematic weaknesses efficiently, the framework approaches error discovery as an iterative optimization process rather than relying on static benchmarks or random probing. The evaluation utilized an English Wikipedia database containing 7.1 million documents and 28.8 million paragraphs across 13 major categories. The system iteratively navigates this knowledge base by using text embeddings to retrieve new paragraphs semantically similar to previously observed model failures. It employs a two-stage hierarchical search from document abstracts to specific paragraphs and models failure propagation using a directed graph structure to prune unpromising inquiry paths. Using an automated generator to formulate and rephrase multiple-choice questions, the authors evaluated eight prominent language models across both reasoning-focused and standard architectures.
The findings show that Stochastic Error Ascent is substantially more effective and economical than existing automated discovery baselines. It uncovered 40.7 times more errors than Automated Capability Discovery and achieved a 26.7% higher error detection rate than AutoBencher, while reducing the financial cost per discovered error by 599 times and 9 times, respectively. Human validation of 1,000 generated questions confirmed a 100% accuracy and relevance pass rate. Furthermore, error clustering revealed distinct failure modes across model families: systems such as GPT-4o, DeepSeek-V3, and o1-mini exhibited concentrated errors in arts and culture, while other architectures struggled heavily with empirical domains like health, sciences, and chronological reasoning. When models were augmented with retrieved facts on their identified deficiency areas, GPT-4o resolved only 28.6% of errors, demonstrating strong internal memory-context conflicts where models favor incorrect pre-trained assumptions over provided context.
These results demonstrate that standard retrieval-augmented generation may be insufficient to prevent hallucinations when language models possess entrenched incorrect priors, introducing operational risks for enterprises relying on automated outputs. The strong clustering of failures across model families indicates that common data curation practices have created shared blind spots across industry models. Organizations deploying language models should adopt targeted error-discovery auditing to map domain-specific blind spots before deployment, rather than relying on generic benchmarks, and refine training data to remediate identified chronological and contextual weaknesses.
Decision-makers should consider the framework's current operating boundaries. The methodology is currently validated on textual data from structured encyclopedia sources, and attempts to generalize the discovery pipeline to multimodal domains or train lightweight predictor models to forecast failures have yielded limited accuracy. Nonetheless, the high empirical precision and dramatic cost reductions demonstrate strong reliability for auditing text-based language model factual integrity.
- Paper: When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories, Alex Troy Mallen et al. (2022). Its PopQA results establish how factual recall degrades for less popular entities, motivating the source’s search for model knowledge deficiencies across a large knowledge base.
- Paper: HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models, Junyi Li et al. (2023). HaluEval provides an earlier large-scale framework for generating and assessing hallucinations, giving useful precedent for the source’s automated error discovery and validation.
- Paper: FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation, Sewon Min et al. (2023). FActScore establishes automated, retrieval-grounded factuality evaluation, helping explain the evaluation foundations for measuring the source’s discovered errors.
- Paper: Steering Knowledge Selection Behaviours in LLMs via SAE-Based Representation Engineering, Yu Zhao et al. (2025). SPARE turns the source’s finding of persistent memory-context conflicts into a targeted method for steering whether models rely on retrieved context or internal knowledge.
