Discovering Knowledge Deficiencies of Language Models on Massive Knowledge Base

Linxin SongXuwei DingJieyu ZhangTaiwei ShiRyotaro ShimizuRahul GuptaYang LiuJian KangJie-Yu Zhao

article2025arXiv8 citationsBest Paper Award at MRL Workshop 2025

Proposes stochastic error ascent, a scalable framework that treats knowledge failure discovery in closed-weight language models as an optimization problem, uncovering up to forty times more factual errors across massive knowledge bases while drastically reducing query costs.

Listen

Large language models frequently fail to retain and generate factual knowledge accurately, leading to hallucinations and misinformation in high-stakes environments such as healthcare, law, and scientific research. Exhaustively evaluating these systems against massive knowledge repositories is computationally and financially prohibitive, especially for proprietary models accessible only via programming interfaces. The article addresses this challenge by introducing Stochastic Error Ascent, an automated framework designed to systematically uncover knowledge deficiencies in language models under strict query and budgetary constraints.

To identify systematic weaknesses efficiently, the framework approaches error discovery as an iterative optimization process rather than relying on static benchmarks or random probing. The evaluation utilized an English Wikipedia database containing 7.1 million documents and 28.8 million paragraphs across 13 major categories. The system iteratively navigates this knowledge base by using text embeddings to retrieve new paragraphs semantically similar to previously observed model failures. It employs a two-stage hierarchical search from document abstracts to specific paragraphs and models failure propagation using a directed graph structure to prune unpromising inquiry paths. Using an automated generator to formulate and rephrase multiple-choice questions, the authors evaluated eight prominent language models across both reasoning-focused and standard architectures.

The findings show that Stochastic Error Ascent is substantially more effective and economical than existing automated discovery baselines. It uncovered 40.7 times more errors than Automated Capability Discovery and achieved a 26.7% higher error detection rate than AutoBencher, while reducing the financial cost per discovered error by 599 times and 9 times, respectively. Human validation of 1,000 generated questions confirmed a 100% accuracy and relevance pass rate. Furthermore, error clustering revealed distinct failure modes across model families: systems such as GPT-4o, DeepSeek-V3, and o1-mini exhibited concentrated errors in arts and culture, while other architectures struggled heavily with empirical domains like health, sciences, and chronological reasoning. When models were augmented with retrieved facts on their identified deficiency areas, GPT-4o resolved only 28.6% of errors, demonstrating strong internal memory-context conflicts where models favor incorrect pre-trained assumptions over provided context.

These results demonstrate that standard retrieval-augmented generation may be insufficient to prevent hallucinations when language models possess entrenched incorrect priors, introducing operational risks for enterprises relying on automated outputs. The strong clustering of failures across model families indicates that common data curation practices have created shared blind spots across industry models. Organizations deploying language models should adopt targeted error-discovery auditing to map domain-specific blind spots before deployment, rather than relying on generic benchmarks, and refine training data to remediate identified chronological and contextual weaknesses.

Decision-makers should consider the framework's current operating boundaries. The methodology is currently validated on textual data from structured encyclopedia sources, and attempts to generalize the discovery pipeline to multimodal domains or train lightweight predictor models to forecast failures have yielded limited accuracy. Nonetheless, the high empirical precision and dramatic cost reductions demonstrate strong reliability for auditing text-based language model factual integrity.

Cover for Discovering Knowledge Deficiencies of Language Models on Massive Knowledge Base

Abstract

Large language models (LLMs) possess impressive linguistic capabilities but often fail to faithfully retain factual knowledge, leading to hallucinations and unreliable outputs. Understanding LLMs' knowledge deficiencies by exhaustively evaluating against full-scale knowledge bases is computationally prohibitive, especially for closed-weight models. We propose stochastic error ascent (SEA), a scalable and efficient framework for discovering knowledge deficiencies (errors) in closed-weight LLMs under a strict query budget. Rather than naively probing all knowledge candidates, SEA formulates error discovery as a stochastic optimization process: it iteratively retrieves new high-error candidates by leveraging the semantic similarity to previously observed failures. To further enhance search efficiency and coverage, SEA employs hierarchical retrieval across document and paragraph levels, and constructs a relation directed acyclic graph to model error propagation and identify systematic failure modes. Empirically, SEA uncovers 40.7x more knowledge errors than Automated Capability Discovery and 26.7% more than AutoBencher, while reducing the cost-per-error by 599x and 9x, respectively. Human evaluation confirms the high quality of generated questions, while ablation and convergence analyses validate the contribution of each component in SEA. Further analysis on the discovered errors reveals correlated failure patterns across LLM families and recurring deficits, highlighting the need for better data coverage and targeted fine-tuning in future LLM development.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 Vulnerability Discovery for Large Language Models
  • 3.1 Stochastic Error Ascent
  • 4 Comparing Stochastic Error Ascent with Baselines
  • 4.1 Comparing with ACD
  • 4.2 Comparing with AutoBencher
  • 5 Analyzing Stochastic Error Ascent
  • 6 Analyzing LLMs from the Discovery Results
  • 7 Conclusion
  • 8 Future works
  • References
  • A Extra Analysis and Case Studies
  • B Prompt of SEA

Knowls

  1. Knowl 1 — SEA searches for model errors by iteratively following semantic similarity

    algorithm

    Stochastic Error Ascent (SEA) searches a large paragraph collection for facts a black-box language model answers incorrectly, under a fixed inference budget. It maintains a set of evaluated paragraphs, a set of high-error source paragraphs, and a pool of paragraphs not yet selected as sources. First, it evaluates a random initial batch. For each later batch, SEA uses a sentence-embedding model to retrieve paragraphs semantically similar to the current source errors, evaluates the retrieved candidates, and adds them to the evaluated set. It promotes paragraphs whose error rate exceeds the source threshold into the source set, removes promoted paragraphs from the search pool to avoid revisiting them, and updates its relation graph and pruned source set. The loop stops when the inference cost reaches the budget; the output is the evaluated paragraph subset. In the reported implementation, retrieval is hierarchical: retrieve candidate documents by their abstracts, then retrieve paragraphs within those documents. With k=50k=50, SEA takes the top 50 candidates for each source paragraph and randomly selects 40 retrieved paragraphs for the next evaluation batch. The initial batch contains 40 paragraphs sampled across 13 Wikipedia categories. SEA uses an error threshold of 0.50.5 and a source-pruning threshold of 0.50.5; source error rates are calculated over 25 generated questions. The method was run for 20 steps in the reported experiments.

  2. Knowl 2 — Budget-constrained objective for knowledge-deficiency discovery

    equation

    Let KK be a collection of knowledge-base paragraphs, ff a black-box language model being evaluated, and S⊆KS\subseteq K a selected paragraph subset. A question generator gg produces question-answer pairs g(p)g(p) for each paragraph pp; let g(S)=⋃p∈Sg(p)g(S)=\bigcup_{p\in S}g(p). For a generated pair (x,y)(x,y), xx is a multiple-choice question and yy its gold answer. The model’s error rate on SS is the fraction of generated questions answered incorrectly. The discovery objective is to maximize this rate while keeping model-inference cost below budget CC; cost may be measured in API calls or token cost. The paper also rephrases questions into semantically equivalent variants to reduce errors caused by prompt sensitivity.

    TS(f)=1∣g(S)∣∑p∈S∑(x,y)∈g(p)1[f(x)≠y],max⁡S⊆KTS(f)subject tocost⁡f(g(S))<C.T_S(f)=\frac{1}{|g(S)|}\sum_{p\in S}\sum_{(x,y)\in g(p)}\mathbf{1}[f(x)\ne y], \qquad \max_{S\subseteq K} T_S(f)\quad\text{subject to}\quad \operatorname{cost}_f(g(S))<C.

    Here, ∣g(S)∣|g(S)| is the number of generated question-answer pairs, f(x)f(x) is the model’s answer to question xx, and 1[⋅]\mathbf{1}[\cdot] is 1 when its condition is true and 0 otherwise.

  3. Knowl 3 — Relation DAG tracks error propagation and prunes weak sources

    model/method

    SEA represents high-error source paragraphs discovered at successive search steps as nodes in a directed acyclic graph. It adds directed edges from a source paragraph at one step to its most semantically similar newly discovered source paragraphs at the next step, treating these links as potential error-propagation paths. To score a source paragraph pp, SEA averages the individual error rates of all source-paragraph descendants reachable from pp. It prunes pp when this cumulative descendant error is below threshold γ\gamma; the experiments set γ=0.5\gamma=0.5. Newly promoted source paragraphs are removed from the remaining search pool, preventing them from being selected again as new sources. This graph-based pruning is intended to retain sources associated with continuing high-error regions and discard sources whose descendants are less error-inducing.

  4. Knowl 4 — SEA discovers more errors than ACD and AutoBencher at lower cost per error

    empirical result

    Across the reported evaluations, SEA found 40.7 times as many errors as Automated Capability Discovery (ACD) under the same budget, with a 599-fold lower cost per error. Against AutoBencher, compared using an equal number of questions, SEA’s mean error rate was 0.380.38 versus 0.300.30, a 26.7% relative increase, and its cost per error was 9 times lower. On DeepSeek-V3 specifically, SEA’s error rate was 0.420.42, compared with AutoBencher’s 0.260.26. For the ACD comparison, SEA’s discovered source paragraphs were compared with ACD’s failed tasks; for the AutoBencher comparison, the methods were evaluated on the same question count, with AutoBencher producing 2,000 questions across 13 Wikipedia-category benchmarks. ACD failed to run on o1-mini because of an OpenAI prompt-usage-policy violation.

  5. Knowl 5 — Evaluation uses a 7.1-million-document Wikipedia collection and eight models

    experimental setup

    The knowledge base contains 7.1 million English Wikipedia documents and 28.8 million paragraphs across 13 top-level categories and their subcategories; each page is treated as a document, its abstract is used for document-level retrieval, and its sections are treated as paragraphs. The evaluation covers DeepSeek-R1, DeepSeek-R1-Distill-Llama-70B, DeepSeek-V3, Llama-3.3-70B-Instruct, Qwen2.5-72B-Instruct, gpt-4o, gpt-4o-mini, and o1-mini. GPT-4o generates and rephrases multiple-choice questions from paragraphs, and the same question template is used for each test model. SEA uses mGTE embeddings for hierarchical retrieval; the page-title and abstract embeddings are indexed using FAISS. Decoding uses temperature 0.10.1 and top-p 0.90.9. In the main experiments, the initial batch and subsequent retrieved batches each contain 40 paragraphs, the retrieval parameter is k=50k=50, and the source and pruning thresholds are both 0.50.5. The reported 20-step runs evaluate 20,000 questions in total.

  6. Knowl 6 — Human review found all sampled generated questions reliable

    empirical result

    Five college-level student evaluators checked 1,000 questions randomly sampled from SEA’s 20 search steps by comparing their gold answers with the corresponding source paragraphs. All 1,000 questions passed: each question’s answer was present in, and supported by, the paragraph used to generate it. The paper reports this as a 100% human pass rate for the sampled generated questions.

  7. Knowl 7 — Search errors rise over iterations, and retrieval and pruning both contribute

    empirical result

    Across eight evaluated models, SEA’s cumulative error rate rose sharply in early search steps and then tended toward a plateau over 20 iterations, while the per-step batches continued to expose challenging questions. The trajectories varied by model: o1-mini and DeepSeek-V3 showed steeper per-step increases, while gpt-4o-mini was comparatively flat; DeepSeek-R1’s cumulative error increased more gradually. Ablations on Qwen2.5-72B-Instruct, gpt-4o, Llama-3.3-70B-Instruct, and DeepSeek-V3 compared full SEA with a version without source pruning and a random-selection version. Random selection’s cumulative error barely increased, and the gap between full SEA and the unpruned version grew after several steps. The authors interpret these results as evidence that error-related retrieval and pruning each contribute to continued discovery.

  8. Knowl 8 — Models share some failure patterns, but subset difficulty is asymmetric

    empirical result

    Cross-validation evaluated models on error-focused paragraph subsets discovered for other models. The reported correlation between testee and provider results is asymmetric: gpt-4o-mini’s results correlated 0.9170.917 with gpt-4o’s, while the reverse-direction correlation was 0.4230.423. The authors report stronger alignment within model families than across families, with the exception of o1-mini’s behavior. Testee accuracy also varied by provider subset: subsets discovered for gpt-4o and gpt-4o-mini were generally less challenging, whereas DeepSeek-V3’s subset was more difficult, and o1-mini tended to have lower accuracy. DeepSeek-R1 was omitted from this cross-validation because of budget limits.

  9. Knowl 9 — Discovered errors form recurring, model-associated knowledge clusters

    empirical result

    The authors visualized discovered source-error paragraphs by embedding them and reducing the embeddings with t-SNE. In the category-conditioned search, gpt-4o, DeepSeek-V3, and o1-mini errors overlapped substantially, especially in culture and the arts; gpt-4o-mini and DeepSeek-R1 had more distinctive clusters, including history and events and society and social sciences. DeepSeek-V3, DeepSeek-R1-Distill-Llama-70B, and Qwen2.5-72B-Instruct also shared errors in health and fitness and natural and physical sciences. An analysis of two clusters found that gpt-4o, DeepSeek-V3, and o1-mini errors in culture and the arts involved chronological analysis, location details, pattern recognition, data synthesis, and relational reasoning. Errors for Qwen2.5-72B-Instruct, Llama-3.3-70B-Instruct, and DeepSeek-R1-Distill-Llama-70B clustered in health and fitness and natural and physical sciences, with reported difficulties in chronological or historical data, contextual interpretation, trends, assumptions, and contextual associations. These patterns are based on the paper’s selected clusters and should not be read as an exhaustive taxonomy of model failures.

  10. Knowl 10 — Providing retrieved facts only partly resolves detected errors

    empirical result

    To test whether gpt-4o would use explicit evidence for questions it had answered incorrectly, the authors sampled 1,000 incorrect questions from its SEA question-answer set and retested them with their corresponding retrieved factual context included in the prompt. Accuracy reached only 28.6%. The authors interpret this result as evidence that the model often did not adopt the supplied context and continued to rely on parametric knowledge, indicating a memory-context conflict under this test setup.

  11. Knowl 11 — Search scope depends on text quality, initialization, and inference cost

    limitation

    The reported SEA evaluation searches English text from Wikipedia and does not establish performance on images, videos, or other modalities; the authors identify reliable question-answer synthesis from multimodal material as a barrier to extension. They also note that the search results depend on the initial set and that querying closed-weight models constrains search scope through cost. As one possible expansion strategy, they trained BERT classifiers on 4,402 paragraphs collected during gpt-4o search, labeling a paragraph by whether its generated-question accuracy was below 0.50.5 and splitting the data 8:1:1 for training, validation, and testing. The reported test accuracy was 66.22% for bert-base-uncased and 67.85% for bert-large-uncased; the authors characterize this as insufficiently accurate for reliably extending the search.

Coverage note — The five illustrative free-response misinformation cases and the full prompt templates are omitted as examples or implementation materials rather than standalone findings; modality and search-scope limitations are included.

References

  1. 1.David M Blei, Andrew Y Ng, and Michael I Jordan. Latent dirichlet allocation. Journal of machine Learning research, 3(Jan):993–1022, 2003.
  2. 2.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, Dave Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert-Voss, William Hebgen Guss, Alex Nichol, Alex Paino, Nikolas Tezak, Jie Tang, Igor Babuschkin, Suchir Balaji, Shantanu Jain, William Saunders, Christopher Hesse, Andrew N. Carr, Jan Leike, Josh Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, and Wojciech Zaremba. Evaluating large language models trained on code, 2021. URL https://arxiv.org/abs/2107.03374.
  3. 3.DeepSeek-AI, Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, Xiaokang Zhang, Xingkai Yu, Yu Wu, Z. F. Wu, Zhibin Gou, Zhihong Shao, Zhuoshu Li, Ziyi Gao, Aixin Liu, Bing Xue, Bingxuan Wang, Bochao Wu, Bei Feng, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, Damai Dai, Deli Chen, Dongjie Ji, Erhang Li, Fangyun Lin, Fucong Dai, Fuli Luo, Guangbo Hao, Guanting Chen, Guowei Li, H. Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Qu, Hui Li, Jianzhong Guo, Jiashi Li, Jiawei Wang, Jingchang Chen, Jingyang Yuan, Junjie Qiu, Junlong Li, J. L. Cai, Jiaqi Ni, Jian Liang, Jin Chen, Kai Dong, Kai Hu, Kaige Gao, Kang Guan, Kexin Huang, Kuai Yu, Lean Wang, Lecong Zhang, Liang Zhao, Litong Wang, Liyue Zhang, Lei Xu, Leyi Xia, Mingchuan Zhang, Minghua Zhang, Minghui Tang, Meng Li, Miaojun Wang, Mingming Li, Ning Tian, Panpan Huang, Peng Zhang, Qiancheng Wang, Qinyu Chen, Qiushi Du, Ruiqi Ge, Ruisong Zhang, Ruizhe Pan, Runji Wang, R. J. Chen, R. L. Jin, Ruyi Chen, Shanghao Lu, Shangyan Zhou, Shanhuang Chen, Shengfeng Ye, Shiyu Wang, Shuiping Yu, Shunfeng Zhou, Shuting Pan, S. S. Li, Shuang Zhou, Shaoqing Wu, Shengfeng Ye, Tao Yun, Tian Pei, Tianyu Sun, T. Wang, Wangding Zeng, Wanjia Zhao, Wen Liu, Wenfeng Liang, Wenjun Gao, Wenqin Yu, Wentao Zhang, W. L. Xiao, Wei An, Xiaodong Liu, Xiaohan Wang, Xiaokang Chen, Xiaotao Nie, Xin Cheng, Xin Liu, Xin Xie, Xingchao Liu, Xinyu Yang, Xinyuan Li, Xuecheng Su, Xuheng Lin, X. Q. Li, Xiangyue Jin, Xiaojin Shen, Xiaosha Chen, Xiaowen Sun, Xiaoxiang Wang, Xinnan Song, Xinyi Zhou, Xianzu Wang, Xinxia Shan, Y. K. Li, Y. Q. Wang, Y. X. Wei, Yang Zhang, Yanhong Xu, Yao Li, Yao Zhao, Yaofeng Sun, Yaohui Wang, Yi Yu, Yichao Zhang, Yifan Shi, Yiliang Xiong, Ying He, Yishi Piao, Yisong Wang, Yixuan Tan, Yiyang Ma, Yiyuan Liu, Yongqiang Guo, Yuan Ou, Yuduan Wang, Yue Gong, Yuheng Zou, Yujia He, Yunfan Xiong, Yuxiang Luo, Yuxiang You, Yuxuan Liu, Yuyang Zhou, Y. X. Zhu, Yanhong Xu, Yanping Huang, Yaohui Li, Yi Zheng, Yuchen Zhu, Yunxian Ma, Ying Tang, Yukun Zha, Yuting Yan, Z. Z. Ren, Zehui Ren, Zhangli Sha, Zhe Fu, Zhean Xu, Zhenda Xie, Zhengyan Zhang, Zhewen Hao, Zhicheng Ma, Zhigang Yan, Zhiyu Wu, Zihui Gu, Zijia Zhu, Zijun Liu, Zilin Li, Ziwei Xie, Ziyang Song, Zizheng Pan, Zhen Huang, Zhipeng Xu, Zhongyu Zhang, and Zhen Zhang. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025. URL https://arxiv.org/abs/2501.12948.
  4. 4.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding, 2019. URL https://arxiv.org/abs/1810.04805.
  5. 5.Matthijs Douze, Alexandr Guzhva, Chengqi Deng, Jeff Johnson, Gergely Szilvasy, Pierre-Emmanuel Mazaré, Maria Lomeli, Lucas Hosseini, and Hervé Jégou. The faiss library. arXiv preprint arXiv:2401.08281, 2024.
  6. 6.Zorik Gekhman, Jonathan Herzig, Roee Aharoni, Chen Elkind, and Idan Szpektor. Trueteacher: Learning factual consistency evaluation with large language models, 2023. URL https://arxiv.org/abs/2305.11171.
  7. 7.Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Arthur Hinsvark, Arun Rao, Aston Zhang, Aurelien Rodriguez, Austen Gregerson, Ava Spataru, Baptiste Roziere, Bethany Biron, Binh Tang, Bobbie Chern, Charlotte Caucheteux, Chaya Nayak, Chloe Bi, Chris Marra, Chris McConnell, Christian Keller, Christophe Touret, Chunyang Wu, Corinne Wong, Cristian Canton Ferrer, Cyrus Nikolaidis, Damien Allonsius, Daniel Song, Danielle Pintz, Danny Livshits, Danny Wyatt, David Esiobu, Dhruv Choudhary, Dhruv Mahajan, Diego Garcia-Olano, Diego Perino, Dieuwke Hupkes, Egor Lakomkin, Ehab AlBadawy, Elina Lobanova, Emily Dinan, Eric Michael Smith, Filip Radenovic, Francisco Guzmán, Frank Zhang, Gabriel Synnaeve, Gabrielle Lee, Georgia Lewis Anderson, Govind Thattai, Graeme Nail, Gregoire Mialon, Guan Pang, Guillem Cucurell, Hailey Nguyen, Hannah Korevaar, Hu Xu, Hugo Touvron, Iliyan Zarov, Imanol Arrieta Ibarra, Isabel Kloumann, Ishan Misra, Ivan Evtimov, Jack Zhang, Jade Copet, Jaewon Lee, Jan Geffert, Jana Vranes, Jason Park, Jay Mahadeokar, Jeet Shah, Jelmer van der Linde, Jennifer Billock, Jenny Hong, Jenya Lee, Jeremy Fu, Jianfeng Chi, Jianyu Huang, Jiawen Liu, Jie Wang, Jiecao Yu, Joanna Bitton, Joe Spisak, Jongsoo Park, Joseph Rocca, Joshua Johnstun, Joshua Saxe, Junteng Jia, Kalyan Vasuden Alwala, Karthik Prasad, Kartikeya Upasani, Kate Plawiak, Ke Li, Kenneth Heafield, Kevin Stone, Khalid El-Arini, Krithika Iyer, Kshitiz Malik, Kuenley Chiu, Kunal Bhalla, Kushal Lakhotia, Lauren Rantala-Yeary, Laurens van der Maaten, Lawrence Chen, Liang Tan, Liz Jenkins, Louis Martin, Lovish Madaan, Lubo Malo, Lukas Blecher, Lukas Landzaat, Luke de Oliveira, Madeline Muzzi, Mahesh Pasupuleti, Mannat Singh, Manohar Paluri, Marcin Kardas, Maria Tsimpoukelli, Mathew Oldham, Mathieu Rita, Maya Pavlova, Melanie Kambadur, Mike Lewis, Min Si, Mitesh Kumar Singh, Mona Hassan, Naman Goyal, Narjes Torabi, Nikolay Bashlykov, Nikolay Bogoychev, Niladri Chatterji, Ning Zhang, Olivier Duchenne, Onur Çelebi, Patrick Alrassy, Pengchuan Zhang, Pengwei Li, Petar Vasic, Peter Weng, Prajjwal Bhargava, Pratik Dubal, Praveen Krishnan, Punit Singh Koura, Puxin Xu, Qing He, Qingxiao Dong, Ragavan Srinivasan, Raj Ganapathy, Ramon Calderer, Ricardo Silveira Cabral, Robert Stojnic, Roberta Raileanu, Rohan Maheswari, Rohit Girdhar, Rohit Patel, Romain Sauvestre, Ronnie Polidoro, Roshan Sumbaly, Ross Taylor, Ruan Silva, Rui Hou, Rui Wang, Saghar Hosseini, Sahana Chennabasappa, Sanjay Singh, Sean Bell, Seohyun Sonia Kim, Sergey Edunov, Shaoliang Nie, Sharan Narang, Sharath Raparthy, Sheng Shen, Shengye Wan, Shruti Bhosale, Shun Zhang, Simon Vandenhende, Soumya Batra, Spencer Whitman, Sten Sootla, Stephane Collot, Suchin Gururangan, Sydney Borodinsky, Tamar Herman, Tara Fowler, Tarek Sheasha, Thomas Georgiou, Thomas Scialom, Tobias Speckbacher, Todor Mihaylov, Tong Xiao, Ujjwal Karn, Vedanuj Goswami, Vibhor Gupta, Vignesh Ramanathan, Viktor Kerkez, Vincent Gonguet, Virginie Do, Vish Vogeti, Vítor Albiero, Vladan Petrovic, Weiwei Chu, Wenhan Xiong, Wenyin Fu, Whitney Meers, Xavier Martinet, Xiaodong Wang, Xiaofang Wang, Xiaoqing Ellen Tan, Xide Xia, Xinfeng Xie, Xuchao Jia, Xuewei Wang, Yaelle Goldschlag, Yashesh Gaur, Yasmine Babaei, Yi Wen, Yiwen Song, Yuchen Zhang, Yue Li, Yuning Mao, Zacharie Delpierre Coudert, Zheng Yan, Zhengxing Chen, Zoe Papakipos, Aaditya Singh, Aayushi Srivastava, Abha Jain, Adam Kelsey, Adam Shajnfeld, Adithya Gangidi, Adolfo Victoria, Ahuva Goldstand, Ajay Menon, Ajay Sharma, Alex Boesenberg, Alexei Baevski, Allie Feinstein, Amanda Kallet, Amit Sangani, Amos Teo, Anam Yunus, Andrei Lupu, Andres Alvarado, Andrew Caples, Andrew Gu, Andrew Ho, Andrew Poulton, Andrew Ryan, Ankit Ramchandani, Annie Dong, Annie Franco, Anuj Goyal, Aparajita Saraf, Arkabandhu Chowdhury, Ashley Gabriel, Ashwin Bharambe, Assaf Eisenman, Azadeh Yazdan, Beau James, Ben Maurer, Benjamin Leonhardi, Bernie Huang, Beth Loyd, Beto De Paola, Bhargavi Paranjape, Bing Liu, Bo Wu, Boyu Ni, Braden Hancock, Bram Wasti, Brandon Spence, Brani Stojkovic, Brian Gamido, Britt Montalvo, Carl Parker, Carly Burton, Catalina Mejia, Ce Liu, Changhan Wang, Changkyu Kim, Chao Zhou, Chester Hu, Ching-Hsiang Chu, Chris Cai, Chris Tindal, Christoph Feichtenhofer, Cynthia Gao, Damon Civin, Dana Beaty, Daniel Kreymer, Daniel Li, David Adkins, David Xu, Davide Testuggine, Delia David, Devi Parikh, Diana Liskovich, Didem Foss, Dingkang Wang, Duc Le, Dustin Holland, Edward Dowling, Eissa Jamil, Elaine Montgomery, Eleonora Presani, Emily Hahn, Emily Wood, Eric-Tuan Le, Erik Brinkman, Esteban Arcaute, Evan Dunbar, Evan Smothers, Fei Sun, Felix Kreuk, Feng Tian, Filippos Kokkinos, Firat Ozgenel, Francesco Caggioni, Frank Kanayet, Frank Seide, Gabriela Medina Florez, Gabriella Schwarz, Gada Badeer, Georgia Swee, Gil Halpern, Grant Herman, Grigory Sizov, Guangyi, Zhang, Guna Lakshminarayanan, Hakan Inan, Hamid Shojanazeri, Han Zou, Hannah Wang, Hanwen Zha, Haroun Habeeb, Harrison Rudolph, Helen Suk, Henry Aspegren, Hunter Goldman, Hongyuan Zhan, Ibrahim Damlaj, Igor Molybog, Igor Tufanov, Ilias Leontiadis, Irina-Elena Veliche, Itai Gat, Jake Weissman, James Geboski, James Kohli, Janice Lam, Japhet Asher, Jean-Baptiste Gaya, Jeff Marcus, Jeff Tang, Jennifer Chan, Jenny Zhen, Jeremy Reizenstein, Jeremy Teboul, Jessica Zhong, Jian Jin, Jingyi Yang, Joe Cummings, Jon Carvill, Jon Shepard, Jonathan McPhie, Jonathan Torres, Josh Ginsburg, Junjie Wang, Kai Wu, Kam Hou U, Karan Saxena, Kartikay Khandelwal, Katayoun Zand, Kathy Matosich, Kaushik Veeraraghavan, Kelly Michelena, Keqian Li, Kiran Jagadeesh, Kun Huang, Kunal Chawla, Kyle Huang, Lailin Chen, Lakshya Garg, Lavender A, Leandro Silva, Lee Bell, Lei Zhang, Liangpeng Guo, Licheng Yu, Liron Moshkovich, Luca Wehrstedt, Madian Khabsa, Manav Avalani, Manish Bhatt, Martynas Mankus, Matan Hasson, Matthew Lennie, Matthias Reso, Maxim Groshev, Maxim Naumov, Maya Lathi, Meghan Keneally, Miao Liu, Michael L. Seltzer, Michal Valko, Michelle Restrepo, Mihir Patel, Mik Vyatskov, Mikayel Samvelyan, Mike Clark, Mike Macey, Mike Wang, Miquel Jubert Hermoso, Mo Metanat, Mohammad Rastegari, Munish Bansal, Nandhini Santhanam, Natascha Parks, Natasha White, Navyata Bawa, Nayan Singhal, Nick Egebo, Nicolas Usunier, Nikhil Mehta, Nikolay Pavlovich Laptev, Ning Dong, Norman Cheng, Oleg Chernoguz, Olivia Hart, Omkar Salpekar, Ozlem Kalinli, Parkin Kent, Parth Parekh, Paul Saab, Pavan Balaji, Pedro Rittner, Philip Bontrager, Pierre Roux, Piotr Dollar, Polina Zvyagina, Prashant Ratanchandani, Pritish Yuvraj, Qian Liang, Rachad Alao, Rachel Rodriguez, Rafi Ayub, Raghotham Murthy, Raghu Nayani, Rahul Mitra, Rangaprabhu Parthasarathy, Raymond Li, Rebekkah Hogan, Robin Battey, Rocky Wang, Russ Howes, Ruty Rinott, Sachin Mehta, Sachin Siby, Sai Jayesh Bondu, Samyak Datta, Sara Chugh, Sara Hunt, Sargun Dhillon, Sasha Sidorov, Satadru Pan, Saurabh Mahajan, Saurabh Verma, Seiji Yamamoto, Sharadh Ramaswamy, Shaun Lindsay, Shaun Lindsay, Sheng Feng, Shenghao Lin, Shengxin Cindy Zha, Shishir Patil, Shiva Shankar, Shuqiang Zhang, Shuqiang Zhang, Sinong Wang, Sneha Agarwal, Soji Sajuyigbe, Soumith Chintala, Stephanie Max, Stephen Chen, Steve Kehoe, Steve Satterfield, Sudarshan Govindaprasad, Sumit Gupta, Summer Deng, Sungmin Cho, Sunny Virk, Suraj Subramanian, Sy Choudhury, Sydney Goldman, Tal Remez, Tamar Glaser, Tamara Best, Thilo Koehler, Thomas Robinson, Tianhe Li, Tianjun Zhang, Tim Matthews, Timothy Chou, Tzook Shaked, Varun Vontimitta, Victoria Ajayi, Victoria Montanez, Vijai Mohan, Vinay Satish Kumar, Vishal Mangla, Vlad Ionescu, Vlad Poenaru, Vlad Tiberiu Mihailescu, Vladimir Ivanov, Wei Li, Wenchen Wang, Wenwen Jiang, Wes Bouaziz, Will Constable, Xiaocheng Tang, Xiaojian Wu, Xiaolan Wang, Xilun Wu, Xinbo Gao, Yaniv Kleinman, Yanjun Chen, Ye Hu, Ye Jia, Ye Qi, Yenda Li, Yilin Zhang, Ying Zhang, Yossi Adi, Youngjin Nam, Yu, Wang, Yu Zhao, Yuchen Hao, Yundi Qian, Yunlu Li, Yuzi He, Zach Rait, Zachary DeVito, Zef Rosnbrick, Zhaoduo Wen, Zhenyu Yang, Zhiwei Zhao, and Zhiyu Ma. The llama 3 herd of models, 2024. URL https://arxiv.org/abs/2407.21783.
  8. 8.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. In International Conference on Learning Representations, 2021a. URL https://openreview.net/forum?id=d7KBjmI3GmQ.
  9. 9.Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. Measuring mathematical problem solving with the math dataset. NeurIPS, 2021b.
  10. 10.Yue Huang, Chujie Gao, Siyuan Wu, Haoran Wang, Xiangqi Wang, Yujun Zhou, Yanbo Wang, Jiayi Ye, Jiawen Shi, Qihui Zhang, Yuan Li, Han Bao, Zhaoyi Liu, Tianrui Guan, Dongping Chen, Ruoxi Chen, Kehan Guo, Andy Zou, Bryan Hooi Kuen-Yew, Caiming Xiong, Elias Stengel-Eskin, Hongyang Zhang, Hongzhi Yin, Huan Zhang, Huaxiu Yao, Jaehong Yoon, Jieyu Zhang, Kai Shu, Kaijie Zhu, Ranjay Krishna, Swabha Swayamdipta, Taiwei Shi, Weijia Shi, Xiang Li, Yiwei Li, Yuexing Hao, Yuexing Hao, Zhihao Jia, Zhize Li, Xiuying Chen, Zhengzhong Tu, Xiyang Hu, Tianyi Zhou, Jieyu Zhao, Lichao Sun, Furong Huang, Or Cohen Sasson, Prasanna Sattigeri, Anka Reuel, Max Lamparth, Yue Zhao, Nouha Dziri, Yu Su, Huan Sun, Heng Ji, Chaowei Xiao, Mohit Bansal, Nitesh V. Chawla, Jian Pei, Jianfeng Gao, Michael Backes, Philip S. Yu, Neil Zhenqiang Gong, Pin-Yu Chen, Bo Li, and Xiangliang Zhang. On the trustworthiness of generative foundation models: Guideline, assessment, and perspective, 2025. URL https://arxiv.org/abs/2502.14296.
  11. 11.Naman Jain, King Han, Alex Gu, Wen-Ding Li, Fanjia Yan, Tianjun Zhang, Sida Wang, Armando Solar-Lezama, Koushik Sen, and Ion Stoica. Livecodebench: Holistic and contamination free evaluation of large language models for code, 2024. URL https://arxiv.org/abs/2403.07974.
  12. 12.Che Jiang, Biqing Qi, Xiangyu Hong, Dayuan Fu, Yang Cheng, Fandong Meng, Mo Yu, Bowen Zhou, and Jie Zhou. On large language models’ hallucination with regard to known facts, 2024. URL https://arxiv.org/abs/2403.20009.
  13. 13.Zhengbao Jiang, Frank F Xu, Jun Araki, and Graham Neubig. How can we know what language models know? Transactions of the Association for Computational Linguistics, 8:423–438, 2020.
  14. 14.Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, et al. Dynabench: Rethinking benchmarking in nlp. arXiv preprint arXiv:2104.14337, 2021.
  15. 15.Xiang Lisa Li, Evan Zheran Liu, Percy Liang, and Tatsunori Hashimoto. Autobencher: Creating salient, novel, difficult datasets for language models. arXiv preprint arXiv:2407.08351, 2024a.
  16. 16.Yuangang Li, Jiaqi Li, Zhuo Xiao, Tiankai Yang, Yi Nian, Xiyang Hu, and Yue Zhao. Nlp-adbench: Nlp anomaly detection benchmark. arXiv preprint arXiv:2412.04784, 2024b.
  17. 17.Renjie Liang, Li Li, Chongzhi Zhang, Jing Wang, Xizhou Zhu, and Aixin Sun. Tvr-ranking: A dataset for ranked video moment retrieval with imprecise queries, 2024. URL https://arxiv.org/abs/2407.06597.
  18. 18.Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, et al. Deepseek-v3 technical report. arXiv preprint arXiv:2412.19437, 2024.
  19. 19.Cong Lu, Shengran Hu, and Jeff Clune. Automated capability discovery via model self-exploration. arXiv preprint arXiv:2502.07577, 2025.
  20. 20.Potsawee Manakul, Adian Liusie, and Mark J. F. Gales. Selfcheckgpt: Zero-resource black-box hallucination detection for generative large language models, 2023. URL https://arxiv.org/abs/2303.08896.
  21. 21.Sara Merken. Ai ’hallucinations’ in court papers spell trouble for lawyers, 2025. URL https://www.reuters.com/technology/artificial-intelligence/ai-hallucinations-court-papers-spell-trouble-lawyers-2025-02-18/. Accessed: 2025-03-21.
  22. 22.OpenAI. Gpt-4o mini: advancing cost-efficient intelligence. https://openai.com/index/gpt-4o-mini-advancing-cost-efficient-intelligence/, 2024a.
  23. 23.OpenAI. Openai o1 system card, 2024b. URL https://cdn.openai.com/o1-system-card.pdf.
  24. 24.OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mohammad Bavarian, Jeff Belgum, Irwan Bello, Jake Berdine, Gabriel Bernadett-Shapiro, Christopher Berner, Lenny Bogdonoff, Oleg Boiko, Madelaine Boyd, Anna-Luisa Brakman, Greg Brockman, Tim Brooks, Miles Brundage, Kevin Button, Trevor Cai, Rosie Campbell, Andrew Cann, Brittany Carey, Chelsea Carlson, Rory Carmichael, Brooke Chan, Che Chang, Fotis Chantzis, Derek Chen, Sully Chen, Ruby Chen, Jason Chen, Mark Chen, Ben Chess, Chester Cho, Casey Chu, Hyung Won Chung, Dave Cummings, Jeremiah Currier, Yunxing Dai, Cory Decareaux, Thomas Degry, Noah Deutsch, Damien Deville, Arka Dhar, David Dohan, Steve Dowling, Sheila Dunning, Adrien Ecoffet, Atty Eleti, Tyna Eloundou, David Farhi, Liam Fedus, Niko Felix, Simón Posada Fishman, Juston Forte, Isabella Fulford, Leo Gao, Elie Georges, Christian Gibson, Vik Goel, Tarun Gogineni, Gabriel Goh, Rapha Gontijo-Lopes, Jonathan Gordon, Morgan Grafstein, Scott Gray, Ryan Greene, Joshua Gross, Shixiang Shane Gu, Yufei Guo, Chris Hallacy, Jesse Han, Jeff Harris, Yuchen He, Mike Heaton, Johannes Heidecke, Chris Hesse, Alan Hickey, Wade Hickey, Peter Hoeschele, Brandon Houghton, Kenny Hsu, Shengli Hu, Xin Hu, Joost Huizinga, Shantanu Jain, Shawn Jain, Joanne Jang, Angela Jiang, Roger Jiang, Haozhun Jin, Denny Jin, Shino Jomoto, Billie Jonn, Heewoo Jun, Tomer Kaftan, Łukasz Kaiser, Ali Kamali, Ingmar Kanitscheider, Nitish Shirish Keskar, Tabarak Khan, Logan Kilpatrick, Jong Wook Kim, Christina Kim, Yongjik Kim, Jan Hendrik Kirchner, Jamie Kiros, Matt Knight, Daniel Kokotajlo, Łukasz Kondraciuk, Andrew Kondrich, Aris Konstantinidis, Kyle Kosic, Gretchen Krueger, Vishal Kuo, Michael Lampe, Ikai Lan, Teddy Lee, Jan Leike, Jade Leung, Daniel Levy, Chak Ming Li, Rachel Lim, Molly Lin, Stephanie Lin, Mateusz Litwin, Theresa Lopez, Ryan Lowe, Patricia Lue, Anna Makanju, Kim Malfacini, Sam Manning, Todor Markov, Yaniv Markovski, Bianca Martin, Katie Mayer, Andrew Mayne, Bob McGrew, Scott Mayer McKinney, Christine McLeavey, Paul McMillan, Jake McNeil, David Medina, Aalok Mehta, Jacob Menick, Luke Metz, Andrey Mishchenko, Pamela Mishkin, Vinnie Monaco, Evan Morikawa, Daniel Mossing, Tong Mu, Mira Murati, Oleg Murk, David Mély, Ashvin Nair, Reiichiro Nakano, Rajeev Nayak, Arvind Neelakantan, Richard Ngo, Hyeonwoo Noh, Long Ouyang, Cullen O’Keefe, Jakub Pachocki, Alex Paino, Joe Palermo, Ashley Pantuliano, Giambattista Parascandolo, Joel Parish, Emy Parparita, Alex Passos, Mikhail Pavlov, Andrew Peng, Adam Perelman, Filipe de Avila Belbute Peres, Michael Petrov, Henrique Ponde de Oliveira Pinto, Michael, Pokorny, Michelle Pokrass, Vitchyr H. Pong, Tolly Powell, Alethea Power, Boris Power, Elizabeth Proehl, Raul Puri, Alec Radford, Jack Rae, Aditya Ramesh, Cameron Raymond, Francis Real, Kendra Rimbach, Carl Ross, Bob Rotsted, Henri Roussez, Nick Ryder, Mario Saltarelli, Ted Sanders, Shibani Santurkar, Girish Sastry, Heather Schmidt, David Schnurr, John Schulman, Daniel Selsam, Kyla Sheppard, Toki Sherbakov, Jessica Shieh, Sarah Shoker, Pranav Shyam, Szymon Sidor, Eric Sigler, Maddie Simens, Jordan Sitkin, Katarina Slama, Ian Sohl, Benjamin Sokolowsky, Yang Song, Natalie Staudacher, Felipe Petroski Such, Natalie Summers, Ilya Sutskever, Jie Tang, Nikolas Tezak, Madeleine B. Thompson, Phil Tillet, Amin Tootoonchian, Elizabeth Tseng, Preston Tuggle, Nick Turley, Jerry Tworek, Juan Felipe Cerón Uribe, Andrea Vallone, Arun Vijayvergiya, Chelsea Voss, Carroll Wainwright, Justin Jay Wang, Alvin Wang, Ben Wang, Jonathan Ward, Jason Wei, CJ Weinmann, Akila Welihinda, Peter Welinder, Jiayi Weng, Lilian Weng, Matt Wiethoff, Dave Willner, Clemens Winter, Samuel Wolrich, Hannah Wong, Lauren Workman, Sherwin Wu, Jeff Wu, Michael Wu, Kai Xiao, Tao Xu, Sarah Yoo, Kevin Yu, Qiming Yuan, Wojciech Zaremba, Rowan Zellers, Chong Zhang, Marvin Zhang, Shengjia Zhao, Tianhao Zheng, Juntang Zhuang, William Zhuk, and Barret Zoph. Gpt-4 technical report, 2024. URL https://arxiv.org/abs/2303.08774.
  25. 25.Fabio Petroni, Tim Rocktäschel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, Alexander H Miller, and Sebastian Riedel. Language models as knowledge bases? arXiv preprint arXiv:1909.01066, 2019.
  26. 26.Long Phan, Alice Gatti, Ziwen Han, Nathaniel Li, Josephina Hu, Hugh Zhang, Chen Bo Calvin Zhang, Mohamed Shaaban, John Ling, Sean Shi, et al. Humanity’s last exam. arXiv preprint arXiv:2501.14249, 2025.
  27. 27.Qwen, :, An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, Keqin Bao, Kexin Yang, Le Yu, Mei Li, Mingfeng Xue, Pei Zhang, Qin Zhu, Rui Men, Runji Lin, Tianhao Li, Tianyi Tang, Tingyu Xia, Xingzhang Ren, Xuancheng Ren, Yang Fan, Yang Su, Yichang Zhang, Yu Wan, Yuqiong Liu, Zeyu Cui, Zhenru Zhang, and Zihan Qiu. Qwen2.5 technical report, 2025. URL https://arxiv.org/abs/2412.15115.
  28. 28.Nils Reimers and Iryna Gurevych. Sentence-bert: Sentence embeddings using siamese bert-networks. arXiv preprint arXiv:1908.10084, 2019.
  29. 29.Taiwei Shi, Kai Chen, and Jieyu Zhao. Safer-instruct: Aligning language models with automated preference data, 2024. URL https://arxiv.org/abs/2311.08685.
  30. 30.Taiwei Shi, Zhuoer Wang, Longqi Yang, Ying-Chun Lin, Zexue He, Mengting Wan, Pei Zhou, Sujay Jauhar, Sihao Chen, Shan Xia, Hongfei Zhang, Jieyu Zhao, Xiaofeng Xu, Xia Song, and Jennifer Neville. Wildfeedback: Aligning llms with in-situ user interactions and feedback, 2025. URL https://arxiv.org/abs/2408.15549.
  31. 31.Taylor Shin, Yasaman Razeghi, Robert L Logan IV, Eric Wallace, and Sameer Singh. Autoprompt: Eliciting knowledge from language models with automatically generated prompts. arXiv preprint arXiv:2010.15980, 2020.
  32. 32.Adi Simhi, Jonathan Herzig, Idan Szpektor, and Yonatan Belinkov. Distinguishing ignorance from error in llm hallucinations. arXiv preprint arXiv:2410.22071, 2024.
  33. 33.Linxin Song, Jieyu Zhang, Lechao Cheng, Pengyuan Zhou, Tianyi Zhou, and Irene Li. Nlpbench: Evaluating large language models on solving nlp problems. arXiv preprint arXiv:2309.15630, 2023.
  34. 34.Zhaochen Su, Jun Zhang, Xiaoye Qu, Tong Zhu, Yanshu Li, Jiashuo Sun, Juntao Li, Min Zhang, and Yu Cheng. Conflictbank: A benchmark for evaluating the influence of knowledge conflicts in llm, 2024. URL https://arxiv.org/abs/2408.12076.
  35. 35.Laurens Van der Maaten and Geoffrey Hinton. Visualizing data using t-sne. Journal of machine learning research, 9(11), 2008.
  36. 36.Siyuan Wang, Zhuohan Long, Zhihao Fan, Zhongyu Wei, and Xuanjing Huang. Benchmark self-evolving: A multi-agent framework for dynamic llm evaluation, 2024a. URL https://arxiv.org/abs/2402.11443.
  37. 37.Xiaoxuan Wang, Ziniu Hu, Pan Lu, Yanqiao Zhu, Jieyu Zhang, Satyen Subramaniam, Arjun R. Loomba, Shichang Zhang, Yizhou Sun, and Wei Wang. Scibench: Evaluating college-level scientific problem-solving abilities of large language models, 2024b. URL https://arxiv.org/abs/2307.10635.
  38. 38.Wikipedia contributors. Wikipedia, the free encyclopedia, 2004. URL https://www.wikipedia.org/. [Online; accessed 22-July-2004].
  39. 39.Siyuan Wu, Yue Huang, Chujie Gao, Dongping Chen, Qihui Zhang, Yao Wan, Tianyi Zhou, Xiangliang Zhang, Jianfeng Gao, Chaowei Xiao, and Lichao Sun. Unigen: A unified framework for textual dataset generation using large language models, 2024. URL https://arxiv.org/abs/2406.18966.
  40. 40.Tiankai Yang, Yi Nian, Shawn Li, Ruiyao Xu, Yuangang Li, Jiaqi Li, Zhuo Xiao, Xiyang Hu, Ryan Rossi, Kaize Ding, et al. Ad-llm: Benchmarking large language models for anomaly detection. arXiv preprint arXiv:2412.11142, 2024.
  41. 41.Lei Yu, Meng Cao, Jackie CK Cheung, and Yue Dong. Mechanistic understanding and mitigation of language model non-factual hallucinations. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (eds.), Findings of the Association for Computational Linguistics: EMNLP 2024, pp. 7943–7956, Miami, Florida, USA, November 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.findings-emnlp.466. URL https://aclanthology.org/2024.findings-emnlp.466/.
  42. 42.Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, et al. Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 9556–9567, 2024.
  43. 43.Zhiyuan Zeng, Yizhong Wang, Hannaneh Hajishirzi, and Pang Wei Koh. Evaltree: Profiling language model weaknesses via hierarchical capability trees, 2025. URL https://arxiv.org/abs/2503.08893.
  44. 44.Jieyu Zhang, Weikai Huang, Zixian Ma, Oscar Michel, Dong He, Tanmay Gupta, Wei-Chiu Ma, Ali Farhadi, Aniruddha Kembhavi, and Ranjay Krishna. Task me anything. arXiv preprint arXiv:2406.11775, 2024a.
  45. 45.Jieyu Zhang, Le Xue, Linxin Song, Jun Wang, Weikai Huang, Manli Shu, An Yan, Zixian Ma, Juan Carlos Niebles, Caiming Xiong, et al. Provision: Programmatically scaling vision-centric instruction data for multimodal language models. arXiv preprint arXiv:2412.07012, 2024b.
  46. 46.Xin Zhang, Yanzhao Zhang, Dingkun Long, Wen Xie, Ziqi Dai, Jialong Tang, Huan Lin, Baosong Yang, Pengjun Xie, Fei Huang, et al. mgte: Generalized long-context text representation and reranking models for multilingual text retrieval. arXiv preprint arXiv:2407.19669, 2024c.
  47. 47.Yue Zhang, Yafu Li, Leyang Cui, Deng Cai, Lemao Liu, Tingchen Fu, Xinting Huang, Enbo Zhao, Yu Zhang, Yulong Chen, Longyue Wang, Anh Tuan Luu, Wei Bi, Freda Shi, and Shuming Shi. Siren’s song in the ai ocean: A survey on hallucination in large language models, 2023. URL https://arxiv.org/abs/2309.01219.
  48. 48.Yu Zhao, Alessio Devoto, Giwon Hong, Xiaotang Du, Aryo Pradipta Gema, Hongru Wang, Xuanli He, Kam-Fai Wong, and Pasquale Minervini. Steering knowledge selection behaviours in llms via sae-based representation engineering, 2025. URL https://arxiv.org/abs/2410.15999.
  49. 49.Kaijie Zhu, Jiaao Chen, Jindong Wang, Neil Zhenqiang Gong, Diyi Yang, and Xing Xie. Dyval: Dynamic evaluation of large language models for reasoning tasks, 2024a. URL https://arxiv.org/abs/2309.17167.
  50. 50.Kaijie Zhu, Jindong Wang, Qinlin Zhao, Ruochen Xu, and Xing Xie. Dynamic evaluation of large language models by meta probing agents, 2024b. URL https://arxiv.org/abs/2402.14865.

Citation

MLA
Song, L., et al. “Discovering Knowledge Deficiencies of Language Models on Massive Knowledge Base”. arXiv, 2025, http://arxiv.org/abs/2503.23361v1.
APA
Song, L., Ding, X., Zhang, J., Shi, T., Shimizu, R., Gupta, R., Liu, Y., Kang, J., & Zhao, J. (2025). Discovering Knowledge Deficiencies of Language Models on Massive Knowledge Base. arXiv. http://arxiv.org/abs/2503.23361v1
Chicago
Song, L., X. Ding, J. Zhang, et al. 2025. “Discovering Knowledge Deficiencies of Language Models on Massive Knowledge Base”. arXiv. http://arxiv.org/abs/2503.23361v1.
Harvard
Song, L. et al. (2025) “Discovering Knowledge Deficiencies of Language Models on Massive Knowledge Base”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2503.23361v1.
Vancouver
1. Song L, Ding X, Zhang J, Shi T, Shimizu R, Gupta R, Liu Y, Kang J, Zhao J (2025) Discovering Knowledge Deficiencies of Language Models on Massive Knowledge Base. arXiv

BibTeX

@article{song2025discovering,
  title = {Discovering Knowledge Deficiencies of Language Models on Massive Knowledge Base},
  author = {Song, Linxin and Ding, Xuwei and Zhang, Jieyu and Shi, Taiwei and Shimizu, Ryotaro and Gupta, Rahul and Liu, Yang and Kang, Jian and Zhao, Jieyu},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2503.23361v1},
  eprint = {2503.23361}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/