When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories

Alex Troy MallenAkari AsaiVictor ZhongRajarshi DasHannaneh HajishirziDaniel Khashabi

article2022ACL1,454 citations

Reveals that scaling language models fails to resolve factual errors on long-tail knowledge and introduces an adaptive retrieval strategy on the PopQA benchmark that queries external memory only when needed, significantly cutting inference costs while improving factual accuracy.

Listen

Large language models often struggle to reliably answer questions that require fine-grained, real-world factual knowledge, leading to hallucinations and factual inaccuracies. The article investigates the limits of language models in memorizing factual knowledge, evaluates how scaling model size compares against augmenting models with external search retrieval, and introduces a method to selectively deploy retrieval only when needed.

The authors conducted a large-scale evaluation using ten language models across three model families—ranging from 1.3 billion parameters to large-scale GPT-3 models—tested on open-domain question answering. To evaluate factual knowledge across different frequency levels, the authors introduced PopQA, a new dataset of 14,000 entity-centric questions mapped against Wikipedia monthly page views as a proxy for subject popularity, and cross-evaluated findings on the EntityQuestions dataset. The analysis evaluated both unassisted model memory and semi-parametric configurations augmented with search retrievers such as BM25 and Contriever.

The investigation produced four primary findings. First, a language model's factual accuracy strongly correlates with subject entity popularity; models reliably recall well-known facts but struggle heavily with less frequent, long-tail entities. Second, simply scaling up model parameters fails to resolve long-tail gaps: for the 4,000 least popular questions in PopQA, increasing model size from 6 billion parameters to GPT-3 only improved accuracy from 16% to 19%. Third, external retrieval dramatically improves accuracy on rare entities—allowing a 2.7-billion-parameter model to outperform unassisted GPT-3 on tail knowledge—yet retrieval actually degrades GPT-3's accuracy by 10% on popular entities due to irrelevant or misleading search passages. Fourth, an "Adaptive Retrieval" strategy, which consults external documents only when entity popularity falls below a specific threshold, improved GPT-3 accuracy by up to 5.3% over fixed retrieval pipelines while cutting GPT-3 inference API costs by roughly half.

These findings indicate that relying solely on increasing parameter size is an inefficient and ineffective approach for mastering long-tail world knowledge. Organizations deploying language models should not assume that frontier-scale models are factual across all domains, nor should they blindly attach retrieval pipelines to every prompt. Instead, engineering teams should implement adaptive retrieval systems that evaluate query characteristics beforehand to reduce infrastructure latency, lower operating expenses, and prevent misleading search contexts from confusing the underlying model.

Decision-makers should note that these experiments relied on synthetic, template-generated questions and used monthly Wikipedia page views as a static proxy for entity popularity. While confidence in the core trade-off between model scale, popularity, and retrieval utility is high, further validation on complex, real-world enterprise queries and alternative domain-specific popularity metrics is recommended before final production deployment.

Cover for When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories

Abstract

Despite their impressive performance on diverse tasks, large language models (LMs) still struggle with tasks requiring rich world knowledge, implying the limitations of relying solely on their parameters to encode a wealth of world knowledge. This paper aims to understand LMs' strengths and limitations in memorizing factual knowledge, by conducting large-scale knowledge probing experiments of 10 models and 4 augmentation methods on PopQA, our new open-domain QA dataset with 14k questions. We find that LMs struggle with less popular factual knowledge, and that scaling fails to appreciably improve memorization of factual knowledge in the long tail. We then show that retrieval-augmented LMs largely outperform orders of magnitude larger LMs, while unassisted LMs remain competitive in questions about high-popularity entities. Based on those findings, we devise a simple, yet effective, method for powerful and efficient retrieval-augmented LMs, which retrieves non-parametric memories only when necessary. Experimental results show that this significantly improves models' performance while reducing the inference costs.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Evaluation Setup
  • 3.1 Focus and Task
  • 3.2 Dimensions of Analysis
  • 3.3 Benchmarks
  • 4 Memorization Depends on Popularity and Relationship Type
  • 4.1 Experimental Setup
  • 4.2 Results
  • 5 Non-parametric Memory Complements Parametric Memory
  • 5.1 Experimental Setup
  • 5.2 Results
  • 6 Adaptive Retrieval: Using Retrieval Only Where It Helps
  • 6.1 Method
  • 6.2 Results
  • 7 Discussion and Conclusions
  • References
  • A Details of PopQA Constructions
  • B Experimental Details
  • C Detailed Results
  • C.1 LM results
  • C.2 Retrieval-augmented LM results
  • C.3 Qualitative Results

Knowls

  1. Knowl 1 — Adaptive Retrieval for Factual Question Answering

    algorithm

    Adaptive Retrieval is a hybrid inference strategy that dynamically decides whether to query an external non-parametric retrieval system or rely on a language model's parametric knowledge based on the popularity of the subject entity and its relationship type.

    Input: Input question qq, subject entity SS, relation type RR, entity popularity p(S)p(S), per-relation popularity thresholds {τr}\{\tau_r\}, language model LMLM, retriever RetrieverRetriever
    Output: Predicted answer aa
    if p(S)<τRp(S) < \tau_R then
        D←Retriever(q)D \leftarrow Retriever(q)
        a←LM(prompt(D,q))a \leftarrow LM(\text{prompt}(D, q))
    else
        a←LM(prompt(q))a \leftarrow LM(\text{prompt}(q))
    end if
    return aa

    For each relation type rr, the popularity threshold τr\tau_r is determined on a tuning/development set by performing a brute-force search over observed entity popularities to maximize adaptive accuracy:

    τr=arg⁡max⁡τ1∣Qr∣∑qi∈QrI(yi∈pred(qi,τ))\tau_r = \arg\max_{\tau} \frac{1}{|Q_r|} \sum_{q_i \in Q_r} \mathbb{I}\left( y_i \in \text{pred}(q_i, \tau) \right)

    where QrQ_r is the set of evaluation questions belonging to relation rr, yiy_i is the set of valid reference answers for question qiq_i, I(⋅)\mathbb{I}(\cdot) is the indicator function, and pred(qi,τ)\text{pred}(q_i, \tau) outputs the prediction of the retrieval-augmented system when p(Si)<τp(S_i) < \tau and the prediction of the unassisted parametric model (or a parametric context generator such as GenRead) when p(Si)≥τp(S_i) \ge \tau.

  2. Knowl 2 — PopQA Benchmark Construction and Sampling

    model/method

    PopQA is an open-domain, entity-centric question answering dataset of 14,267 questions designed to probe parametric and non-parametric memory across 16 Wikidata relationship types with a controlled long-tail popularity distribution.

    To prevent the dataset from being completely dominated by obscure entities while still covering the long tail, candidate knowledge triples (S,R,O)(S, R, O) sampled from Wikidata are filtered according to the exact-match term frequency ff of the subject entity's aliases in an 800 MB random slice of the C4 corpus. A candidate triple is retained if and only if:

    f>exp⁡(8Rrand−6)f > \exp(8R_{\text{rand}} - 6)

    where Rrand∼U(0,1)R_{\text{rand}} \sim \mathcal{U}(0, 1) is a uniform pseudo-random variable. Sampling stops once 2,000 triples have been collected for any single relationship type.

    Each triple is converted into a natural language question using a handcrafted template specific to relationship RR (e.g., "What is [subj]'s occupation?", "Who was the director of [subj]?"). Entity popularity p(S)p(S) is quantified as the monthly Wikipedia page views of the subject entity SS via the Wikipedia API.

  3. Knowl 3 — Correlation Between Subject Entity Popularity and Parametric Memorization

    empirical result

    Closed-book language model question-answering accuracy exhibits a consistent, positive correlation with the subject entity's popularity, measured by log monthly Wikipedia page views. Across 10 autoregressive language models from three model families—OPT (1.3B, 2.7B, 6.7B, 13B), GPT-Neo (GPT-Neo 1.3B, GPT-Neo 2.7B, GPT-J 6B, GPT-NeoX 20B), and GPT-3 (davinci-002, davinci-003):

    • The Pearson correlation between log⁡10(popularity)\log_{10}(\text{popularity}) and question-answering accuracy is positive across virtually all evaluated relation types.
    • Correlation strength grows with model capacity: GPT-Neo 1.3B displays a weak correlation (r≈0.1r \approx 0.1), whereas GPT-3 davinci-003 exhibits a strong correlation (r≈0.4r \approx 0.4).
    • Subject entity popularity measured from Wikipedia page views serves as an input-level predictor of whether an LM has memorized a factual triple in its parameters.
  4. Knowl 4 — Language Model Parameter Scaling and the Long Tail of Factual Knowledge

    empirical result

    Increasing model parameter size from 1.3B to ~175B improves closed-book question answering accuracy almost exclusively on high-popularity entities (log⁡10(popularity)>4\log_{10}(\text{popularity}) > 4), while yielding negligible accuracy gains on long-tail entities.

    When evaluated on the 4,000 least popular entity questions in PopQA:

    • GPT-J (6B parameters) attains 16% accuracy.
    • GPT-NeoX (20B parameters) attains 16% accuracy.
    • GPT-3 davinci-003 (~175B parameters) attains 19% accuracy.

    While scaling reduces the entity popularity threshold required for reliable memorization, it fails to extend parametric recall deep into the long tail at practical model scales.

  5. Knowl 5 — Retrieval Augmentation for Tail Entity Knowledge

    empirical result

    Augmenting language models with non-parametric retrieval (prepending the top-1 retrieved passage from a Wikipedia index using BM25 or Contriever to the input prompt) substantially improves QA accuracy on long-tail entities across all model sizes.

    On the least popular questions in PopQA:

    • Retrieval-augmented smaller models outperform much larger unassisted parametric models; for example, Contriever-augmented GPT-Neo 2.7B surpasses vanilla GPT-3 davinci-003 on the 4,000 least popular queries.
    • For extremely low-popularity entities, static term-based retrieval (BM25) outperforms dense retrieval (Contriever), whereas Contriever achieves higher accuracy overall across the entire dataset.
  6. Knowl 6 — Retrieval-Induced Errors on Well-Memorized Entities

    empirical result

    For questions about popular entities that a large language model has already memorized in its parameters, retrieval augmentation can degrade performance by introducing irrelevant, ambiguous, or distracting context.

    On PopQA evaluated with GPT-3 davinci-003 and Contriever:

    • In 10% of all questions, unassisted GPT-3 answers correctly, but the Contriever-augmented version produces an incorrect prediction.
    • For this failure subset, the top-1 retrieval recall (recall@1, measuring whether a gold answer string appears in the retrieved passage) is only 0.14, compared to an average recall@1 of 0.42 across the entire benchmark.
    • Qualitative error analysis reveals that retrieval failure typically arises from homonyms or namesake entities (e.g., retrieving an article about a different person sharing the subject's name), which misleads the LM into overriding its correct parametric belief.
  7. Knowl 7 — Contriever Retrieval Outcome Contingency on PopQA with GPT-3

    data/table

    The interaction between unassisted parametric generation by GPT-3 davinci-003 and Contriever retrieval augmentation on the PopQA dataset is characterized by the four-way breakdown of model success and failure alongside retriever recall@1:

    Contriever-Augmented LM Succeeded Contriever-Augmented LM Failed
    LM Succeeded 0.83 (24%) 0.14 (10%)
    LM Failed 0.88 (17%) 0.11 (49%)

    Each entry reports the retriever recall@1 (proportion of top-1 retrieved documents containing a gold answer string) for that category, with the proportion of total PopQA questions falling into that quadrant given in parentheses. Retrieval proves beneficial on 17% of queries where retriever recall@1 is high (0.88), but causes regressions on 10% of queries where poor retrieval quality (0.14 recall@1) misleads an otherwise correct model.

  8. Knowl 8 — Accuracy and Efficiency Gains of Adaptive Retrieval

    empirical result

    Adaptive Retrieval achieves higher QA accuracy and lower inference computational costs compared to universal retrieval augmentation.

    • Accuracy: Pairing GPT-3 davinci-003 adaptively with GenRead and Contriever on PopQA achieves 46.5% accuracy, representing a 5.3% absolute improvement over any non-adaptive retrieval baseline.
    • Query Reduction: Because larger models memorize popular entities, they require retrieval far less frequently. GPT-3 davinci-003 with BM25 triggers retrieval on only 40% of PopQA questions, and GPT-NeoX 20B skips retrieval on more than 20% of questions.
    • Cost and Latency: Adaptive Retrieval reduces GPT-3 API costs by up to 50% on PopQA and 15% on EntityQuestions, and reduces inference latency on open-source models (such as GPT-J 6B and GPT-NeoX 20B) by up to 9% compared to retrieving for every query.
  9. Knowl 9 — Impact of Relationship Type and Surface Artifacts on Knowledge Memorization

    empirical result

    The degree to which language models memorize factual knowledge and the correlation between accuracy and entity popularity vary significantly across relationship types:

    • Constrained / Artifact-Heavy Relations: Relationship types such as "country" and "sport" exhibit high baseline accuracy across all model sizes and little-to-no correlation with subject entity popularity. Models exploit surface-level artifacts in subject entity names (e.g., extracting nationality from a person's name) or guess dominant category labels (e.g., predicting the most frequent sport or country).
    • Open-Domain Relations: Relations with broad answer spaces and minimal surface clues (such as "author", "director", "occupation", and "screenwriter") have low baseline zero-shot accuracy and high positive correlation between entity popularity and model accuracy, indicating genuine dependence on parametric factual memorization.
  10. Knowl 10 — Evaluation Constraints of Popularity Metrics and Synthetic QA Data

    limitation

    The methodology and findings have several explicit limitations:

    • Synthetic Formulation: PopQA is constructed by verbalizing structured Wikidata triples (S,R,O)(S, R, O) using fixed natural language templates, which may not capture the linguistic variability, nuance, and complexity found in organic, naturally occurring user queries.
    • Time-Dependent Popularity Proxy: Subject entity popularity is estimated using Wikipedia monthly page views, which fluctuates over time and is an indirect proxy for the true token frequency of entities in private and public pretraining corpora.
    • Knowledge Scope: The investigation is restricted to entity-centric factual knowledge triples and does not evaluate other categories of world knowledge, such as commonsense, procedural, or temporal reasoning.

Coverage note — None was omitted; all primary contributions—including PopQA construction, empirical analyses of memorization vs. scale and popularity, retrieval dynamics on long-tail vs. popular entities, the Adaptive Retrieval method, and efficiency benchmarks—are covered.

References

  1. 1.Nancy E Adams. 2015. Bloom’s taxonomy of cognitive learning objectives. Journal of the Medical Library Association.
  2. 2.Akari Asai, Matt Gardner, and Hannaneh Hajishirzi. 2022. Evidentiality-guided generation for knowledge-intensive NLP tasks. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies.
  3. 3.Payal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen, et al. 2016. MS MARCO: A human generated machine reading comprehension dataset.
  4. 4.Sidney Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, Michael Pieler, Usvsn Sai Prashanth, Shivanshu Purohit, Laria Reynolds, Jonathan Tow, Ben Wang, and Samuel Weinbach. 2022. GPT-NeoX-20B: An open-source autoregressive language model. In Proceedings of BigScience Episode #5 – Workshop on Challenges & Perspectives in Creating Large Language Models.
  5. 5.Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George Bm Van Den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, Diego De Las Casas, Aurelia Guy, Jacob Menick, Roman Ring, Tom Hennigan, Saffron Huang, Loren Maggiore, Chris Jones, Albin Cassirer, Andy Brock, Michela Paganini, Geoffrey Irving, Oriol Vinyals, Simon Osindero, Karen Simonyan, Jack Rae, Erich Elsen, and Laurent Sifre. 2022. Improving language models by retrieving from trillions of tokens. In Proceedings of the 39th International Conference on Machine Learning.
  6. 6.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. In Advances in Neural Information Processing systems.
  7. 7.Boxi Cao, Hongyu Lin, Xianpei Han, Le Sun, Lingyong Yan, Meng Liao, Tong Xue, and Jin Xu. 2021. Knowledgeable or educated guess? revisiting language models as knowledge bases. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing.
  8. 8.Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramèr, and Chiyuan Zhang. 2022. Quantifying memorization across neural language models.
  9. 9.Hung-Ting Chen, Michael JQ Zhang, and Eunsol Choi. 2022. Rich knowledge sources bring complex knowledge conflicts: Recalibrating models to reflect conflicting evidence. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing.
  10. 10.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. PaLM: Scaling language modeling with pathways.
  11. 11.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies.
  12. 12.Paolo Ferragina and Ugo Scaiella. 2010. TAGME: on-the-fly annotation of short text fragments (by wikipedia entities). In Proceedings of the 19th ACM international conference on Information and knowledge management.
  13. 13.Thibault Févry, Livio Baldini Soares, Nicholas FitzGerald, Eunsol Choi, and Tom Kwiatkowski. 2020. Entities as experts: Sparse memory access with entity supervision. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing.
  14. 14.Junxian He, Graham Neubig, and Taylor Berg-Kirkpatrick. 2021. Efficient nearest neighbor language models. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing.
  15. 15.Jie Huang, Hanyin Shao, and Kevin Chen-Chuan Chang. 2022. Are large pre-trained language models leaking your personal information? In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing.
  16. 16.Gautier Izacard, Mathilde Caron, Lucas Hosseini, Sebastian Riedel, Piotr Bojanowski, Armand Joulin, and Edouard Grave. 2022a. Unsupervised dense information retrieval with contrastive learning. Transactions on Machine Learning Research.
  17. 17.Gautier Izacard and Edouard Grave. 2021. Leveraging passage retrieval with generative models for open domain question answering. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics.
  18. 18.Gautier Izacard, Patrick Lewis, Maria Lomeli, Lucas Hosseini, Fabio Petroni, Timo Schick, Jane Dwivedi-Yu, Armand Joulin, Sebastian Riedel, and Edouard Grave. 2022b. Few-shot learning with retrieval augmented language models.
  19. 19.Joel Jang, Seonghyeon Ye, Changho Lee, Sohee Yang, Joongbo Shin, Janghoon Han, Gyeonghun Kim, and Minjoon Seo. 2022. Temporalwiki: A lifelong benchmark for training and evaluating ever-evolving language models. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing.
  20. 20.Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield Dodds, Nova DasSarma, Eli Tran-Johnson, et al. 2022. Language models (mostly) know what they know.
  21. 21.Nikhil Kandpal, Haikang Deng, Adam Roberts, Eric Wallace, and Colin Raffel. 2022. Large language models struggle to learn long-tail knowledge.
  22. 22.Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense passage retrieval for open-domain question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing.
  23. 23.Jungo Kasai, Keisuke Sakaguchi, Yoichi Takahashi, Ronan Le Bras, Akari Asai, Xinyan Yu, Dragomir Radev, Noah A Smith, Yejin Choi, and Kentaro Inui. 2022. Realtime QA: What’s the answer right now?
  24. 24.Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis. 2020. Generalization through memorization: Nearest neighbor language models. In International Conference on Learning Representations.
  25. 25.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019. Natural questions: A benchmark for question answering research. Transactions of the Association for Computational Linguistics.
  26. 26.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. In Advances in Neural Information Processing Systems.
  27. 27.Shayne Longpre, Kartik Perisetla, Anthony Chen, Nikhil Ramesh, Chris DuBois, and Sameer Singh. 2021. Entity-based knowledge conflicts in question answering. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7052–7063, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  28. 28.Sewon Min, Weijia Shi, Mike Lewis, Xilun Chen, Wen-tau Yih, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2022. Nonparametric masked language model.
  29. 29.Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel Bowman. 2022. BBQ: A hand-built bias benchmark for question answering. In Findings of the Association for Computational Linguistics: ACL 2022.
  30. 30.Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019. Language models as knowledge bases? In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing.
  31. 31.Nina Poerner, Ulli Waltinger, and Hinrich Schütze. 2020. E-BERT: Efficient-yet-effective entity embeddings for BERT. In Findings of the Association for Computational Linguistics: EMNLP 2020.
  32. 32.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research.
  33. 33.Yasaman Razeghi, Robert L Logan IV, Matt Gardner, and Sameer Singh. 2022. Impact of pretraining term frequencies on few-shot reasoning.
  34. 34.Adam Roberts, Colin Raffel, and Noam Shazeer. 2020. How much knowledge can you pack into the parameters of a language model? In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing.
  35. 35.Stephen Robertson, Hugo Zaragoza, et al. 2009. The probabilistic relevance framework: Bm25 and beyond. Foundations and Trends in Information Retrieval.
  36. 36.Christopher Sciavolino, Zexuan Zhong, Jinhyuk Lee, and Danqi Chen. 2021. Simple entity-centric questions challenge dense retrievers. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing.
  37. 37.Kurt Shuster, Spencer Poff, Moya Chen, Douwe Kiela, and Jason Weston. 2021. Retrieval augmentation reduces hallucination in conversation. In Findings of the Association for Computational Linguistics: EMNLP 2021.
  38. 38.Zhiqing Sun, Xuezhi Wang, Yi Tay, Yiming Yang, and Denny Zhou. 2022. Recitation-augmented language models.
  39. 39.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed H. Chi, Quoc V Le, and Denny Zhou. 2022. Chain of thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems.
  40. 40.Wenhao Yu, Dan Iter, Shuohang Wang, Yichong Xu, Mingxuan Ju, Soumya Sanyal, Chenguang Zhu, Michael Zeng, and Meng Jiang. 2022. Generate rather than retrieve: Large language models are strong context generators.
  41. 41.Aohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang, Hanyu Lai, Ming Ding, Zhuoyi Yang, Yifan Xu, Wendi Zheng, Xiao Xia, et al. 2022. Glm-130b: An open bilingual pre-trained model.
  42. 42.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. 2022. Opt: Open pre-trained transformer language models.

Citation

MLA
Mallen, A., et al. “When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 9802–22, https://doi.org/10.18653/v1/2023.acl-long.546.
APA
Mallen, A., Asai, A., Zhong, V., Das, R., Khashabi, D., & Hajishirzi, H. (2023). When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 9802–9822. https://doi.org/10.18653/v1/2023.acl-long.546
Chicago
Mallen, A., A. Asai, V. Zhong, R. Das, D. Khashabi, and H. Hajishirzi. 2023. “When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 9802–22. https://doi.org/10.18653/v1/2023.acl-long.546.
Harvard
Mallen, A. et al. (2023) “When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 9802–9822. Available at: https://doi.org/10.18653/v1/2023.acl-long.546.
Vancouver
1. Mallen A, Asai A, Zhong V, Das R, Khashabi D, Hajishirzi H (2023) When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 9802–9822

BibTeX

@inproceedings{mallen-etal-2023-trust,
    title = "When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories",
    author = "Mallen, Alex  and
      Asai, Akari  and
      Zhong, Victor  and
      Das, Rajarshi  and
      Khashabi, Daniel  and
      Hajishirzi, Hannaneh",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.546/",
    doi = "10.18653/v1/2023.acl-long.546",
    pages = "9802--9822"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/