Astute RAG: Overcoming Imperfect Retrieval Augmentation and Knowledge Conflicts for Large Language Models

Fei WangXingchen WanRuoxi SunJiefeng ChenSercan . Arik

article2025ACL79 citations

Proposes Astute RAG, a source-aware framework that resolves knowledge conflicts between large language models and noisy retrieved documents to prevent performance degradation during retrieval failures.

Listen

Retrieval-augmented generation (RAG) is widely adopted to enhance large language models by supplying external information from search engines or corporate databases. However, real-world retrieval is often imperfect, frequently introducing irrelevant, misleading, or conflicting text. When retrieved data contradicts a model's internal learned knowledge, language models often falter and produce inaccurate outputs, posing significant operational and reliability risks for automated decision-making systems.

The article evaluates the negative impacts of imperfect retrieval and knowledge conflicts on language model performance and demonstrates a new, training-free methodology called ASTUTE RAG. This approach aims to systematically resolve conflicts between external search results and internal model knowledge to generate more accurate and dependable answers.

To evaluate this challenge, the authors conducted a controlled study on over 1,000 realistic question-answer pairs spanning general, biomedical, and specialized long-tail domains, using Google Search to retrieve live web snippets. They tested proprietary and open-source models—including Claude 3.5 Sonnet, Gemini 1.5 Pro, and Mistral—across multiple baseline techniques. The ASTUTE RAG framework was introduced and evaluated through a three-step process: adaptively generating internal passages from the model, consolidating internal and external knowledge while tracking information sources, and finalizing answers based on source credibility and agreement.

The analysis yielded several critical findings. First, imperfect retrieval is pervasive; approximately 70% of retrieved web passages failed to contain the direct correct answer, leading to knowledge conflicts in roughly 19% of cases. Second, when conflicts occurred, internal model knowledge and external search results corrected each other in near-equal proportions (about 47% versus 53%), proving that neither source is universally reliable on its own. Third, ASTUTE RAG consistently outperformed all competing methods across all evaluated models, achieving overall accuracy gains of 4.1% to 6.9% over the strongest baselines. Finally, in worst-case scenarios where all retrieved documents were unhelpful or misleading, ASTUTE RAG was the only evaluated method that matched or exceeded the accuracy of models operating entirely without retrieval, resolving knowledge conflicts correctly in approximately 80% of conflicting cases.

These findings indicate that simply feeding raw search results into language models introduces severe vulnerabilities, often making system performance worse than using an unaugmented model. ASTUTE RAG mitigates this risk by effectively cross-checking external evidence against internal memory. Because it functions entirely via structured prompting without requiring fine-tuning or heavy computational overhead—incurring less than a 5% increase in token usage—it provides a cost-effective safety mechanism for enterprise applications.

For organizations deploying retrieval-based artificial intelligence, adopting a source-aware consolidation mechanism like ASTUTE RAG is recommended to protect against search noise and hallucinations. Teams should implement explicit conflict-resolution steps before generating final responses rather than relying solely on upstream search rerankers. Future work should focus on validating this approach on long-form document inputs and testing how well it performs with less capable, smaller language models.

arXiv: 2410.07176
Cover for Astute RAG: Overcoming Imperfect Retrieval Augmentation and Knowledge Conflicts for Large Language Models

Abstract

Retrieval augmented generation (RAG), while effectively integrating external knowledge to address the inherent limitations of large language models (LLMs), can be hindered by imperfect retrieval that contain irrelevant, misleading, or even malicious information. Previous studies have rarely connected the behavior of RAG through joint analysis, particularly regarding error propagation coming from imperfect retrieval and potential conflicts between LLMs' internal knowledge and external sources. Through comprehensive and controlled analyses under realistic conditions, we find that imperfect retrieval augmentation is inevitable, common, and harmful. We identify the knowledge conflicts between LLM-internal and external knowledge from retrieval as a bottleneck to overcome imperfect retrieval in the post-retrieval stage of RAG. To address this, we propose Astute RAG, a novel RAG approach designed to be resilient to imperfect retrieval augmentation. It adaptively elicits essential information from LLMs' internal knowledge, iteratively consolidates internal and external knowledge with source-awareness, and finalizes the answer according to information reliability. Our experiments with Gemini and Claude demonstrate the superior performance of Astute RAG compared to previous robustness-enhanced RAG approaches. Specifically, Astute RAG is the only RAG method that achieves performance comparable to or even surpassing conventional use of LLMs under the worst-case scenario. Further analysis reveals the effectiveness of Astute RAG in resolving knowledge conflicts, thereby improving the trustworthiness of RAG.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 The Pitfall of RAG
  • 4 Astute RAG
  • 4.1 Overview
  • 4.2 Adaptive Generation of Internal Knowledge
  • 4.3 Iterative Source-aware Knowledge Consolidation
  • 4.4 Answer Finalization
  • 5 Experiments
  • 5.1 Experimental Settings
  • 5.2 Main Results
  • 5.3 Analyses
  • 6 Conclusion
  • References
  • A Prompt Template for Astute RAG
  • B Data Collection
  • C Comparison with Answer Refinement
  • D Comparison with Context Filtering
  • E Comparison with Context Compression
  • F Influence of passage ordering.

Knowls

  1. Knowl 1 — ASTUTE RAG consolidates internal and retrieved knowledge before answering

    algorithm

    ASTUTE RAG is a training-free, post-retrieval procedure for a query qq, retrieved external passages EE, and a prediction-only language model MM. It first elicits passages from the model’s internal knowledge, then consolidates the two sources while retaining source information, and finally selects an answer using the reliability of the resulting evidence groups. The procedure can be implemented as follows:

    Input: Query q, retrieved passages E, language model M, maximum generated passages m_hat, consolidation rounds t (t >= 1)
    Output: Final answer a
    I = Ask M for at most m_hat accurate, relevant internal-knowledge passages for q
    D0 = E followed by I
    Label every passage in D0 as external or internal
    if t == 1:
        Ask M to consolidate D0 and answer q
    else:
        Set the current consolidated set D to D0
        Repeat t - 1 times:
            Ask M to consolidate q using both D0 and D
            In the new set, cluster mutually consistent information
            Keep conflicting accounts in separate groups
            Remove information irrelevant to q
            Preserve each group’s source labels and original passage provenance
            Set D to the new consolidated set
        Ask M to consolidate and answer using D0 and the latest D
    For each resulting evidence group, ask M to propose an answer and confidence
    Select the answer best supported by source credibility, cross-source agreement,
    frequency of supporting information, and information completeness
    Return the selected answer

    The consolidation instructions require grouping consistent documents into concise summaries, separating conflicting accounts, excluding irrelevant content, and recording whether each resulting document came from internal memory or external retrieval. Answer finalization compares candidate answers grounded in those groups. With the default settings t=1t=1 and m^=1\hat m=1, ASTUTE RAG makes two model calls per instance: one for internal passage generation and one for consolidation plus answering. In general, the described procedure uses t+1t+1 calls.

  2. Knowl 2 — Internal passages are generated adaptively with reliability constraints

    model/method

    ASTUTE RAG elicits the language model’s internal knowledge as passages that can supplement missing information in retrieved documents. The generation instruction asks for accurate, query-relevant information and directs the model to say “I don’t know” when its information is unclear or uncertain, rather than inventing an answer. The model is asked for at most m^\hat m passages, each covering distinct information, and may generate fewer—including none—when it has little reliable information. Thus the generated count mm satisfies 0≤m≤m^0\leq m\leq\hat m and is chosen adaptively rather than fixed in advance.

    In the reported passage-count comparison, increasing the maximum from m^=1\hat m=1 to m^=2\hat m=2 raised the average number of generated passages from 0.690.69 to 1.241.24. Accuracy changed from 52.2052.20 to 52.2052.20 on NQ, 84.1084.10 to 85.1685.16 on TriviaQA, 60.1460.14 to 60.8460.84 on BioASQ, and 44.3844.38 to 43.2643.26 on PopQA; overall accuracy changed from 61.7161.71 to 62.0062.00. The results show that the adaptive procedure generated fewer passages than the allowed maximum on average, and that allowing a second passage produced only a small overall accuracy change.

  3. Knowl 3 — Real-world web retrieval frequently omits answer-bearing information

    experimental setup

    The paper’s controlled retrieval analysis used 1,042 short-form question–answer pairs drawn from NQ, TriviaQA, BioASQ, and PopQA. For each question, Google Search returned up to 30 results; the first 10 accessible websites were selected, and the passage from each site was the paragraph corresponding to the search-result snippet. Questions and answers were not chosen or annotated based on the retrieved results. Retrieval precision was the proportion of passages that directly contained the true answer.

    Under this setup, roughly 70% of retrieved passages did not directly contain a true answer, even though the passages came from a real-world search engine and web corpus. The paper reports this as evidence that imperfect retrieval is common in realistic RAG settings, not merely an artifact of synthetic corruption.

  4. Knowl 4 — Internal and external answers are both often the correct side of a conflict

    empirical result

    In the paper’s Claude 3.5 Sonnet analysis of real-world retrieval, the answers from the model without retrieval and with retrieval had four outcomes: both correct in 45.4%45.4\% of instances, internal answer correct but RAG answer incorrect in 9.1%9.1\%, internal answer incorrect but RAG answer correct in 10.1%10.1\%, and both incorrect in 35.4%35.4\%. Thus, the two answers conflicted in 19.2%19.2\% of instances. Within those conflicting cases, internal knowledge was correct in 47.4%47.4\% and external retrieved knowledge was correct in 52.6%52.6\%.

    The near-even division means neither source can generally be treated as the reliable one in advance: retrieval can correct an LLM’s answer, and the LLM’s internal knowledge can correct retrieval-based errors.

  5. Knowl 5 — Lower retrieval precision generally coincides with more knowledge conflicts

    empirical result

    Across the studied models—Claude, Gemini, Mistral-Large, and Mistral-Nemo—the conflict rate between answers with and without retrieval was generally highest when retrieval precision was low but nonzero, with the peak occurring around 10%10\% precision in the reported analysis. Conflict rates generally fell as precision increased, although the trend had fluctuations. At 0%0\% precision, conflict rates were notably lower than at low-but-nonzero precision; the paper suggests this may occur because such results tend to be irrelevant rather than factually incorrect. The observed trend was shared across the studied models despite their different training processes.

  6. Knowl 6 — ASTUTE RAG improves overall accuracy across four language models

    data/table

    The following zero-shot results report answer accuracy (percent) on four short-form QA datasets and their overall average. The experiments used temperature 00, a maximum output of 1,024 tokens, and string-match accuracy. For each model, the comparison baseline shown is the strongest non-ASTUTE method by overall accuracy in the reported results. ASTUTE RAG has the highest overall score for every model, including the 12B-parameter Mistral-Nemo.

    Model Method NQ TriviaQA BioASQ PopQA Overall
    Claude 3.5 Sonnet No RAG 47.1 82.0 50.4 29.8 54.5
    RAG 44.4 76.7 58.0 36.0 55.5
    InstructRAG 47.1 83.0 58.0 41.0 58.8
    ASTUTE RAG 52.2 84.1 60.1 44.4 61.7
    Gemini 1.5 Pro No RAG 44.8 80.2 45.8 25.3 51.3
    RAG 42.7 76.0 55.2 33.7 53.7
    Self-Route 47.5 79.9 58.0 38.2 57.6
    ASTUTE RAG 50.2 81.6 58.0 40.5 59.2
    Mistral-Large (128B) No RAG 46.8 79.5 43.7 24.7 51.1
    RAG 43.1 77.4 55.9 36.0 54.7
    USC 51.2 80.9 61.5 36.0 59.5
    ASTUTE RAG 50.2 82.7 58.4 42.1 59.9
    Mistral-Nemo (12B) No RAG 29.8 67.8 34.3 23.0 40.2
    RAG 39.3 66.8 49.0 32.6 48.3
    Self-Route 41.4 73.5 51.8 30.9 51.2
    ASTUTE RAG 42.7 73.9 49.3 32.6 51.3
  7. Knowl 7 — ASTUTE RAG remains robust in low-quality retrieval and conflicting-answer cases

    empirical result

    Across retrieval-precision buckets, ASTUTE RAG achieved higher accuracy than the compared RAG alternatives and did not give up its gains when retrieval quality was high. Near zero precision, the other RAG variants fell below the no-retrieval baseline, whereas ASTUTE RAG was the only evaluated RAG method that did not.

    On instances where the answers with and without RAG conflicted, ASTUTE RAG selected the correct answer in approximately 80%80\% of cases, the strongest result among the compared methods. It also improved performance on instances where neither source alone yielded the correct answer, consistent with its ability to combine partially correct information.

    A separate worst-case test used the English refined subset of RGB and supplied five top negative passages per instance. In this setting, RAG was more than 50 accuracy points below the no-RAG model. Baseline robustness methods improved on ordinary RAG but remained substantially below no RAG; ASTUTE RAG was the only tested RAG method reported as reaching performance close to the no-RAG baseline.

  8. Knowl 8 — Providing source labels improves ASTUTE RAG’s overall accuracy

    empirical result

    A Gemini 1.5 Pro ablation compared ASTUTE RAG, which labels passages as internal-memory or external-retrieval evidence during consolidation and answer generation, with a version that omits these source labels. Accuracy is reported as a percentage.

    Method NQ TriviaQA BioASQ PopQA Overall
    ASTUTE RAG 50.2 81.6 58.0 40.5 59.2
    ASTUTE RAG (No Source) 48.1 82.3 57.7 39.9 58.6

    Source labels improved overall accuracy by 0.60.6 points and improved scores on NQ, BioASQ, and PopQA, while TriviaQA decreased by 0.70.7 points. The result supports source awareness as a useful part of consolidation, though the effect was not positive on every dataset.

  9. Knowl 9 — Additional consolidation rounds yield diminishing gains at modest reported cost

    empirical result

    The main experiments used one consolidation round (t=1t=1) to limit inference cost. In the reported analysis, increasing the number of rounds from t=1t=1 to t=2t=2 and then t=3t=3 continued to improve accuracy, but the gain from each additional round diminished. The paper also reports that ASTUTE RAG increased average token use by less than 5%5\% while improving performance by more than 11%11\% compared with the RAG baseline. These cost figures are reported as an aggregate comparison rather than dataset-by-dataset values.

  10. Knowl 10 — ASTUTE RAG depends on capable instruction-following and reasoning models

    limitation

    The paper states that ASTUTE RAG’s effectiveness depends on the underlying LLM having strong instruction-following and reasoning abilities. Its applicability may therefore be more limited with less sophisticated models. The experiments also did not extend to longer input contexts; the authors identify evaluation with longer inputs as an important future direction because imperfect retrieval and knowledge conflicts may become more pronounced in that setting.

Coverage note — The knowls omit detailed results for the secondary comparisons with answer refinement, context filtering, context compression, and passage-ordering strategies, as well as the plotted ASQA results; these analyses are ancillary to the core method, real-world conflict findings, and principal robustness evaluations.

References

  1. 1.Marco Alessio, Guglielmo Faggioli, Nicola Ferro, Franco Maria Nardini, Raffaele Perego, et al. 2024. Improving rag systems via sentence clustering and reordering. In RAG@ SIGIR 2024 workshop: The Information Retrieval’s Role in RAG Systems, ACM, pages 1–10.
  2. 2.Akari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil, and Hannaneh Hajishirzi. 2023. Self-rag: Learning to retrieve, generate, and critique through self-reflection. In The Twelfth International Conference on Learning Representations.
  3. 3.Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al. 2022. Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073.
  4. 4.Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George Bm Van Den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, et al. 2022. Improving language models by retrieving from trillions of tokens. In International conference on machine learning, pages 2206–2240. PMLR.
  5. 5.Jiawei Chen, Hongyu Lin, Xianpei Han, and Le Sun. 2024a. Benchmarking large language models in retrieval-augmented generation. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 17754–17762.
  6. 6.Xinyun Chen, Renat Aksitov, Uri Alon, Jie Ren, Kefan Xiao, Pengcheng Yin, Sushant Prakash, Charles Sutton, Xuezhi Wang, and Denny Zhou. 2024b. Universal self-consistency for large language models. In ICML 2024 Workshop on In-Context Learning.
  7. 7.Florin Cuconasu, Giovanni Trappolini, Federico Siciliano, Simone Filice, Cesare Campagnano, Yoelle Maarek, Nicola Tonellotto, and Fabrizio Silvestri. 2024. The power of noise: Redefining retrieval for rag systems. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 719–729.
  8. 8.Sunhao Dai, Chen Xu, Shicheng Xu, Liang Pang, Zhenhua Dong, and Jun Xu. 2024. Unifying bias and unfairness in information retrieval: A survey of challenges and opportunities with large language models. arXiv preprint arXiv:2404.11457.
  9. 9.Feiteng Fang, Yuelin Bai, Shiwen Ni, Min Yang, Xiaojun Chen, and Ruifeng Xu. 2024. Enhancing noise robustness of retrieval-augmented language models with adaptive adversarial training. arXiv preprint arXiv:2405.20978.
  10. 10.Michael Glass, Gaetano Rossiello, Md Faisal Mahbub Chowdhury, Ankita Naik, Pengshan Cai, and Alfio Gliozzo. 2022. Re2g: Retrieve, rerank, generate. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2701–2715.
  11. 11.Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang. 2020. Retrieval augmented language model pre-training. In International conference on machine learning, pages 3929–3938. PMLR.
  12. 12.Soyeong Jeong, Jinheon Baek, Sukmin Cho, Sung Ju Hwang, and Jong C Park. 2024. Adaptive-rag: Learning to adapt retrieval-augmented large language models through question complexity. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 7029–7043.
  13. 13.Zhengbao Jiang, Frank F Xu, Luyu Gao, Zhiqing Sun, Qian Liu, Jane Dwivedi-Yu, Yiming Yang, Jamie Callan, and Graham Neubig. 2023. Active retrieval augmented generation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 7969–7992.
  14. 14.Zhuoran Jin, Pengfei Cao, Yubo Chen, Kang Liu, Xiaojian Jiang, Jiexin Xu, Li Qiuxia, and Jun Zhao. 2024. Tug-of-war between knowledge: Exploring and resolving knowledge conflicts in retrieval-augmented language models. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), pages 16867–16878.
  15. 15.Mandar Joshi, Eunsol Choi, Daniel S Weld, and Luke Zettlemoyer. 2017. Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1601–1611.
  16. 16.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al. 2019. Natural questions: a benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:453–466.
  17. 17.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in Neural Information Processing Systems, 33:9459–9474.
  18. 18.Shayne Longpre, Kartik Perisetla, Anthony Chen, Nikhil Ramesh, Chris DuBois, and Sameer Singh. 2021. Entity-based knowledge conflicts in question answering. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing.
  19. 19.Xinbei Ma, Yeyun Gong, Pengcheng He, Hai Zhao, and Nan Duan. 2023. Query rewriting in retrieval-augmented large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 5303–5315.
  20. 20.Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Daniel Khashabi, and Hannaneh Hajishirzi. 2023. When not to trust language models: Investigating effectiveness of parametric and non-parametric memories. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 9802–9822.
  21. 21.Ruotong Pan, Boxi Cao, Hongyu Lin, Xianpei Han, Jia Zheng, Sirui Wang, Xunliang Cai, and Le Sun. 2024. Not all contexts are equal: Teaching llms credibility-aware generation. arXiv preprint arXiv:2404.06809.
  22. 22.Quang Hieu Pham, Hoang Ngo, Anh Tuan Luu, and Dat Quoc Nguyen. 2024. Who’s who: Large language models meet knowledge conflicts in practice. arXiv preprint arXiv:2410.15737.
  23. 23.Rulin Shao, Jacqueline He, Akari Asai, Weijia Shi, Tim Dettmers, Sewon Min, Luke Zettlemoyer, and Pang Wei Koh. 2024. Scaling retrieval-based language models with a trillion-token datastore. arXiv preprint arXiv:2407.12854.
  24. 24.Ivan Stelmakh, Yi Luan, Bhuwan Dhingra, and Ming-Wei Chang. 2022. Asqa: Factoid questions meet long-form answers. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 8273–8288.
  25. 25.Hongjin Su, Howard Yen, Mengzhou Xia, Weijia Shi, Niklas Muennighoff, Han-yu Wang, Haisu Liu, Quan Shi, Zachary S Siegel, Michael Tang, et al. 2024. Bright: A realistic and challenging benchmark for reasoning-intensive retrieval. arXiv preprint arXiv:2407.12883.
  26. 26.Hexiang Tan, Fei Sun, Wanli Yang, Yuanzhuo Wang, Qi Cao, and Xueqi Cheng. 2024. Blinded by generated contexts: How language models merge generated and retrieved contexts for open-domain qa? arXiv preprint arXiv:2401.11911.
  27. 27.Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, and Iryna Gurevych. 2024. Beir: A heterogeneous benchmark for zero-shot evaluation of information retrieval models. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2).
  28. 28.George Tsatsaronis, Georgios Balikas, Prodromos Malakasiotis, Ioannis Partalas, Matthias Zschunke, Michael R Alvers, Dirk Weissenborn, Anastasia Krithara, Sergios Petridis, Dimitris Polychronopoulos, et al. 2015. An overview of the bioasq large-scale biomedical semantic indexing and question answering competition. BMC bioinformatics, 16:1–28.
  29. 29.Haotian Wang, Xiyuan Du, Weijiang Yu, Qianglong Chen, Kun Zhu, Zheng Chu, Lian Yan, and Yi Guan. 2023a. Apollo’s oracle: Retrieval-augmented reasoning in multi-agent debates. arXiv preprint arXiv:2312.04854.
  30. 30.Zhiruo Wang, Jun Araki, Zhengbao Jiang, Md Rizwan Parvez, and Graham Neubig. 2023b. Learning to filter context for retrieval-augmented generation. arXiv preprint arXiv:2311.08377.
  31. 31.Zilong Wang, Zifeng Wang, Long Le, Huaixiu Steven Zheng, Swaroop Mishra, Vincent Perot, Yuwei Zhang, Anush Mattapalli, Ankur Taly, Jingbo Shang, et al. 2024. Speculative rag: Enhancing retrieval augmented generation through drafting. arXiv preprint arXiv:2407.08223.
  32. 32.Zhepei Wei, Wei-Lin Chen, and Yu Meng. 2024. Instructrag: Instructing retrieval-augmented generation with explicit denoising. arXiv preprint arXiv:2406.13629.
  33. 33.Chong Xiang, Tong Wu, Zexuan Zhong, David Wagner, Danqi Chen, and Prateek Mittal. 2024. Certifiably robust rag against retrieval corruption. arXiv preprint arXiv:2405.15556.
  34. 34.Jian Xie, Kai Zhang, Jiangjie Chen, Renze Lou, and Yu Su. 2024. Adaptive chameleon or stubborn sloth: Revealing the behavior of large language models in knowledge conflicts. In The Twelfth International Conference on Learning Representations.
  35. 35.Fangyuan Xu, Weijia Shi, and Eunsol Choi. 2023. Re-comp: Improving retrieval-augmented lms with compression and selective augmentation. arXiv preprint arXiv:2310.04408.
  36. 36.Peng Xu, Wei Ping, Xianchao Wu, Lawrence McAfee, Chen Zhu, Zihan Liu, Sandeep Subramanian, Evelina Bakhturina, Mohammad Shoeybi, and Bryan Catanzaro. 2024. Retrieval meets long context large language models. In The Twelfth International Conference on Learning Representations.
  37. 37.Shi-Qi Yan, Jia-Chen Gu, Yun Zhu, and Zhen-Hua Ling. 2024. Corrective retrieval augmented generation. arXiv preprint arXiv:2401.15884.
  38. 38.Xiao Yang, Kai Sun, Hao Xin, Yushi Sun, Nikita Bhalla, Xiangsen Chen, Sajal Choudhary, Rongze Daniel Gui, Ziran Will Jiang, Ziyu Jiang, et al. 2024. Crag–comprehensive rag benchmark. arXiv preprint arXiv:2406.04744.
  39. 39.Chanwoong Yoon, Taewhoo Lee, Hyeon Hwang, Minbyul Jeong, and Jaewoo Kang. 2024. Compact: Compressing retrieved documents actively for question answering. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 21424–21439.
  40. 40.Ori Yoran, Tomer Wolfson, Ori Ram, and Jonathan Berant. 2024. Making retrieval-augmented language models robust to irrelevant context. In The Twelfth International Conference on Learning Representations.
  41. 41.Wenhao Yu, Dan Iter, Shuohang Wang, Yichong Xu, Mingxuan Ju, Soumya Sanyal, Chenguang Zhu, Michael Zeng, and Meng Jiang. 2023a. Generate rather than retrieve: Large language models are strong context generators. In The Eleventh International Conference on Learning Representations.
  42. 42.Wenhao Yu, Hongming Zhang, Xiaoman Pan, Kaixin Ma, Hongwei Wang, and Dong Yu. 2023b. Chain-of-note: Enhancing robustness in retrieval-augmented language models. arXiv preprint arXiv:2311.09210.
  43. 43.Yue Yu, Wei Ping, Zihan Liu, Boxin Wang, Jiaxuan You, Chao Zhang, Mohammad Shoeybi, and Bryan Catanzaro. 2024. Rankrag: Unifying context ranking with retrieval-augmented generation in llms. arXiv preprint arXiv:2407.02485.
  44. 44.Yunxiang Zhang, Muhammad Khalifa, Lajanugen Logeswaran, Moontae Lee, Honglak Lee, and Lu Wang. 2023. Merging generated and retrieved knowledge for open-domain qa. In The 2023 Conference on Empirical Methods in Natural Language Processing.
  45. 45.Zheng Zhao, Emilio Monti, Jens Lehmann, and Haytham Assem. 2024. Enhancing contextual understanding in large language models through contrastive decoding. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 4225–4237.
  46. 46.Wei Zou, Runpeng Geng, Binghui Wang, and Jinyuan Jia. 2024. Poisonedrag: Knowledge poisoning attacks to retrieval-augmented generation of large language models. arXiv preprint arXiv:2402.07867.

Citation

MLA
Wang, F., et al. “Astute RAG: Overcoming Imperfect Retrieval Augmentation and Knowledge Conflicts for Large Language Models”. arXiv, 2024, http://arxiv.org/abs/2410.07176v2.
APA
Wang, F., Wan, X., Sun, R., Chen, J., & Arık, S. Ö. (2024). Astute RAG: Overcoming Imperfect Retrieval Augmentation and Knowledge Conflicts for Large Language Models. arXiv. http://arxiv.org/abs/2410.07176v2
Chicago
Wang, F., X. Wan, R. Sun, J. Chen, and S. Ö. Arık. 2024. “Astute RAG: Overcoming Imperfect Retrieval Augmentation and Knowledge Conflicts for Large Language Models”. arXiv. http://arxiv.org/abs/2410.07176v2.
Harvard
Wang, F. et al. (2024) “Astute RAG: Overcoming Imperfect Retrieval Augmentation and Knowledge Conflicts for Large Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2410.07176v2.
Vancouver
1. Wang F, Wan X, Sun R, Chen J, Arık SÖ (2024) Astute RAG: Overcoming Imperfect Retrieval Augmentation and Knowledge Conflicts for Large Language Models. arXiv

BibTeX

@article{wang2024astute,
  title = {Astute RAG: Overcoming Imperfect Retrieval Augmentation and Knowledge Conflicts for Large Language Models},
  author = {Wang, Fei and Wan, Xingchen and Sun, Ruoxi and Chen, Jiefeng and Arık, Sercan Ö.},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2410.07176v2},
  eprint = {2410.07176}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/