Enhancing Noise Robustness of Retrieval-Augmented Language Models with Adaptive Adversarial Training

Feiteng FangYuelin BaiShiwen NiMin YangXiaojun ChenRuifeng Xu

article2024ACL92 citations

Presents an adaptive adversarial training framework that dynamically adjusts model training and applies multi-task learning to prevent large language models from being misled by irrelevant, superficial, or counterfactual retrieved contexts.

Listen

Large language models often struggle with factual inaccuracies and outdated knowledge. To address these issues, organizations increasingly use retrieval-augmented generation to pull relevant facts from external databases before generating answers. However, real-world search mechanisms frequently retrieve imperfect or noisy information. When these models encounter inaccurate or irrelevant context, their accuracy drops significantly, creating serious operational risks for automated decision-making and knowledge systems.

The article evaluates how different categories of retrieval noise impair model performance and demonstrates a new training framework designed to strengthen model resilience against these errors.

The researchers established a benchmark called RAG-Bench using 3,000 test queries drawn from three standard question-answering datasets (Natural Questions, TriviaQA, and WebQ). They categorized retrieval noise into three real-world types: irrelevant text, superficially relevant text lacking the right answer, and counterfactual text containing incorrect facts. They then developed Retrieval-augmented Adaptive Adversarial Training (RAAT), an approach that dynamically identifies the noise types causing the highest generation errors during fine-tuning and prioritizes those examples for model updates. This method also integrates multi-task learning to teach the model to classify noise types internally.

The evaluation revealed several key findings. First, existing leading models experience notable accuracy declines ranging from about 0.2% to 13.4% when exposed to noisy context. Second, counterfactual and superficially relevant noise cause far more degradation than completely irrelevant noise. Third, applying the RAAT method to an open-source 7-billion parameter language model achieved an average exact match score of 82.2% and an F1 score of 86.4% across all noise settings, outperforming standard fine-tuning baselines by roughly 2.1 to 2.5 percentage points. Finally, ablation testing confirmed that both the adaptive regularization and the auxiliary noise-classification tasks are essential for maintaining stable performance across diverse noise conditions.

These findings indicate that retrieval-augmented systems can achieve higher reliability without relying solely on upstream search improvements. By enabling language models to withstand misleading and conflicting evidence, organizations can reduce compliance and operational risks associated with automated outputs. The results challenge the traditional assumption that simple offline data augmentation is sufficient, showing that dynamic, noise-aware adversarial adaptation is necessary for robust performance.

Organizations deploying retrieval-augmented systems should adopt noise-aware training strategies that specifically target counterfactual and partially relevant distractors. For future research and implementation, development teams should validate this method across broader natural language tasks beyond question answering and explore joint training frameworks that optimize both the search retriever and the language generator simultaneously.

The findings are bounded by the study's focus on open-domain question-answering benchmarks and single-model fine-tuning. While confidence in the reported performance gains on standard benchmarks is high, readers should exercise caution before generalizing these results to complex multi-step reasoning tasks or domain-specific enterprise databases without further pilot testing.

Cover for Enhancing Noise Robustness of Retrieval-Augmented Language Models with Adaptive Adversarial Training

Abstract

Large Language Models (LLMs) exhibit substantial capabilities yet encounter challenges, including hallucination, outdated knowledge, and untraceable reasoning processes. Retrieval-augmented generation (RAG) has emerged as a promising solution, integrating knowledge from external databases to mitigate these challenges. However, inappropriate retrieved passages can potentially hinder the LLMs’ capacity to generate comprehensive and high-quality responses. Prior RAG studies on the robustness of retrieval noises often confine themselves to a limited set of noise types, deviating from real-world retrieval environments and limiting practical applicability. In this study, we initially investigate retrieval noises and categorize them into three distinct types, reflecting real-world environments. We analyze the impact of these various retrieval noises on the robustness of LLMs. Subsequently, we propose a novel RAG approach known as Retrieval-augmented Adaptive Adversarial Training (RAAT). RAAT leverages adaptive adversarial training to dynamically adjust the model’s training process in response to retrieval noises. Concurrently, it employs multi-task learning to ensure the model’s capacity to internally recognize noisy contexts. Extensive experiments demonstrate that the LLaMA-2 7B model trained using RAAT exhibits significant improvements in F1 and EM scores under diverse noise conditions. For reproducibility, we release our code and data at: https://github.com/calubkk/RAAT.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Methodology
  • 3.1 Problem Setup
  • 3.2 Diverse Retrieval Noises
  • 3.3 Retrieval-augmented Adaptive Adversarial Training
  • 3.4 Incorporating Noise Awareness
  • 4 Experiments
  • 4.1 Dataset Construction
  • 4.2 Evaluation Metrics
  • 4.3 Baseline Methods
  • 4.4 Implementation Details
  • 5 Experimental Results
  • 5.1 Main Results
  • 5.2 Ablation Study
  • 5.3 Further Discussion
  • 6 Conclusion
  • 7 Limitations
  • Acknowledgments
  • References
  • A Has the Model Truly Attained Noise Awareness?

Knowls

  1. Knowl 1 — Retrieval-augmented Adaptive Adversarial Training (RAAT) Framework

    model/method

    Retrieval-augmented Adaptive Adversarial Training (RAAT) is a fine-tuning method designed to fortify Retrieval-Augmented Language Models (RALMs) against retrieval noises. Given an input query xx and a target answer yy, RAAT defines a discrete data augmentation space consisting of four context configurations: DA={dag,dar,dai,dac}DA = \{d_{ag}, d_{ar}, d_{ai}, d_{ac}\} where dagd_{ag} denotes input formatted with the golden retrieval context only, dard_{ar} includes additional relevant retrieval noise, daid_{ai} includes additional irrelevant retrieval noise, and dacd_{ac} includes additional counterfactual retrieval noise.

    For an augmented input x′=da(x)x' = da(x), the autoregressive token generation loss is computed as: L′(θ,x′,y)=−1∣y∣∑t=1∣y∣log⁡Pθ(yt∣x′,y<t)\mathcal{L}'(\theta, x', y) = -\frac{1}{|y|} \sum_{t=1}^{|y|} \log P_\theta(y_t \mid x', y_{<t}) where θ\theta represents the parameters of the language model.

    In each iteration, the model evaluates L′\mathcal{L}' across all four candidate augmentations and identifies the maximum loss Lmax⁡′=max⁡da∈DAL′(θ,da(x),y)\mathcal{L}'_{\max} = \max_{da \in DA} \mathcal{L}'(\theta, da(x), y) and the minimum loss Lmin⁡′=min⁡da∈DAL′(θ,da(x),y)\mathcal{L}'_{\min} = \min_{da \in DA} \mathcal{L}'(\theta, da(x), y). To prevent overfitting to any single noise profile and stabilize optimization, RAAT applies a variance-regularization term: Lreg=∥Lmax⁡′−Lmin⁡′∥22\mathcal{L}_{\text{reg}} = \|\mathcal{L}'_{\max} - \mathcal{L}'_{\min}\|_2^2

    The adaptive adversarial generation loss is formulated as: Lada=Lmax⁡′+wreg⋅Lreg\mathcal{L}_{\text{ada}} = \mathcal{L}'_{\max} + w_{\text{reg}} \cdot \mathcal{L}_{\text{reg}} where wregw_{\text{reg}} is a weighting hyperparameter (set to 0.10.1).

    Concurrently, to encourage the model to internally recognize context quality, RAAT introduces an auxiliary classification task via a linear layer attached to the LLM representation to classify each of the four contexts into its respective type (labels 11 through 44) using a cross-entropy classification loss Lcls\mathcal{L}_{\text{cls}}. The final joint multi-task objective is: LRAAT=wada⋅Lada+wcls⋅Lcls\mathcal{L}_{\text{RAAT}} = w_{\text{ada}} \cdot \mathcal{L}_{\text{ada}} + w_{\text{cls}} \cdot \mathcal{L}_{\text{cls}} where wada=2w_{\text{ada}} = 2 and wcls=1w_{\text{cls}} = 1.

  2. Knowl 2 — Taxonomy of Retrieval Noise in Retrieval-Augmented Generation

    definition

    In retrieval-augmented generation for question answering, retrieval noise present in non-golden retrieved contexts cnoisyc_{\text{noisy}} is classified into three distinct categories:

    1. Relevant retrieval noise (crc_r): Contexts that are superficially relevant or topically aligned with the query xx, but do not contain the factual information required to produce the correct answer yy. These contexts tend to mislead models by appearing plausible while lacking the target answer.
    2. Irrelevant retrieval noise (cic_i): Contexts that have low semantic relevance to the query xx, originating from erroneous retrievals and consisting of completely off-topic material.
    3. Counterfactual retrieval noise (ccc_c): Contexts that are topically related to query xx and closely resemble the true context, but contain factually incorrect or corrupted information (such as target entities substituted with false alternatives).
  3. Knowl 3 — RAG-Bench Benchmark for Retrieval Noise Robustness

    experimental setup

    RAG-Bench is an evaluation benchmark constructed to assess the noise robustness of retrieval-augmented language models across three open-domain question answering datasets: Natural Questions (NQ), TriviaQA, and WebQ.

    For each query, Dense Passage Retrieval (DPR) retrieves the top-10 Wikipedia passages. Queries are filtered to retain only those containing at least two golden retrieval contexts (passages containing the ground-truth answer). The three noise types are constructed as follows:

    • Relevant noise (crc_r): The top-ranked retrieved passage from the DPR top-10 that is not a golden context.
    • Irrelevant noise (cic_i): A passage randomly selected from the retrieved results of a different, unrelated query.
    • Counterfactual noise (ccc_c): Generated by taking one of the query's golden retrieval passages and substituting its answer entity with an incorrect entity.

    The benchmark dataset split contains:

    • Training set: 4,500 samples (1,500 sampled uniformly from each of NQ, TriviaQA, and WebQ).
    • Validation set: 300 non-overlapping samples (100 per dataset).
    • Test set: 3,000 samples (1,000 sampled uniformly from each test set of NQ, TriviaQA, and WebQ).
  4. Knowl 4 — Comparative Evaluation of Retrieval-Augmented Models on RAG-Bench

    data/table

    The table below details Exact Match (EM) and F1 scores on the 3,000-sample RAG-Bench test set under four test settings: Golden Only (only the correct passage is provided in context), Golden & $c_i$ (golden passage plus irrelevant noise), Golden & $c_r$ (golden passage plus relevant noise), Golden & $c_c$ (golden passage plus counterfactual noise), and their average (Avg). Models evaluated include zero-shot LLMs (LLaMA-2 7B/13B/70B, Qwen 7B/14B, ChatGPT-3.5) and fine-tuned LLaMA-2 7B baselines (RALMgolden, RetRobust, RALMretrieved, RALMmultiple, and RAAT).

    Method Golden Only Golden cic_i Golden crc_r Golden ccc_c Avg
    F1 EM F1 EM F1 EM F1 EM F1 EM
    LLaMA27B_{\text{7B}} 65.56 51.80 56.14 42.87 53.10 39.73 51.81 38.37 56.68 43.19
    Qwen7B_{\text{7B}} 62.57 47.07 61.48 46.06 55.50 40.50 53.26 36.90 58.20 42.63
    LLaMA213B_{\text{13B}} 69.27 55.00 63.25 49.47 62.27 47.97 62.07 47.17 64.22 49.90
    Qwen14B_{\text{14B}} 67.45 51.43 66.71 51.20 61.88 46.16 58.65 41.30 63.67 47.52
    LLaMA270B_{\text{70B}} 71.43 56.56 70.05 55.13 65.97 51.33 63.91 48.27 67.84 52.82
    ChatGPT3.5_{3.5} 73.98 60.50 72.24 60.30 70.65 56.89 69.00 54.64 71.47 58.10
    RALMgolden_{\text{golden}} 80.31 74.03 79.33 72.73 73.26 66.33 73.08 65.40 76.50 69.62
    RetRobust 80.10 73.80 79.25 72.97 74.81 68.30 75.46 68.43 77.41 70.88
    RALMretrieved_{\text{retrieved}} 80.04 73.40 81.09 74.80 75.99 69.10 73.10 65.67 77.55 70.74
    RALMmultiple_{\text{multiple}} 85.47 80.17 85.27 81.20 83.07 78.33 83.25 79.23 84.27 79.73
    RAAT 87.15 83.07 86.80 82.73 85.14 81.00 86.29 82.10 86.35 82.23

    RAAT based on LLaMA-2 7B outperforms all baseline models across all noise categories, exceeding the strongest fine-tuning baseline (RALMmultiple) by an average of 2.08% in F1 score and 2.50% in EM score.

  5. Knowl 5 — Vulnerability Profile of Language Models to Retrieval Noise Types

    empirical result

    Empirical evaluation of zero-shot large language models (LLaMA-2 7B/13B/70B, Qwen 7B/14B, and ChatGPT-3.5) on retrieval-augmented question answering reveals distinct sensitivity patterns across retrieval noise types:

    1. Counterfactual noise (ccc_c) causes the largest performance drops across all models, leading to reductions of up to 13.75% in F1 (e.g., LLaMA-2 7B drops from 65.56% to 51.81% F1, and 51.80% to 38.37% EM).
    2. Relevant noise (crc_r) causes the second largest degradation (e.g., a 12.46% F1 reduction in LLaMA-2 7B), demonstrating that superficial topical relevance without answer content actively misleads LLMs.
    3. Irrelevant noise (cic_i) has a comparatively minor impact on capable LLMs (e.g., ChatGPT-3.5 drops by only 1.74% F1 and 0.20% EM when irrelevant noise is added).
    4. Model scaling reduces vulnerability: across identical model families, larger models exhibit smaller relative drops when exposed to retrieval noise (for example, relevant noise causes a 12.46% F1 reduction in LLaMA-2 7B versus only a 7.00% reduction in LLaMA-2 13B).
  6. Knowl 6 — Ablation Analysis of RAAT Components

    data/table

    The ablation study assesses the individual contributions of the noise-aware classification loss Lcls\mathcal{L}_{\text{cls}} and the variance-regularization term Lreg\mathcal{L}_{\text{reg}} in the RAAT objective function when fine-tuning LLaMA-2 7B on RAG-Bench:

    Method Golden Only Golden cic_i Golden crc_r Golden ccc_c Avg
    F1 EM F1 EM F1 EM F1 EM F1 EM
    RAAT 87.15 83.07 86.80 82.73 85.14 81.00 86.29 82.10 86.35 82.23
    RAAT w/o Lcls\mathcal{L}_{\text{cls}} 86.76 82.77 86.45 82.27 84.69 80.63 85.54 81.20 85.86 81.71
    RAAT w/o Lreg\mathcal{L}_{\text{reg}} 86.87 83.03 83.92 79.86 84.69 80.57 87.02 82.80 85.63 81.57

    Removing Lcls\mathcal{L}_{\text{cls}} results in a consistent overall drop of 0.49% F1 and 0.52% EM. Removing Lreg\mathcal{L}_{\text{reg}} causes a severe performance drop specifically in the presence of irrelevant retrieval noise (EM drops by 2.87% on Golden & $c_i$), confirming that regularizing loss variance prevents the model from overfitting to higher-loss noise types at the expense of lower-loss noise types.

  7. Knowl 7 — Adversarial Noise Selection Distribution During RAAT Training

    empirical result

    Tracking the sample selection in RAAT across 9,000 parameter updates (covering 4,500 training queries) shows that the adaptive min-max optimization dynamically prioritizes more difficult retrieval noises based on generation loss magnitude:

    • Counterfactual retrieval noise (ccc_c): Selected in 3,165 updates (35.17%)
    • Relevant retrieval noise (crc_r): Selected in 2,706 updates (30.07%)
    • Golden context (cagc_{ag}): Selected in 1,758 updates (19.53%)
    • Irrelevant retrieval noise (cic_i): Selected in 1,371 updates (15.23%)

    This distribution aligns directly with the measured difficulty of noise types, demonstrating that RAAT automatically focuses model updates on counterfactual and relevant retrieval noise while allocating fewer updates to irrelevant noise and golden contexts.

  8. Knowl 8 — Internal Representation Clustering Under RAAT

    empirical result

    t-SNE dimensionality reduction performed on the hidden state vectors of the last token across golden, relevant, irrelevant, and counterfactual contexts shows that:

    • Baseline models fine-tuned with standard instruction tuning (RALMgolden) or noisy training (RetRobust) produce overlapping and entangled representations across different noise types, lacking intrinsic noise discernment.
    • Models fine-tuned with RAAT exhibit distinct and well-separated representation clusters. Specifically, counterfactual retrieval noise samples form a distinct cluster separated from golden, relevant, and irrelevant context representations, confirming that the auxiliary classification objective Lcls\mathcal{L}_{\text{cls}} successfully induces internal noise awareness.
  9. Knowl 9 — Limitations of RAAT and Current Evaluation Benchmark

    limitation

    The methodology and benchmark have two key limitations:

    1. Benchmark Scope: The evaluation benchmark is constructed exclusively from three open-domain question answering datasets using Wikipedia as the knowledge source, which does not cover other diverse NLP tasks or broader enterprise knowledge base domains.
    2. Decoupled Generation Optimization: RAAT operates strictly at the reader/generator LLM stage to provide robustness against imperfect inputs, and does not perform joint or end-to-end training of the retriever and the language model.

Coverage note — None was omitted; all core contributions, formulations, empirical results, ablations, analyses, and limitations have been captured.

References

  1. 1.Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. 2023. Qwen technical report. arXiv preprint arXiv:2309.16609.
  2. 2.Tao Bai, Jinqi Luo, Jun Zhao, Bihan Wen, and Qian Wang. 2021. Recent advances in adversarial training for adversarial robustness. arXiv preprint arXiv:2102.01356.
  3. 3.Jonathan Berant, Andrew Chou, Roy Frostig, and Percy Liang. 2013. Semantic parsing on freebase from question-answer pairs. In Proceedings of the 2013 conference on empirical methods in natural language processing, pages 1533–1544.
  4. 4.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  5. 5.Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al. 2023. Sparks of artificial general intelligence: Early experiments with gpt-4. arXiv preprint arXiv:2303.12712.
  6. 6.Danqi Chen, Adam Fisch, Jason Weston, and Antoine Bordes. 2017. Reading wikipedia to answer opendomain questions. arXiv preprint arXiv:1704.00051.
  7. 7.Jiawei Chen, Hongyu Lin, Xianpei Han, and Le Sun. 2023. Benchmarking large language models in retrieval-augmented generation. arXiv preprint arXiv:2309.01431.
  8. 8.Antonia Creswell, Murray Shanahan, and Irina Higgins. 2022. Selection-inference: Exploiting large language models for interpretable logical reasoning. arXiv preprint arXiv:2205.09712.
  9. 9.Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, and Haofen Wang. 2023. Retrieval-augmented generation for large language models: A survey. arXiv preprint arXiv:2312.10997.
  10. 10.Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy. 2014. Explaining and harnessing adversarial examples. arXiv preprint arXiv:1412.6572.
  11. 11.Almog Gueta, Elad Venezian, Colin Raffel, Noam Slonim, Yoav Katz, and Leshem Choshen. 2023. Knowledge is a region in weight space for fine-tuned language models. arXiv preprint arXiv:2302.04863.
  12. 12.Maor Ivgi and Jonathan Berant. 2021. Achieving model robustness through discrete adversarial training. arXiv preprint arXiv:2104.05062.
  13. 13.Neel Jain, Ping-yeh Chiang, Yuxin Wen, John Kirchenbauer, Hong-Min Chu, Gowthami Somepalli, Brian R Bartoldson, Bhavya Kailkhura, Avi Schwarzschild, Aniruddha Saha, et al. 2023. Neftune: Noisy embeddings improve instruction finetuning. arXiv preprint arXiv:2310.05914.
  14. 14.Robin Jia and Percy Liang. 2017. Adversarial examples for evaluating reading comprehension systems. arXiv preprint arXiv:1707.07328.
  15. 15.Mandar Joshi, Eunsol Choi, Daniel S Weld, and Luke Zettlemoyer. 2017. Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension. arXiv preprint arXiv:1705.03551.
  16. 16.Nikhil Kandpal, Haikang Deng, Adam Roberts, Eric Wallace, and Colin Raffel. 2023. Large language models struggle to learn long-tail knowledge. In International Conference on Machine Learning, pages 15696–15707. PMLR.
  17. 17.Vladimir Karpukhin, Barlas Oğuz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense passage retrieval for open-domain question answering. arXiv preprint arXiv:2004.04906.
  18. 18.Alexey Kurakin, Ian Goodfellow, and Samy Bengio. 2016. Adversarial machine learning at scale. arXiv preprint arXiv:1611.01236.
  19. 19.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al. 2019. Natural questions: a benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:453–466.
  20. 20.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in Neural Information Processing Systems, 33:9459–9474.
  21. 21.Daliang Li, Ankit Singh Rawat, Manzil Zaheer, Xin Wang, Michal Lukasik, Andreas Veit, Felix Yu, and Sanjiv Kumar. 2022. Large language models with controllable working memory. arXiv preprint arXiv:2211.05110.
  22. 22.Ke Liang, Yue Liu, Sihang Zhou, Wenxuan Tu, Yi Wen, Xihong Yang, Xiangjun Dong, and Xinwang Liu. 2023. Knowledge graph contrastive learning based on relation-symmetrical structure. IEEE Transactions on Knowledge and Data Engineering.
  23. 23.Xi Victoria Lin, Xilun Chen, Mingda Chen, Weijia Shi, Maria Lomeli, Rich James, Pedro Rodriguez, Jacob Kahn, Gergely Szilvasy, Mike Lewis, et al. 2023. Ra-dit: Retrieval-augmented dual instruction tuning. arXiv preprint arXiv:2310.01352.
  24. 24.Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. 2017. Towards deep learning models resistant to adversarial attacks. arXiv preprint arXiv:1706.06083.
  25. 25.Yuning Mao, Pengcheng He, Xiaodong Liu, Yelong Shen, Jianfeng Gao, Jiawei Han, and Weizhu Chen. 2020. Generation-augmented retrieval for open-domain question answering. arXiv preprint arXiv:2009.08553.
  26. 26.Takeru Miyato, Andrew M Dai, and Ian Goodfellow. 2016. Adversarial training methods for semi-supervised text classification. arXiv preprint arXiv:1605.07725.
  27. 27.John X Morris, Eli Lifland, Jin Yong Yoo, Jake Grigsby, Di Jin, and Yanjun Qi. 2020. Textattack: A framework for adversarial attacks, data augmentation, and adversarial training in nlp. arXiv preprint arXiv:2005.05909.
  28. 28.Shiwen Ni, Jiawen Li, Min Yang, and Hung-Yu Kao. 2023. Dropattack: A random dropped weight attack adversarial training for natural language understanding. IEEE/ACM Transactions on Audio, Speech, and Language Processing.
  29. 29.Sebastian Ruder. 2017. An overview of multi-task learning in deep neural networks. arXiv preprint arXiv:1706.05098.
  30. 30.Xinyue Shen, Zeyuan Chen, Michael Backes, and Yang Zhang. 2023. In chatgpt we trust? measuring and characterizing the reliability of chatgpt. arXiv preprint arXiv:2304.08979.
  31. 31.Freda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales, David Dohan, Ed H Chi, Nathanael Schärli, and Denny Zhou. 2023. Large language models can be easily distracted by irrelevant context. In International Conference on Machine Learning, pages 31210–31227. PMLR.
  32. 32.Nandan Thakur, Luiz Bonifacio, Xinyu Zhang, Odunayo Ogundepo, Ehsan Kamalloo, David Alfonso-Hermelo, Xiaoguang Li, Qun Liu, Boxing Chen, Mehdi Rezagholizadeh, et al. 2023. Nomiracl: Knowing when you don’t know for robust multilingual retrieval-augmented generation. arXiv preprint arXiv:2312.11361.
  33. 33.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  34. 34.Boxin Wang, Shuohang Wang, Yu Cheng, Zhe Gan, Ruoxi Jia, Bo Li, and Jingjing Liu. 2020. Infobert: Improving robustness of language models from an information theoretic perspective. arXiv preprint arXiv:2010.02329.
  35. 35.Yi Wu, David Bamman, and Stuart Russell. 2017. Adversarial training for relation extraction. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 1778–1783.
  36. 36.Michihiro Yasunaga, Jungo Kasai, and Dragomir Radev. 2017. Robust multilingual part-of-speech tagging via adversarial training. arXiv preprint arXiv:1711.04903.
  37. 37.Xunjian Yin, Baizhou Huang, and Xiaojun Wan. 2023. Alcuna: large language models meet new knowledge. arXiv preprint arXiv:2310.14820.
  38. 38.Ori Yoran, Tomer Wolfson, Ori Ram, and Jonathan Berant. 2023. Making retrieval-augmented language models robust to irrelevant context. arXiv preprint arXiv:2310.01558.
  39. 39.Wenhao Yu, Hongming Zhang, Xiaoman Pan, Kaixin Ma, Hongwei Wang, and Dong Yu. 2023. Chain-of-note: Enhancing robustness in retrieval-augmented language models. arXiv preprint arXiv:2311.09210.
  40. 40.Chen Zhu, Yu Cheng, Zhe Gan, Siqi Sun, Tom Goldstein, and Jingjing Liu. 2019. Freelb: Enhanced adversarial training for natural language understanding. arXiv preprint arXiv:1909.11764.
  41. 41.Yutao Zhu, Huaying Yuan, Shuting Wang, Jiongnan Liu, Wenhan Liu, Chenlong Deng, Zhicheng Dou, and Ji-Rong Wen. 2023. Large language models for information retrieval: A survey. arXiv preprint arXiv:2308.07107.
  42. 42.Andy Zou, Zifan Wang, J Zico Kolter, and Matt Fredrikson. 2023. Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043.

Citation

MLA
Fang, F., et al. “Enhancing Noise Robustness of Retrieval-Augmented Language Models with Adaptive Adversarial Training”. ACL 2024, Main Conference, 2024, http://arxiv.org/abs/2405.20978v1.
APA
Fang, F., Bai, Y., Ni, S., Yang, M., Chen, X., & Xu, R. (2024). Enhancing Noise Robustness of Retrieval-Augmented Language Models with Adaptive Adversarial Training. ACL 2024, Main Conference. http://arxiv.org/abs/2405.20978v1
Chicago
Fang, F., Y. Bai, S. Ni, M. Yang, X. Chen, and R. Xu. 2024. “Enhancing Noise Robustness of Retrieval-Augmented Language Models with Adaptive Adversarial Training”. ACL 2024, Main Conference. http://arxiv.org/abs/2405.20978v1.
Harvard
Fang, F. et al. (2024) “Enhancing Noise Robustness of Retrieval-Augmented Language Models with Adaptive Adversarial Training”, ACL 2024, Main Conference [Preprint]. Available at: http://arxiv.org/abs/2405.20978v1.
Vancouver
1. Fang F, Bai Y, Ni S, Yang M, Chen X, Xu R (2024) Enhancing Noise Robustness of Retrieval-Augmented Language Models with Adaptive Adversarial Training. ACL 2024, Main Conference

BibTeX

@article{fang2024enhancing,
  title = {Enhancing Noise Robustness of Retrieval-Augmented Language Models with Adaptive Adversarial Training},
  author = {Fang, Feiteng and Bai, Yuelin and Ni, Shiwen and Yang, Min and Chen, Xiaojun and Xu, Ruifeng},
  year = {2024},
  journal = {ACL 2024, Main Conference},
  url = {http://arxiv.org/abs/2405.20978v1},
  eprint = {2405.20978}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/