Chain-of-Note: Enhancing Robustness in Retrieval-Augmented Language Models

Wenhao YuHongming ZhangXiaoman PanPeixin CaoKaixin MaJian LiHongwei WangDong Yu

article2024EMNLP195 citations

Proposes a sequential note-taking framework for retrieval-augmented language models that systematically evaluates the relevance of retrieved documents, allowing models to filter noise, rely on internal knowledge when appropriate, and admit when an answer is unknown.

Listen

Retrieval-augmented language models combine external document search with generative artificial intelligence to provide up-to-date answers and reduce factual inaccuracies. In real-world applications, however, automated retrieval frequently pulls noisy or irrelevant documents that distract the model, cause it to override its own accurate internal knowledge, or trigger misleading responses. Furthermore, standard retrieval systems struggle to recognize when information is entirely missing, often producing confident hallucinations instead of admitting knowledge gaps.

The article evaluates whether introducing a structured note-taking mechanism can enhance model robustness against irrelevant information and improve its ability to decline unanswerable questions. Specifically, it introduces Chain-of-Note, a framework where the model systematically evaluates each retrieved document by generating sequential reading notes to judge its relevance and credibility before formulating a final answer.

To test this concept, the authors prompted GPT-4 to generate 10,000 training examples with reading notes from search queries, fine-tuning an open-source LLaMA-2 7B model. The approach was evaluated across four open-domain question answering datasets under various noise levels and against real-time questions outside the model's pre-training knowledge base. Experiments were also conducted using GPT-4 directly to compare the approach against standard step-by-step reasoning prompts.

The findings show that generating reading notes consistently improves overall accuracy and substantially bolsters reliability under adverse conditions. First, on datasets containing completely noisy documents, the method improved exact match accuracy by an average of about 7.9 points over standard retrieval systems. Second, on completely new, real-time queries outside the training scope, it increased the rejection rate by over 10.5 points, enabling the model to respond with "unknown" rather than guessing incorrectly. Third, on larger systems like GPT-4, the method outperformed standard chain-of-thought prompting by about 2.0 to 4.0 percentage points across various benchmarks.

These results demonstrate that requiring a language model to explicitly assess source relevance mitigates operational and compliance risks tied to factual errors and hallucinations. Rather than blindly trusting retrieved text, the model learns to filter noise, infer answers using inherent knowledge when external data is incomplete, or decline answering when information is absent. This transparency also provides an interpretable reasoning trail for why specific evidence was accepted or dismissed.

For practical deployment, organizations should adopt this note-taking structure or utilize a "hybrid training" approach to balance accuracy and operational costs. The article found that standard note generation increases inference time from about 0.6 seconds to roughly 12 seconds per query on benchmark hardware. However, a hybrid training strategy—training models on both direct answers and note generation—internalizes the reasoning capability, matching standard speed (about 0.6 seconds) while preserving most robustness gains. Decision-makers should validate this method on their proprietary domain data before large-scale deployment to ensure generated notes remain concise and effective.

arXiv: 2311.09210
Cover for Chain-of-Note: Enhancing Robustness in Retrieval-Augmented Language Models

Abstract

Retrieval-augmented language model (RALM) represents a significant advancement in mitigating factual hallucination by leveraging external knowledge sources. However, the reliability of the retrieved information is not always guaranteed, and the retrieval of irrelevant data can mislead the response generation. Moreover, standard RALMs frequently neglect their intrinsic knowledge due to the interference from retrieved information. In instances where the retrieved information is irrelevant, RALMs should ideally utilize their intrinsic knowledge or, in the absence of both intrinsic and retrieved knowledge, opt to respond with "unknown" to avoid hallucination. In this paper, we introduce CHAIN-OF-NOTE (CON), a novel approach to improve robustness of RALMs in facing noisy, irrelevant documents and in handling unknown scenarios. The core idea of CON is to generate sequential reading notes for each retrieved document, enabling a thorough evaluation of their relevance to the given question and integrating this information to formulate the final answer. Our experimental results show that GPT-4, when equipped with CON, outperforms the CHAIN-OF-THOUGHT approach. Besides, we utilized GPT-4 to create 10K CON data, subsequently trained on LLaMa-2 7B model. Our experiments across four open-domain QA benchmarks show that fine-tuned RALMs equipped with CON significantly outperform standard fine-tuned RALMs.

Table of Contents

  • 1 Introduction
  • 2 Proposed Method
  • 2.1 Overview
  • 2.2 Background of Existing RALMs
  • 2.3 The CHAIN-OF-NOTE Framework
  • 2.3.1 CHAIN-OF-NOTE Format Design
  • 2.3.2 Data Collection and Model Training
  • 2.3.3 Hybrid Training for Better Efficiency
  • 3 Experiments
  • 3.1 Experimental Settings and Evaluations
  • 3.1.1 Datasets and Splits
  • 3.1.2 Baseline Methods
  • 3.1.3 Evaluation Metrics
  • 3.2 Evaluation on Overall QA Performance
  • 3.3 Evaluation on Noise Robustness
  • 3.4 Evaluation on Unknown Robustness
  • 3.5 Evaluation on Hybrid Training Strategy
  • 4 Related Work
  • 5 Conclusion
  • 6 Limitations
  • References
  • A Appendix
  • A.1 More Related Work
  • A.1.1 Retrieval-Augmented Language Models
  • A.2 Chain-of-X Approaches in Large Language Models
  • A.3 Dataset Information
  • A.4 Implementation Details
  • A.5 Instruction Prompts
  • A.6 Case Studies
  • A.7 Licenses

Knowls

  1. Knowl 1 — Chain-of-Note Framework for Robust Retrieval-Augmented Generation

    model/method

    The Chain-of-Note (CoN) framework improves the robustness of retrieval-augmented language models (RALMs) against noisy, irrelevant retrieved documents and out-of-domain/unanswerable queries. Given an input query xx and kk retrieved documents [d1,…,dk][d_1, \dots, d_k], standard RALMs learn a direct mapping f:{x,d1,…,dk}→yf: \{x, d_1, \dots, d_k\} \to y, which frequently suffers from surface-level processing, hallucination under noise, and failure to acknowledge knowledge gaps.

    Instead, CoN introduces intermediate sequential reading notes [yd1,…,ydk][y_{d_1}, \dots, y_{d_k}] before generating the final answer yy, learning the mapping: f:{x,d1,…,dk}→{yd1,…,ydk,y}f: \{x, d_1, \dots, d_k\} \to \{y_{d_1}, \dots, y_{d_k}, y\} where ydiy_{d_i} is a reading note evaluating the relevance, reliability, and factual utility of the ii-th document did_i.

    CoN formats notes into three distinct operational scenarios:

    1. Direct Answering (Type a): When a document directly contains the answer to xx, the reading note extracts and summarizes the critical information, leading directly to the answer.
    2. Contextual Inference (Type b): When retrieved documents do not directly provide the complete answer but provide useful context, the note synthesizes this context with the model's parametric knowledge to deduce the answer.
    3. Unknown Acknowledgment (Type c): When retrieved documents are irrelevant or insufficient and the model's internal knowledge cannot resolve the query, the note identifies the lack of evidence and the model outputs an explicit "unknown" response to prevent hallucination.
  2. Knowl 2 — Hybrid Training for Efficient Implicit Chain-of-Note Reasoning

    model/method

    Generating explicit sequential reading notes increases decoding latency significantly during inference. To eliminate this runtime overhead while preserving the robustness benefits of Chain-of-Note (CoN), a hybrid training strategy is used.

    During fine-tuning, training batches are equally partitioned:

    • 50% Standard RALM format: The model is trained to directly generate the final answer yy given (x,d1,…,dk)(x, d_1, \dots, d_k) without notes, using a standard instruction prompt.
    • 50% Chain-of-Note format: The model is trained to generate explicit sequential reading notes followed by the final answer, {yd1,…,ydk,y}\{y_{d_1}, \dots, y_{d_k}, y\}, using a CoN instruction prompt.

    During inference, the model is queried exclusively using the standard RALM prompt without explicit note generation. This forces the model to perform implicit CoN evaluation within its internal hidden representations, retaining high noise robustness while operating at the low latency of standard direct-generation RALMs.

  3. Knowl 3 — Chain-of-Note Dataset Generation and Training Pipeline

    experimental setup

    To train open-source models with Chain-of-Note (CoN) capabilities without manual annotation, synthetic training data is collected using GPT-4 (gpt-4-1106).

    1. Data Source: 10,000 questions are randomly sampled from the Natural Questions (NQ) training set.
    2. Prompting: GPT-4 is prompted using in-context few-shot exemplars covering the three reading note types (direct answering, contextual inference, and unknown handling) to generate reading notes for Dense Passage Retrieval (DPR) passages.
    3. Fine-Tuning Setup: A LLaMA-2 7B backbone is fine-tuned on the concatenated sequence of instruction, question, documents, generated notes, and answers. Fine-tuning uses the DeepSpeed library with the ZeRO optimizer, running for 3 epochs with a batch size of 128 in bfloat16 precision. A learning rate of 5×10−65 \times 10^{-6} (selected from {1×10−6,2×10−6,5×10−6,1×10−5,2×10−5}\{1\times 10^{-6}, 2\times 10^{-6}, 5\times 10^{-6}, 1\times 10^{-5}, 2\times 10^{-5}\}) is used. Greedy decoding is applied during inference.
  4. Knowl 4 — Overall Open-Domain QA Benchmark Performance

    empirical result

    Chain-of-Note (CoN) improves overall open-domain QA performance over standard retrieval-augmented models across three benchmarks (Natural Questions [NQ], TriviaQA, and Web Questions [WebQ]) when evaluated on full test splits with DPR retrieval.

    Models NQ TriviaQA WebQ Average
    EM F1 EM F1 EM F1 EM F1
    Backbone: LLaMA-2 7B
    QA fine-tune w/o IR 28.80 37.53 63.19 68.61 28.30 42.77 35.98 44.27
    SAIL 36.20 44.23 73.20 80.92 27.92 40.65 45.77 55.27
    Retrieve-Read 47.39 55.81 74.92 81.53 29.58 43.51 48.49 56.97
    + CHAIN-OF-NOTE (ours) 48.92 57.53 76.27 82.25 32.33 46.68 50.46 58.78
    Backbone: GPT-4-1106 (Accuracy)
    QA prompt w/o IR 54.0 74.2 56.2 61.5
    Retrieve-Read 61.8 70.6 56.8 63.1
    + Chain-of-Thought 63.6 71.2 58.4 64.4
    + CHAIN-OF-NOTE (ours) 63.8 74.6 58.8 65.7

    When evaluated on LLaMA-2 7B, CoN achieves an average Exact Match (EM) improvement of +1.97 over standard Retrieve-Read. On NQ, when DPR retrieves relevant documents, CoN yields a +1.2 EM improvement, but when DPR fails to retrieve relevant documents, the gain increases to +2.3 EM, showing CoN's advantage in handling unhelpful retrieval outputs. On GPT-4-1106, CoN outperforms Chain-of-Thought prompting by +1.3 average accuracy.

  5. Knowl 5 — Noise Robustness Across Controlled Document Noise Ratios

    empirical result

    To evaluate noise robustness, subsets of test sets are constructed where the ratio rr of noisy (irrelevant) documents among the top-kk retrieved documents (k=5k=5) varies from 0%0\% (all relevant) to 100%100\% (all noisy/irrelevant). CoN substantially mitigates the degradation caused by irrelevant context in LLaMA-2 7B:

    Models Noise NQ TriviaQA WebQ Average
    Ratio EM F1 EM F1 EM F1 EM F1
    Retrieve-Read 100% 34.28 41.74 55.30 61.67 29.58 46.34 39.72 49.92
    + CHAIN-OF-NOTE 100% 41.83 49.58 64.30 70.00 36.85 53.07 47.66 57.55
    (+7.55) (+7.84) (+9.00) (+8.33) (+7.27) (+6.73) (+7.94) (+7.63)
    Retrieve-Read 80% 54.28 61.03 73.83 80.02 35.46 52.70 54.52 64.58
    + CHAIN-OF-NOTE 80% 56.63 63.23 75.89 81.24 40.60 56.54 57.70 67.00
    Retrieve-Read 60% 61.44 67.94 78.44 83.65 37.01 54.16 58.96 68.58
    + CHAIN-OF-NOTE 60% 63.43 69.33 78.79 84.07 41.26 56.91 61.16 70.10
    Retrieve-Read 40% 64.62 71.12 80.56 86.76 38.40 55.60 61.19 71.16
    + CHAIN-OF-NOTE 40% 65.91 72.22 81.72 87.11 42.16 58.15 63.26 72.49
    Retrieve-Read 20% 67.21 73.69 81.73 87.89 39.95 56.66 62.96 72.75
    + CHAIN-OF-NOTE 20% 70.00 76.08 82.86 88.24 44.36 60.13 65.74 74.82
    Retrieve-Read 0% 69.23 75.57 83.34 89.44 42.24 58.59 64.93 74.53
    + CHAIN-OF-NOTE 0% 73.28 79.86 83.52 88.94 46.16 62.38 67.65 77.06

    Under 100%100\% noisy documents, the standard RALM's average EM drops to 39.72 (below the 48.49 full test set baseline). CoN achieves 47.66 EM, providing a +7.94 EM (+7.63 F1) improvement and enabling the model to filter out irrelevant text and rely on parametric knowledge.

  6. Knowl 6 — Unknown Robustness on Real-Time Queries

    empirical result

    Unknown robustness measures the model's ability to refrain from hallucinating and instead reject questions that lie entirely outside its pre-training temporal knowledge scope and lack supporting evidence in retrieved documents. Evaluating LLaMA-2 7B on RealTimeQA (using questions dated after May 2023, post-dating LLaMA-2's pre-training cutoff) yields:

    Models EM F1 RR (%)
    Retrieve-Read 15.6 19.9 6.1
    + CHAIN-OF-NOTE (ours) 15.7 20.3 13.0

    where Reject Rate (RR) is defined as number of rejected questionstotal questions\frac{\text{number of rejected questions}}{\text{total questions}}. Chain-of-Note more than doubles the rejection rate from 6.1%6.1\% to 13.0%13.0\% while maintaining comparable exact match and F1 accuracy on queries that remained factually constant over time.

  7. Knowl 7 — Inference Latency and Robustness Trade-Off of Hybrid Training

    empirical result

    Explicit note generation introduces substantial decoding time overhead, which hybrid training completely resolves without sacrificing noise robustness.

    Models Inference Time (seconds/example)
    Retrieve-Read 0.6104
    + CHAIN-OF-NOTE (explicit) 12.0192
    + CHAIN-OF-NOTE (hybrid) 0.6074

    Measured on 8×A1008 \times \text{A100} GPUs:

    • Standard explicit CoN requires 12.0192s per example (a ∼20×\sim 20\times slowdown over standard Retrieve-Read at 0.6104s).
    • Hybrid CoN achieves an inference time of 0.6074s per example, matching standard Retrieve-Read.
    • On the NQ noise robustness curve, hybrid CoN maintains performance close to explicit CoN across all noise ratios: at 100%100\% noise, standard RALM gets 34.3 EM, hybrid CoN gets 36.3 EM, and explicit CoN gets 41.8 EM; at 0%0\% noise, standard RALM gets 69.2 EM, hybrid CoN gets 73.1 EM, and explicit CoN gets 73.3 EM.
  8. Knowl 8 — Inference Overhead and Note-Quality Dependency in Chain-of-Note

    limitation

    The primary limitation of explicit Chain-of-Note is the increased inference cost and latency resulting from autoregressive generation of sequential reading notes for every retrieved passage (12.02 s12.02\text{ s} per example vs. 0.61 s0.61\text{ s} for standard retrieval-augmented generation). While the hybrid training approach reduces this latency during inference, performance slightly drops relative to explicit note generation. Furthermore, system efficacy depends heavily on the conciseness and quality of the generated notes, which can vary depending on the length and structural complexity of the retrieved documents.

Coverage note — None was omitted; all contributed models, dataset construction steps, experimental benchmarks, empirical comparisons across noise/unknown robustness, efficiency analyses, and stated limitations are fully represented.

References

  1. 1.Jonathan Berant, Andrew Chou, Roy Frostig, and Percy Liang. 2013. Semantic parsing on freebase from question-answer pairs. In EMNLP, pages 1533–1544.
  2. 2.Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George Bm Van Den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, et al. 2022. Improving language models by retrieving from trillions of tokens. In International conference on machine learning, pages 2206–2240. PMLR.
  3. 3.Danqi Chen, Adam Fisch, Jason Weston, and Antoine Bordes. 2017. Reading wikipedia to answer open-domain questions. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1870–1879.
  4. 4.Hao Cheng, Yelong Shen, Xiaodong Liu, Pengcheng He, Weizhu Chen, and Jianfeng Gao. 2021. Unitedqa: A hybrid approach for open domain question answering. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 3080–3090.
  5. 5.Antonia Creswell, Murray Shanahan, and Irina Higgins. 2022. Selection-inference: Exploiting large language models for interpretable logical reasoning. arXiv preprint arXiv:2205.09712.
  6. 6.Bhuwan Dhingra, Jeremy R Cole, Julian Martin Eisenschlos, Daniel Gillick, Jacob Eisenstein, and William W Cohen. 2022. Time-aware language models as temporal knowledge bases. Transactions of the Association for Computational Linguistics, 10:257–273.
  7. 7.Shehzaad Dhuliawala, Mojtaba Komeili, Jing Xu, Roberta Raileanu, Xian Li, Asli Celikyilmaz, and Jason Weston. 2023. Chain-of-verification reduces hallucination in large language models. arXiv preprint arXiv:2309.11495.
  8. 8.Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang. 2020. Realm: Retrieval-augmented language model pre-training. arXiv preprint arXiv:2002.08909.
  9. 9.Fan Huang, Haewoon Kwak, and Jisun An. 2023. Chain of explanation: New prompting method to generate quality natural language explanation for implicit hate speech. In Proceedings of the ACM Web Conference 2023, pages 90–93.
  10. 10.Gautier Izacard and Edouard Grave. 2021. Leveraging passage retrieval with generative models for open domain question answering. In EACL, pages 874–880.
  11. 11.Gautier Izacard, Patrick Lewis, Maria Lomeli, Lucas Hosseini, Fabio Petroni, Timo Schick, Jane Dwivedi-Yu, Armand Joulin, Sebastian Riedel, and Edouard Grave. 2022. Few-shot learning with retrieval augmented language models. arXiv preprint arXiv:2208.03299.
  12. 12.Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023. Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12):1–38.
  13. 13.Mandar Joshi, Eunsol Choi, Daniel S Weld, and Luke Zettlemoyer. 2017. Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension. In ACL, pages 1601–1611.
  14. 14.Nikhil Kandpal, Haikang Deng, Adam Roberts, Eric Wallace, and Colin Raffel. 2023. Large language models struggle to learn long-tail knowledge. In International Conference on Machine Learning, pages 15696–15707. PMLR.
  15. 15.Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense passage retrieval for open-domain question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6769–6781.
  16. 16.Jungo Kasai, Keisuke Sakaguchi, Yoichi Takahashi, Ronan Le Bras, Akari Asai, Xinyan Yu, Dragomir Radev, Noah A Smith, Yejin Choi, and Kentaro Inui. 2023. Realtime qa: What’s the answer right now? Advances in Neural Information Processing Systems.
  17. 17.Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis. 2020. Generalization through memorization: Nearest neighbor language models. In International Conference on Learning Representations.
  18. 18.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. arXiv preprint arXiv:2205.11916.
  19. 19.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al. 2019. Natural questions: A benchmark for question answering research. TACL, pages 452–466.
  20. 20.Angeliki Lazaridou, Elena Gribovskaya, Wojciech Stokowiec, and Nikolai Grigorev. 2022. Internet-augmented language models through few-shot prompting for open-domain question answering. arXiv preprint arXiv:2203.05115.
  21. 21.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in Neural Information Processing Systems, 33:9459–9474.
  22. 22.Xianzhi Li, Xiaodan Zhu, Zhiqiang Ma, Xiaomo Liu, and Sameena Shah. 2023. Are chatgpt and gpt-4 general-purpose solvers for financial text analytics? an examination on several typical tasks. arXiv preprint arXiv:2305.05862.
  23. 23.Hongyin Luo, Yung-Sung Chuang, Yuan Gong, Tianhua Zhang, Yoon Kim, Xixin Wu, Danny Fox, Helen Meng, and James Glass. 2023. Sail: Search-augmented instruction learning. arXiv preprint arXiv:2305.15225.
  24. 24.Ji Ma, Ivan Korotkov, Yinfei Yang, Keith Hall, and Ryan McDonald. 2021. Zero-shot neural passage retrieval via domain-targeted synthetic question generation. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 1075–1088.
  25. 25.Kaixin Ma, Hao Cheng, Yu Zhang, Xiaodong Liu, Eric Nyberg, and Jianfeng Gao. 2023. Chain-of-skills: A configurable model for open-domain question answering. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistic.
  26. 26.Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Daniel Khashabi, and Hannaneh Hajishirzi. 2023. When not to trust language models: Investigating effectiveness of parametric and non-parametric memories. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 9802–9822.
  27. 27.OpenAI. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774.
  28. 28.Chengwei Qin, Aston Zhang, Zhuosheng Zhang, Jiaao Chen, Michihiro Yasunaga, and Diyi Yang. 2023. Is chatgpt a general-purpose natural language processing task solver? arXiv preprint arXiv:2302.06476.
  29. 29.Yingqi Qu, Yuchen Ding, Jing Liu, Kai Liu, Ruiyang Ren, Wayne Xin Zhao, Daxiang Dong, Hua Wu, and Haifeng Wang. 2021. Rocketqa: An optimized training approach to dense passage retrieval for open-domain question answering. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 5835–5847.
  30. 30.Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. 2020. Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pages 3505–3506.
  31. 31.Devendra Singh Sachan, Mike Lewis, Dani Yogatama, Luke Zettlemoyer, Joelle Pineau, and Manzil Zaheer. 2022. Questions are all you need to train a dense passage retriever. arXiv preprint arXiv:2206.10658.
  32. 32.Freda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales, David Dohan, Ed H Chi, Nathanael Schärli, and Denny Zhou. 2023a. Large language models can be easily distracted by irrelevant context. In International Conference on Machine Learning, pages 31210–31227. PMLR.
  33. 33.Freda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang, Suraj Srivats, Soroush Vosoughi, Hyung Won Chung, Yi Tay, Sebastian Ruder, Denny Zhou, et al. 2023b. Language models are multilingual chain-of-thought reasoners. In The Eleventh International Conference on Learning Representations.
  34. 34.Weijia Shi, Sewon Min, Michihiro Yasunaga, Minjoon Seo, Rich James, Mike Lewis, Luke Zettlemoyer, and Wen-tau Yih. 2023c. Replug: Retrieval-augmented black-box language models. arXiv preprint arXiv:2301.12652.
  35. 35.Devendra Singh, Siva Reddy, Will Hamilton, Chris Dyer, and Dani Yogatama. 2021. End-to-end training of multi-document reader and retriever for open-domain question answering. Advances in Neural Information Processing Systems, 34:25968–25981.
  36. 36.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  37. 37.Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal. 2023. Interleaving retrieval with chain-of-thought reasoning for knowledge-intensive multi-step questions. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics.
  38. 38.Tu Vu, Mohit Iyyer, Xuezhi Wang, Noah Constant, Jerry Wei, Jason Wei, Chris Tar, Yun-Hsuan Sung, Denny Zhou, Quoc Le, et al. 2023. Freshllms: Refreshing large language models with search engine augmentation. arXiv preprint arXiv:2310.03214.
  39. 39.Jianing Wang, Qiushi Sun, Nuo Chen, Xiang Li, and Ming Gao. 2023a. Boosting language models reasoning with chain-of-knowledge prompting. arXiv preprint arXiv:2306.06427.
  40. 40.Keheng Wang, Feiyu Duan, Sirui Wang, Peiguang Li, Yunsen Xian, Chuantao Yin, Wenge Rong, and Zhang Xiong. 2023b. Knowledge-driven cot: Exploring faithful reasoning in llms for knowledge-intensive question answering. arXiv preprint arXiv:2308.13259.
  41. 41.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. 2022. Chain of thought prompting elicits reasoning in large language models. arXiv preprint arXiv:2201.11903.
  42. 42.Ori Yoran, Tomer Wolfson, Ori Ram, and Jonathan Berant. 2023. Making retrieval-augmented language models robust to irrelevant context. arXiv preprint arXiv:2310.01558.
  43. 43.Donghan Yu, Chenguang Zhu, Yuwei Fang, Wenhao Yu, Shuohang Wang, Yichong Xu, Xiang Ren, Yiming Yang, and Michael Zeng. 2022. Kg-fid: Infusing knowledge graph in fusion-in-decoder for open-domain question answering. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 4961–4974.
  44. 44.Wenhao Yu, Dan Iter, Shuohang Wang, Yichong Xu, Mingxuan Ju, Soumya Sanyal, Chenguang Zhu, Michael Zeng, and Meng Jiang. 2023a. Generate rather than retrieve: Large language models are strong context generators. International Conference for Learning Representation (ICLR).
  45. 45.Wenhao Yu, Zhihan Zhang, Zhenwen Liang, Meng Jiang, and Ashish Sabharwal. 2023b. Improving language models via plug-and-play retrieval feedback. arXiv preprint arXiv:2305.14002.
  46. 46.Yue Zhang, Yafu Li, Leyang Cui, Deng Cai, Lemao Liu, Tingchen Fu, Xinting Huang, Enbo Zhao, Yu Zhang, Yulong Chen, et al. 2023a. Siren’s song in the ai ocean: A survey on hallucination in large language models. arXiv preprint arXiv:2309.01219.
  47. 47.Zhuosheng Zhang, Aston Zhang, Mu Li, Hai Zhao, George Karypis, and Alex Smola. 2023b. Multimodal chain-of-thought reasoning in language models. arXiv preprint arXiv:2302.00923.
  48. 48.Zexuan Zhong, Tao Lei, and Danqi Chen. 2022. Training language models with memory augmentation. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 5657–5673.
  49. 49.Fengbin Zhu, Wenqiang Lei, Chao Wang, Jianming Zheng, Soujanya Poria, and Tat-Seng Chua. 2021. Retrieving and reading: A comprehensive survey on open-domain question answering. arXiv preprint arXiv:2101.00774.

Citation

MLA
Yu, W., et al. “Chain-of-Note: Enhancing Robustness in Retrieval-Augmented Language Models”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024, pp. 14672–85, https://doi.org/10.18653/v1/2024.emnlp-main.813.
APA
Yu, W., Zhang, H., Pan, X., Cao, P., Ma, K., Li, J., Wang, H., & (于东), D. Y. (2024). Chain-of-Note: Enhancing Robustness in Retrieval-Augmented Language Models. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 14672–14685. https://doi.org/10.18653/v1/2024.emnlp-main.813
Chicago
Yu, W., H. Zhang, X. Pan, et al. 2024. “Chain-of-Note: Enhancing Robustness in Retrieval-Augmented Language Models”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 14672–85. https://doi.org/10.18653/v1/2024.emnlp-main.813.
Harvard
Yu, W. et al. (2024) “Chain-of-Note: Enhancing Robustness in Retrieval-Augmented Language Models”, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 14672–14685. Available at: https://doi.org/10.18653/v1/2024.emnlp-main.813.
Vancouver
1. Yu W, Zhang H, Pan X, Cao P, Ma K, Li J, Wang H, (于东) DY (2024) Chain-of-Note: Enhancing Robustness in Retrieval-Augmented Language Models. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 14672–14685

BibTeX

@inproceedings{yu-etal-2024-chain,
    title = "Chain-of-Note: Enhancing Robustness in Retrieval-Augmented Language Models",
    author = "Yu, Wenhao  and
      Zhang, Hongming  and
      Pan, Xiaoman  and
      Cao, Peixin  and
      Ma, Kaixin  and
      Li, Jian  and
      Wang, Hongwei  and
      Yu, Dong",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.813/",
    doi = "10.18653/v1/2024.emnlp-main.813",
    pages = "14672--14685"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/