Context-DPO: Aligning Language Models for Context-Faithfulness
Baolong Bi 1^{1}1, Shaohan Huang 2^{2}2, Yiwei Wang 3^{3}3, Tianchi Yang 2^{2}2, Zihan Zhang 2^{2}2
Haizhen Huang 2^{2}2, Lingrui Mei 1^{1}1, Junfeng Fang 4^{4}4, Zehao Li 1^{1}1, Furu Wei 2^{2}2
Weiwei Deng 2^{2}2, Feng Sun 2^{2}2, Qi Zhang 2^{2}2, Shenghua Liu 1^{1}1 ∗^{*}∗
1^{1}1 University of Chinese Academy of Sciences 2^{2}2 Microsoft Corporation
3^{3}3 University of California, Merced 4^{4}4 National University of Singapore
[email protected], [email protected]
Haizhen Huang 2^{2}2, Lingrui Mei 1^{1}1, Junfeng Fang 4^{4}4, Zehao Li 1^{1}1, Furu Wei 2^{2}2
Weiwei Deng 2^{2}2, Feng Sun 2^{2}2, Qi Zhang 2^{2}2, Shenghua Liu 1^{1}1 ∗^{*}∗
1^{1}1 University of Chinese Academy of Sciences 2^{2}2 Microsoft Corporation
3^{3}3 University of California, Merced 4^{4}4 National University of Singapore
[email protected], [email protected]
∗^{*}∗ Corresponding Author
Abstract
Reliable responses from large language models (LLMs) require adherence to user instructions and retrieved information. While alignment techniques help LLMs align with human intentions and values, improving context-faithfulness through alignment remains underexplored. To address this, we propose Context-DPO, the first alignment method specifically designed to enhance LLMs' context-faithfulness. We introduce ConFiQA, a benchmark that simulates Retrieval-Augmented Generation (RAG) scenarios with knowledge conflicts to evaluate context-faithfulness. By leveraging faithful and stubborn responses to questions with provided context from ConFiQA, our Context-DPO aligns LLMs through direct preference optimization. Extensive experiments demonstrate that our Context-DPO significantly improves context-faithfulness, achieving 35% to 280% improvements on popular open-source models. Further analysis demonstrates that Context-DPO preserves LLMs' generative capabilities while providing interpretable insights into context utilization.
1. Introduction
With the widespread deployment of Retrieval-Augmented Generation (RAG) ([1]) and various tools ([2]), large language models (LLMs) ([3, 4, 5, 6]) are increasingly expected to generate responses that adhere closely to provided context, including retrieved information and user instructions. Consequently, context-faithfulness ([7, 8, 9]) has become a critical capability for modern LLM applications, especially in scenarios where parametric knowledge is insufficient or outdated. However, this expectation is challenged by knowledge conflicts ([10, 11, 12]). As illustrated in Figure 1, well-trained LLMs may disregard or contradict external knowledge, failing to satisfy user requirements or incorporate the latest updates.
Existing efforts to enhance the context-faithfulness of LLMs primarily focus on external interventions, such as designing prompts to encourage context integration ([7]) or modifying decoding strategies ([13, 14]) to increase the output probability of relevant tokens. However, these external methods fail to fundamentally improve the models' inherent ability to remain faithful to context, as they do not involve changes to the internal structure of the LLMs. In contrast, alignment techniques ([15, 16]), which aim to make pre-trained LLMs behave in line with human intentions and values, have proven effective in enhancing critical capabilities such as factuality ([17]) and safety ([18]). Despite its importance as a core attribute, context-faithfulness remains an underexplored area in alignment research.
In this work, we present the first exploration of aligning LLMs for context-faithfulness, aiming to reliably enhance their adherence to contextual information. To achieve this, we first propose ConFiQA (Context Faithfulness Question Answering), a novel benchmark designed to evaluate context-faithfulness through question-answering tasks based on counterfactual retrieval passages. ConFiQA tests whether models can generate responses consistent with contexts containing counterfactual elements, simulating real-world scenarios with knowledge conflicts in modern RAG systems. We evaluate current popular LLMs on ConFiQA and find that most models exhibit poor performance in context-faithfulness to varying degrees. Furthermore, our results reveal that context-faithfulness tends to decline as model size increases and training becomes more refined.
Therefore, we argue that modern LLMs also require alignment specifically for context-faithfulness. To address this, we propose Context-DPO, which constructs reasoning chains based on single-hop or multi-hop knowledge to generate two types of responses: faithful (grounded in counterfactual context) and stubborn (based on factual reality). Context-DPO uses preference pairs derived from these responses to reward context-faithful behavior and fine-tune the model via the Direct Preference Optimization (DPO) ([19]).
We conduct experiments on our ConFiQA, Natural Questions ([20]) , and MQUAKE{\mathchoice{\text{MQ{\scriptsize U}A{\scriptsize KE}}}{\text{MQ{\scriptsize U}A{\scriptsize KE}}}{\text{MQ{\scriptscriptstyle U}A{\scriptscriptstyle KE}}}{\text{MQUAKE}}}MQUAKE ([21]) datasets, covering counterfactual retrieval-based question-answering tasks and in-context editing tasks that require following user instructions. Extensive results demonstrate that our Context-DPO effectively aligns LLMs to improve context-faithfulness, consistently outperforming all existing baselines without requiring any external prompt modifications. Specifically, the aligned models achieved substantial improvements compared to their original versions: 35% for Llama2-7B-chat, 78% for Llama3-8B, 151% for Mistral-7B and 280% for Qwen2-7B.
We also conduct interpretability analyses to investigate the context-faithfulness of LLMs. By identifying key generating tokens that effectively distinguish between contextual and parametric knowledge, we analyze the logits and ranking distribution in thses key tokens to reveal why the aligned models exhibit improved faithfulness to context. Additionally, further experiments on TruthfulQA ([22]) demonstrate that models aligned using Context-DPO retain their foundational generative capabilities, indicating that this alignment process has no negative impact.
In summary, our contributions are three-fold:
- We propose ConFiQA, a novel benchmark for evaluating context-faithfulness through question-answering tasks based on counterfactual retrieval passages.
- We introduce Context-DPO, the first alignment method to enhance context-faithfulness, with experiments proving its effectiveness in improving LLMs' adherence to context.
- We uncover the underlying reasons for the improved context-faithfulness of aligned models and confirm that this alignment has no negative impact on their generative performance.
2. ConFiQA: Context Faithfulness Question Answering Benchmark
We introduce the ConFiQA benchmark to evaluate the context-faithfulness of LLMs in real-world RAG scenarios involving knowledge conflicts. ConFiQA consists of three datasets: QA (Question-Answering), MR (Multi-hop Reasoning), and MC (Multi-Conflicts). QA features single-hop question-answering tasks with context containing one corresponding counterfactual, while MR and MC involve multi-hop reasoning tasks with context containing one and multiple related counterfactuals, respectively. In this section, we present the data construction pipeline, provide an overview of the datasets, and evaluate the context-faithfulness of popular LLMs using ConFiQA.
2.1 Data Construction Pipeline
Real-World Fact Sampling To ensure the factuality of the subsequently generated context, we collect triples from Wikidata1 ([23]) to guide the generation of real-world facts. Prior to this, we gather popular entities from Wikipedia2 to facilitate triple sampling, ensuring that LLMs have a strong memory of the generated facts. Using 41 manually selected relations (Table 7 in Appendix) and maintaining a one-to-one correspondence between head entities and tail entities for each relation, we ultimately collected 5,042 entities and 30,295 triples.
1.
Wikidata is a publicly accessible, continuously updated knowledge base of factual triples
2.
We collect entities corresponding to the top 1,000 most-visited Wikipedia pages from 2016 to 2023, based on monthly page views, and retained the most popular entities using criteria such as the number of hyperlinks.
Multi-Hop Path Construction
We construct a factual subgraph Gsub\mathcal{G}_{sub}Gsub based on the sampled triples and then extract 2,3,42,3,42,3,4-hop paths Pf={(s1,r1,t1),…,(sn,rn,tn)}n≤4\mathcal{P}^f=\{(s_1, r_1, t_1), \dots, (s_n, r_n, t_n)\}_{n \leq 4}Pf={(s1,r1,t1),…,(sn,rn,tn)}n≤4 from the subgraph. For MR, we randomly select one triple (si,ri,ti)(s_i, r_i, t_i)(si,ri,ti) from the paths and replace tit_iti with a same-type entity ti′t_i'ti′. The subsequent path is then sampled from ti′t_i'ti′ in the subgraph to ensure the remaining path remains factual. For MC, we perform the same replacement for every triple in fatual path Pf\mathcal{P}^fPf, ensuring that each triple becomes counterfactual. In the generated multi-hop paths with counterfactuals Pc\mathcal{P}^cPc, the head entity of the next hop matches the tail entity of the previous hop, and the relation in each triple remains unchanged before and after replacement, thereby maintaining the validity of multi-hop reasoning.
Counterfactual Context Generation
We apply the same tail entity replacement to provide counterfactual triples (s,r,t′)(s, r, t')(s,r,t′) for QA. Using the triples, we generate context that incorporates its corresponding factual information. This is achieved by prompting ChatGPT-4 to generate a description of entity sss, ensuring that the triple's factual information is embedded within the context (details are provided in Appendix B). To avoid issues of context being ignored or contradicted due to knowledge conflicts ([8]), we first generate factual context based on the original triples, and then replace the tail entity ttt with counterfactual t′t't′ in the context. For MR and MC, we sequentially generate context for all triples along the original multi-hop paths and concatenate them, performing all necessary counterfactual replacements. These replacements, which include handling all aliases and morphological variations of the entities, along with other rules3, ensure semantic and logical coherence. The generated context contains counterfactual fragments alongside accurate descriptions of entities, effectively simulating real-world RAG scenarios involving knowledge conflicts and retrieval noise.
3.
Entities and relations in the sampled path are not repeated
2.2 Overview of Datasets
We use ChatGPT-4 to generate questions based on single-hop triples or multi-hop paths. Each question incorporates the head entity of the first hop and the relationships in each subsequent hop, guiding the model to predict the final tail entity (see Table 8 in Appendix for details). For each dataset in our ConFiQA benchmark, we sample 6,000 instances, with the specific format detailed in Table 1. For MR and MC, the data is evenly distributed across 2,3,42,3,42,3,4-hop paths (see examples in Appendix G).
2.3 Evaluation Metrics
We follow the evaluation metrics defined in [24, 7], but given that LLMs' responses may contain negations or refutations of the counterfactual answer, we apply stricter criteria for PcP_cPc compared to the previous PsP_sPs (substitute answers). Specifically, We use the following four metrics to compare the normalized responses with the normalized answers to evaluate the context-faithfulness of LLMs:
- Pc(↑)P_c (\uparrow)Pc(↑): Frequency of responses matching the context-faithful answer or its aliases, excluding negations or the original answer. Context-faithful answers are counterfactual answers derived from the context.
- Po(↓)P_o (\downarrow)Po(↓): Frequency of responses matching the original factual answer or its aliases.
- MR(↓)M_R (\downarrow)MR(↓): Proportion of responses predicting the correct answer but reluctant to update their predictions, calculated as MR=PoPs+PoM_R = \frac{P_o}{P_s+P_o}MR=Ps+PoPo.
- EM(↑)EM (\uparrow)EM(↑): Frequency of responses exactly matching the context-faithful answer.
2.4 Evaluation on ConFiQA
We use our ConFiQA to evaluate the context-faithfulness of popular open-source models (Llama2-7B-chat, Llama2-13B-chat, Mistral-7B-instruct-v0.2, Qwen2-7B-instruct) and close-source models (ChatGPT-4, Gemini-1.5-pro, ChatGPT-4o). The experimental results, presented in Table 2, reveal the following key findings:
- Despite alignment efforts, such as instruct-tuning, to meet human standards, the tested LLMs exhibit significant deficiencies in context-faithfulness. Most models have an MRM_RMR exceeding 50%, particularly the latest ones, indicating that they tend to rely on their own judgments over the provided context.
- A counterintuitive trend is observed: as model size increases (e.g., Llama2-chat from 7B to 13B) or as models become more advanced (e.g., the latest Llama3-8B-instruct compared to earlier versions like Llama2-7B-chat), their context-faithfulness tends to decline.
These findings indicate that current LLMs generally exhibit poor alignment in context-faithfulness. Furthermore, with advancements in data processing and model training, more advanced models tend to become increasingly confident in their parametric knowledge, resulting in worse context-faithfulness when facing conflicts between contextual and parametric knowledge. This poses significant challenges for tasks that require strict adherence to external knowledge, such as RAG or other specialized, closed-domain applications.
3. Context-DPO: Context-Faithful Direct Preference Optimization
Based on the unsatisfactory performance of existing LLMs in context-faithfulness, we argue that it is essential to specifically align LLMs for context-faithfulness. To address this, we propose Context-DPO, the first alignment approach dedicated to enhancing context-faithfulness by creating preference data and aligning LLMs with DPO. The framework of our Context-DPO is shown in Figure 2.
3.1 Preference Data Generation
Leveraging the counterfactual and factual data provided by ConFiQA, we can construct preference data D=(x,yw,yl)\mathcal{D} = {(x, y_w, y_l)}D=(x,yw,yl) efficiently. Specifically, the input xxx is formed by concatenating the counterfactual context Cc\mathcal{C}^cCc and the question Q\mathcal{Q}Q. To generate the counterfactual reasoning chain, each triple in the counterfactual path is transformed into a textual description using a statement template (Table 7) and sequentially concatenated. Finally, the reasoning chain concludes by summarizing the reasoning process to derive the counterfactual answer, which is determined based on the last tail entity in the chain. This process yields faithful responses ywy_wyw grounded in the counterfactual context. Similarly, stubborn responses yly_lyl, grounded in factual reality, are constructed by following the original factual path. This approach to constructing the preference dataset D\mathcal{D}D simulates the reasoning pattern observed in real-world RAG tasks and mirrors the chain-of-thought ([25]) process of LLMs when producing final answers.
3.2 Context-Faithful Alignment with DPO
We leverage the generated preference data to perform alignment tuning on LLMs for context-faithfulness. While several frameworks exist for alignment training, including the widely adopted RLHF framework, which involves training a reward model on preference data and optimizing the policy using the Proximal Policy Optimization (PPO) algorithm, we employ DPO for context-faithful alignment. DPO, as a recent approach to preference optimization, enables the policy πθ\pi_\thetaπθ to be learned directly from a fixed preference dataset without requiring an explicit reward model or sampling from the policy during training, as is necessary with PPO. Specifically, our Context-DPO uses the standard cross-entropy objective, and its training objective is formulated as follows:
In this formulation, the model policy πθ\pi_\thetaπθ is initialized using the base reference policy πref\pi_{\mathrm{ref}}πref. The parameter β\betaβ regulates the extent of divergence from πref\pi_{\mathrm{ref}}πref, while σ\sigmaσ represents the logistic function.
4. Experiments
4.1 Experimental Setup
Tasks. We evaluate context-faithfulness using the following two tasks: Retrieval Following and Instruction Following. For Retrieval Following, we adopt the setup described in Section 2.4 to assess faithfulness to retrieved passages containing noise and relevant counterfactuals. In contrast, Instruction Following focuses solely on textual editing instructions, testing whether LLMs can effectively adhere to user commands.
Datasets.
We conduct experiments for Retrieval Following using both our ConFiQA and Natural Questions ([20]). In Natural Questions, the context is modified to support counterfactual answers following by [24]. For the Instruction Following task, we utilize the MQUAKE{\mathchoice{\text{MQ{\scriptsize U}A{\scriptsize KE}}}{\text{MQ{\scriptsize U}A{\scriptsize KE}}}{\text{MQ{\scriptscriptstyle U}A{\scriptscriptstyle KE}}}{\text{MQUAKE}}}MQUAKE dataset ([21]), which provides multi-hop questions and in-context editing instructions to assess context-faithfulness in response to counterfactual edits.
Models and Baselines.
We use current popular open-source LLMs (Llama2-7B-chat, Llama2-13B-chat, Mistral-7B-instruct-v0.2, and Qwen2-7B-instruct) as the base models for our experiments. For the Retrieval Following task, we use two prompt-based baselines: the attributed prompt (Attr) and the combination of opinion-based and instruction-based prompts (O&I) ([7]). Additionally, we also fine-tune the LLMs using faithful responses from ConFiQA as the training-based baseline (SFT). For the Instruction Following task, we follow the approach of IKE ([26]), which evaluates the in-context editing capabilities of both the base model and the Context-DPO-aligned model through contextual editing demonstrations. Detailed implementation and prompt templates for these baselines can be found in Appendix D.
4.2 Performance on Retrieval Following
Experimental results for Retrieval Following are shown in Table 3 and Table 4 on our ConFiQA and Natural Questions datasets, respectively. Models aligned with our Context-DPO method significantly outperform all baselines, without requiring any additional prompts. On all tasks in ConFiQA, Llama2-7B-chat, Llama2-13B-chat, Mistral-7B-instruct-v0.2, and Qwen2-7B-instruct show average improvements of 35.2%, 78.3%, 151.8%, and 280.1%, respectively, in PcP_cPc after alignment with our Context-DPO, compared to their initial models. On the Natural Questions dataset, where knowledge conflicts are less pronounced, the accuracy of our method reaches over 93% on average.
This demonstrates that our Context-DPO method is highly effective in significantly improving the context-faithfulness of LLMs. Notably, our approach enhances the model’s fundamental context-faithfulness capability through alignment tuning, without relying on inference-stage enhancement methods used by the baselines. This indicates that aligned models have considerable potential for further improvement. Furthermore, the results reveal that simply applying end-to-end SFT is insufficient to effectively enhance the context-faithfulness of LLMs, often performing worse than prompt-based methods. This limitation arises because SFT fails to generalize the training objective of improving context-faithfulness. In contrast, DPO proves to be an effective alternative, as it captures the training signal for context-faithfulness more robustly through preference pair comparisons.
4.3 Performance on Instruction Following
We evaluate LLMs' Instruction Following ability with the in-context editing task on MQUAKE{\mathchoice{\text{MQ{\scriptsize U}A{\scriptsize KE}}}{\text{MQ{\scriptsize U}A{\scriptsize KE}}}{\text{MQ{\scriptscriptstyle U}A{\scriptscriptstyle KE}}}{\text{MQUAKE}}}MQUAKE dataset, where instruction-based textual prompts are used to guide the models in editing relevant knowledge to answer questions. The context demonstrations we used are provided in Appendix C. Table 5 presents the few-shot accuracy of models, both before and after alignment with our Context DPO, under varying numbers of demonstration prompts. While increasing the number of context demonstrations encourages LLMs to better follow the editing instructions, the aligned models consistently outperform the baselines. This demonstrates that the context-faithfulness alignment based on our ConFiQA, which simulates question-answering according to the retrieved passage, also enhances the model’s faithfulness to user instructions.
4.4 Validation of the Decoupled Improvement in LLMs' Context-Faithfulness
As mentioned by [8], there may be a trade-off between context-faithfulness and factuality in LLMs. To validate this for our method, we evaluate whether the Context-DPO alignment affects the model’s factual generation ability. Using TruthfulQA ([22]), we employ a multiple-choice task where the LLM selects an answer from a range of correct and incorrect options, evaluated by multiple-choice accuracy (MC1, MC2, and MC3). As shown in Table 6, the performance of the aligned models fluctuates by no more than 1% on average across the MC metrics, compared to the original models. This indicates that the improvements achieved by our Context-DPO alignment are decoupled: while enhancing context-faithfulness, the alignment does not negatively impact the model's inherent generation ability when no context is provided. Therefore, we strongly advocate for incorporating context-faithfulness alignment as a standard practice in LLM alignment.
5. In-depth Exploration of the Metamorphosis in Context-Faithfulness
Figure 3 provides an intuitive visualization of the impact of our Context-DPO on LLMs' context-faithfulness, demonstrating its ability to reduce irrelevant responses (other response) and stubborn reliance on parametric knowledge (stubborn response), ultimately leading to more context-faithful answers (context-faithful response). To further investigate the internal mechanisms behind the effective alignment of LLMs' context-faithfulness by our Context-DPO, we utilize the knowledge token capturing algorithm proposed by [8]. for deeper exploration. The algorithm (detailed in Algorithm 1 in Appendix) identifies the tokens with the highest probability of distinguishing between contextual knowledge and parametric knowledge by matching decoded tokens with their corresponding knowledge strings. Following the Instruction Following task (Section 4.3), we collected 2,000 question-answer instances from the MQUAKE{\mathchoice{\text{MQ{\scriptsize U}A{\scriptsize KE}}}{\text{MQ{\scriptsize U}A{\scriptsize KE}}}{\text{MQ{\scriptscriptstyle U}A{\scriptscriptstyle KE}}}{\text{MQUAKE}}}MQUAKE dataset to capture the logits distribution of key tokens, which effectively highlights the distinction between context-faithful responses and stubborn responses.
We calculate the average logits of key tokens representing context-faithfulness, with the results shown in Figure 4. Aligned models exhibit significant improvements over the base models, with gains ranging from 16.8 to 21.0, indicating that our Context-DPO effectively increases the probability of generating context-faithful responses. We further analyze the softmax-transformed logits distribution of these tokens, as shown in Figure 5. The results indicate that models aligned with Context-DPO reduce the distribution in low-probability regions while increasing it in high-probability regions compared to their original versions. This adjustment further increases the likelihood of decoding context-faithful tokens at key positions, leading to a significant rise in the generation frequency of top-ranked tokens, as illustrated in Figure 6. Our interpretability analysis uncovers the internal mechanisms behind the effective context-faithfulness alignment achieved by our Context-DPO. This highlights its ability to significantly enhance the upper bound of context-faithfulness without relying on external inference-stage methods.
6. Conclusion
In this work, we introduce ConFiQA, a novel benchmark that simulates real-world RAG scenarios and knowledge conflicts, enabling the evaluation of LLMs' context-faithfulness. To address shortcomings in context-faithfaulness for current models, we propose Context-DPO, the first alignment method dedicated to enhancing context-faithfulness. This approach leverages ConFiQA to construct preference data and fine-tunes models using DPO. Experimental results demonstrate that Context-DPO significantly enhances the context-faithfulness of popular LLMs without compromising their inherent generative capabilities. Furthermore, interpretability analysis reveals the mechanisms underlying the improvements in faithfulness. Our work paves the way to develop both effective and accountable context-faithfulness for LLMs.
Limitations
This paper focuses on specific knowledge conflict scenarios to better highlight context-faithfulness in evaluation. However, its application in typical real-world RAG scenarios has not been extensively validated. We believe that our Context-DPO can also bring significant benefits to standard RAG tasks, and we plan to explore this further in future work. Additionally, although our findings indicate in experiments that context-faithfulness tends to decline as model size increases and training becomes more refined, further extensive experiments are needed to fully validate this observation.
Ethical Considerations
Ethical considerations are paramount in our research. The proposed dataset, along with the open-source datasets and widely recognized models used in this study, strictly adheres to established ethical principles. Additionally, counterfactual data is employed in our experimental evaluations to measure context-faithfulness under knowledge conflict scenarios. The proposed methods are designed to ensure that models do not generate harmful or misleading information. Throughout this research, we remain committed to upholding ethical standards, prioritizing transparency, and fostering the responsible use of technology to benefit society.
Appendix
A. Related Work
Hallucinations in LLMs
The outputs of large language models (LLMs) often appear plausible at first glance but may exhibit various issues upon closer inspection, a phenomenon commonly referred to as hallucinations ([27, 28, 29, 30]). These hallucinations cause LLMs to produce content that deviates from user inputs, previously generated context, or factual knowledge, severely undermining their reliability in real-world applications ([31, 32, 33, 34, 35]). Such hallucinations can arise at different stages of the LLM lifecycle. Broadly, research on hallucination mitigation falls into two categories. During the training phase, studies such as [36, 37] have explored methods like training data curation and knowledge grounding to better integrate external knowledge into the model. Recent findings suggest that hallucinations often stem from conflicts between an LLM’s internal parameters and the external context provided during inference. In the inference stage, recent works have proposed methods such as confidence estimation ([38]), knowledge retrieval ([39, 40]), and knowledge editing (KE) ([41]) to generate more accurate outputs. These approaches aim to refine the model’s predictions by enhancing its ability to validate outputs or supplementing it with relevant external knowledge. Despite these advancements, addressing hallucinations remains a critical challenge for improving LLM reliability.
Knowledge Conflicts
Knowledge conflicts ([42, 43]) can be categorized into three types: internal conflicts within the context, conflicts between the memories encoded in model parameters, and conflicts between the context and model parameters. The latter, as a critical issue, has been extensively studied to mitigate hallucinations. Various popular tools ([44, 45, 2]) and retrieval-augmented methods ([1, 46, 47]), such as ChatGPT plugins and New Bing, have been introduced as effective strategies for providing external knowledge evidence. However, integrating external knowledge is not without challenges, as it sometimes conflicts with the parametric knowledge of LLMs ([11, 12]), resulting in inconsistent or unreliable outputs, especially when LLMs exhibit overconfidence in their inherent parametric knowledge. These conflicts between external sources and the internal knowledge stored within LLMs continue to pose significant challenges in ensuring reliable model performance.
Retrieval-Augmented Generation
RAG ([48, 49]) enhances LLMs by retrieving relevant document chunks from external knowledge bases based on semantic similarity. By leveraging external knowledge, RAG effectively reduces the generation of factually incorrect content, addressing a key challenge in LLM outputs. Its integration with LLMs has led to widespread adoption, significantly improving the reliability of LLM-based systems ([50]). However, context-faithfulness ([51, 52]) plays a crucial role in determining the performance of RAG, as the retrieved content may conflict with the internal parametric knowledge of LLMs, particularly when the parametric knowledge is insufficient or outdated. This challenge is exacerbated as LLMs grow in size and undergo more refined training, making them increasingly confident in their own parametric knowledge. Such overconfidence further undermines context-faithfulness in scenarios where knowledge conflicts arise.
In-Context Editing
As one of the most effective Knowledge Editing (KE) methods ([41, 53, 54, 55, 56]), in-context editing (ICE) ([57, 21, 26, 58, 59, 60]) has demonstrated state-of-the-art performance in KE. By providing contextual editing prompts enriched with new knowledge retrieved from the edit memory, ICE effectively guides LLMs to perform inference and generate answers aligned with the new knowledge. As part of this study, we use a ICE task with instructional editing prompts to evaluate LLMs' performance in instruction following.
B. Details of Data Constructing
One of our key objectives is to construct counterfactual contexts that simulate RAG scenarios under knowledge conflicts. This process involves two steps. The first step is to establish factual statements. We begin by collecting popular entities from Wikipedia and extracting factual triples associated with these entities from Wikidata. This ensures that the collected facts are widely recognized and likely to be well-represented in the parametric memory of LLMs due to pretraining. Using rule-based transformations, we convert these triples into factual statements, as illustrated by the templates provided in Table 7. Based on a chain of triples, multi-hop questions are generated using the following prompts and the examples in Table 8.
The second step involves generating contexts. Starting with the original factual triples, we expand the descriptions of entities to create enriched contexts that include irrelevant noise unrelated to the questions. The prompts used for generating these contexts are shown in following table:
The head and tail are derived from the collected factual triple (head, relation, tail), and the fact is constructed from this triple using a cloze template. Subsequently, the related entities in the context, along with their aliases and all associated morphological forms, are edited to reflect the counterfactuals. This editing process is achieved through pre-mapping the relationships and systematically replacing the corresponding entities.
C. Instruction Following Task
In our experiments, in addition to the Retrieval Following task on ConFiQA and Natural Questions, we specifically design an Instruction Following task to evaluate the model's faithfulness to user instructions as context. Specifically, we employ an in-context editing (ICE) task using the MQUAKE{\mathchoice{\text{MQ{\scriptsize U}A{\scriptsize KE}}}{\text{MQ{\scriptsize U}A{\scriptsize KE}}}{\text{MQ{\scriptscriptstyle U}A{\scriptscriptstyle KE}}}{\text{MQUAKE}}}MQUAKE dataset to assess this capability. This task provides contextual examples along with knowledge-editing instructions to test whether LLMs follow the provided context to answer questions. The few-shot prompting used for this task includes:
D. Implementation of Baselines
We follow the previous setup ([7]) and utilize two prompt-based baselines: the attributed prompt (Attr) and a combination of opinion-based and instruction-based prompts (O&I). The prompt templates are as follows:
In addition, we provide our own SFT baseline for comparison, which conducts end-to-end training using data in the format of (context + question, faithful response). Experimental results indicate that SFT fails to effectively improve the context-faithfulness performance of LLMs.
E. Convergence of Context-DPO Training
We configure the training with
batch_size=4 and gradient_accumulation_steps=8 and perform Direct Preference Optimization based on the constructed preference pairs. The convergence results are illustrated in Figure 7.F. Knowledge Token Capturing
The goal of the algorithm ([8]) is to identify parts of LLM outputs that distinguish newly acquired knowledge from the context (e.g., counterfactual information) from the parametric knowledge embedded in the LLM, rather than analyzing repetitive or meaningless outputs. For instance, consider an expected LLM output in an Instruction-Following scenario with injected context, such as “A: United States”, compared to the original parametric output without context injection, which might be “A: United Kingdom”. In this case, capturing “A:” is unnecessary as it lacks factual significance, and focusing on “United” is redundant, as it does not reflect the difference between the outputs. Instead, the focus should be on capturing tokens with distinct factual significance—those that can effectively differentiate between newly introduced contextual knowledge and the model’s inherent parametric knowledge. In this example, a token like “Kingdom” serves as a critical marker, clearly highlighting the key divergence between contextual information and the model’s existing knowledge. The pseudocode of the algorithm is shown in Algorithm 1. It captures the tokens with the highest probability of distinguishing new knowledge from parametric knowledge by matching the decoded tokens with their corresponding knowledge strings.
G. Examples of Data in ConFiQA
We provide example templates from the three sub-datasets of our ConFiQA: QA (Question-Answering), MR (Multi-hop Reasoning), and MC (Multi-Conflicts), which are shown in Table 11, Table 9, and Table 10, respectively. We provide a case study of LLAMA2-7B{\mathchoice{\text{LL{\scriptsize A}M{\scriptsize A}2-7B}}{\text{LL{\scriptsize A}M{\scriptsize A}2-7B}}{\text{LL{\scriptscriptstyle A}M{\scriptscriptstyle A}2-7B}}{\text{LLAMA2-7B}}}LLAMA2-7B, LLAMA3-8B{\mathchoice{\text{LL{\scriptsize A}M{\scriptsize A}3-8B}}{\text{LL{\scriptsize A}M{\scriptsize A}3-8B}}{\text{LL{\scriptscriptstyle A}M{\scriptscriptstyle A}3-8B}}{\text{LLAMA3-8B}}}LLAMA3-8B, MISTRAL-7B{\mathchoice{\text{M{\scriptsize ISTRAL}-7B}}{\text{M{\scriptsize ISTRAL}-7B}}{\text{M{\scriptscriptstyle ISTRAL}-7B}}{\text{MISTRAL-7B}}}MISTRAL-7B, and QWEN2-7B{\mathchoice{\text{Q{\scriptsize WEN}2-7B}}{\text{Q{\scriptsize WEN}2-7B}}{\text{Q{\scriptscriptstyle WEN}2-7B}}{\text{QWEN2-7B}}}QWEN2-7B on the QA task in Appendix H. Here, the green text represents the expected context-faithful output, while the red text represents the stubborn response.
H. Case Study
H.1 LLaMA2-7B
H.2 LLaMA3-8B
H.3 Mistral-7B
H.4 Qwen2-7B
References
[1] Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang. 2020. Realm: Retrieval-augmented language model pre-training. Preprint, arXiv:2002.08909.
[2] Yujia Qin, Shengding Hu, Yankai Lin, Weize Chen, Ning Ding, Ganqu Cui, Zheni Zeng, Yufei Huang, Chaojun Xiao, Chi Han, Yi Ren Fung, Yusheng Su, Huadong Wang, Cheng Qian, Runchu Tian, Kunlun Zhu, Shihao Liang, Xingyu Shen, Bokai Xu, Zhen Zhang, Yining Ye, Bowen Li, Ziwei Tang, Jing Yi, Yuzhang Zhu, Zhenning Dai, Lan Yan, Xin Cong, Yaxi Lu, Weilin Zhao, Yuxiang Huang, Junxi Yan, Xu Han, Xian Sun, Dahai Li, Jason Phang, Cheng Yang, Tongshuang Wu, Heng Ji, Zhiyuan Liu, and Maosong Sun. 2024. Tool learning with foundation models. Preprint, arXiv:2304.08354.
[5] Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurélien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023a. Llama: Open and efficient foundation language models. CoRR, abs/2302.13971.
[6] Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023b. Llama 2: Open foundation and fine-tuned chat models. Preprint, arXiv:2307.09288.
[7] Wenxuan Zhou, Sheng Zhang, Hoifung Poon, and Muhao Chen. 2023. Context-faithful prompting for large language models. arXiv preprint arXiv:2303.11315.
[8] Baolong Bi, Shenghua Liu, Yiwei Wang, Lingrui Mei, Junfeng Fang, Hongcheng Gao, Shiyu Ni, and Xueqi Cheng. 2024c. Is factuality enhancement a free lunch for llms? better factuality can lead to worse context-faithfulness. Authorea Preprints.
[9] Yifei Ming, Senthil Purushwalkam, Shrey Pandit, Zixuan Ke, Xuan-Phi Nguyen, Caiming Xiong, and Shafiq Joty. 2024. Faitheval: Can your language model stay faithful to context, even if" the moon is made of marshmallows". arXiv preprint arXiv:2410.03727.
[10] Fabio Petroni, Patrick Lewis, Aleksandra Piktus, Tim Rocktäschel, Yuxiang Wu, Alexander H Miller, and Sebastian Riedel. 2020. How context affects language models' factual predictions. arXiv preprint arXiv:2005.04611.
[11] Chenglei Si, Zhe Gan, Zhengyuan Yang, Shuohang Wang, Jianfeng Wang, Jordan Boyd-Graber, and Lijuan Wang. 2023. Prompting gpt-3 to be reliable. Preprint, arXiv:2210.09150.
[12] Jian Xie, Kai Zhang, Jiangjie Chen, Renze Lou, and Yu Su. 2024. Adaptive chameleon or stubborn sloth: Revealing the behavior of large language models in knowledge conflicts. Preprint, arXiv:2305.13300.
[13] Weijia Shi, Xiaochuang Han, Mike Lewis, Yulia Tsvetkov, Luke Zettlemoyer, and Scott Wen-tau Yih. 2023. Trusting your evidence: Hallucinate less with context-aware decoding. arXiv preprint arXiv:2305.14739.
[14] Baolong Bi, Shenghua Liu, Yiwei Wang, Lingrui Mei, Hongcheng Gao, Yilong Xu, and Xueqi Cheng. 2024e. Adaptive token biaser: Knowledge editing via biasing key entities. arXiv preprint arXiv:2406.12468.
[15] Yang Liu, Yuanshun Yao, Jean-Francois Ton, Xiaoying Zhang, Ruocheng Guo Hao Cheng, Yegor Klochkov, Muhammad Faaiz Taufiq, and Hang Li. 2023. Trustworthy llms: A survey and guideline for evaluating large language models' alignment. arXiv preprint arXiv:2308.05374.
[16] Tianhao Shen, Renren Jin, Yufei Huang, Chuang Liu, Weilong Dong, Zishan Guo, Xinwei Wu, Yan Liu, and Deyi Xiong. 2023. Large language model alignment: A survey. arXiv preprint arXiv:2309.15025.
[17] Katherine Tian, Eric Mitchell, Huaxiu Yao, Christopher D Manning, and Chelsea Finn. 2023. Fine-tuning language models for factuality. arXiv preprint arXiv:2311.08401.
[18] Bochuan Cao, Yuanpu Cao, Lu Lin, and Jinghui Chen. 2023. Defending against alignment-breaking attacks via robustly aligned llm. arXiv preprint arXiv:2309.14348.
[19] Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2024. Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems, 36.
[20] Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al. 2019. Natural questions: a benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:453–466.
[21] Zexuan Zhong, Zhengxuan Wu, Christopher D Manning, Christopher Potts, and Danqi Chen. 2023. Mquake: Assessing knowledge editing in language models via multi-hop questions. arXiv preprint arXiv:2305.14795.
[22] Stephanie Lin, Jacob Hilton, and Owain Evans. 2021. Truthfulqa: Measuring how models mimic human falsehoods. arXiv preprint arXiv:2109.07958.
[23] Denny Vrandečić and Markus Krötzsch. 2014. Wikidata: a free collaborative knowledgebase. Communications of the ACM, 57(10):78–85.
[24] Shayne Longpre, Kartik Perisetla, Anthony Chen, Nikhil Ramesh, Chris DuBois, and Sameer Singh. 2021. Entity-based knowledge conflicts in question answering. arXiv preprint arXiv:2109.05052.
[25] Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824–24837.
[26] Ce Zheng, Lei Li, Qingxiu Dong, Yuxuan Fan, Zhiyong Wu, Jingjing Xu, and Baobao Chang. 2023. Can we edit factual knowledge by in-context learning? arXiv preprint arXiv:2305.12740.
[27] Jean Kaddour, Joshua Harris, Maximilian Mozes, Herbie Bradley, Roberta Raileanu, and Robert McHardy. 2023. Challenges and applications of large language models. arXiv preprint arXiv:2307.10169.
[28] S. M Towhidul Islam Tonmoy, S M Mehedi Zaman, Vinija Jain, Anku Rani, Vipula Rawte, Aman Chadha, and Amitava Das. 2024. A comprehensive survey of hallucination mitigation techniques in large language models. Preprint, arXiv:2401.01313.
[29] Cunxiang Wang, Xiaoze Liu, Yuanhao Yue, Xiangru Tang, Tianhang Zhang, Cheng Jiayang, Yunzhi Yao, Wenyang Gao, Xuming Hu, Zehan Qi, et al. 2023. Survey on factuality in large language models: Knowledge, retrieval and domain-specificity. arXiv preprint arXiv:2310.07521.
[30] Lingrui Mei, Shenghua Liu, Yiwei Wang, Baolong Bi, Jiayi Mao, and Xueqi Cheng. 2024. " not aligned" is not" malicious": Being careful about hallucinations of large language models' jailbreak. arXiv preprint arXiv:2406.11668.
[31] Anisha Gunjal, Jihan Yin, and Erhan Bas. 2024. Detecting and preventing hallucinations in large vision language models. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 18135–18143.
[32] Jiaxin Zhang, Zhongzhi Li, Mingliang Zhang, Fei Yin, Chenglin Liu, and Yashar Moshfeghi. 2024. Geoeval: benchmark for evaluating llms and multi-modal models on geometry problem-solving. arXiv preprint arXiv:2402.10104.
[33] Fang Liu, Yang Liu, Lin Shi, Houkun Huang, Ruifeng Wang, Zhen Yang, and Li Zhang. 2024. Exploring and evaluating hallucinations in llm-powered code generation. arXiv preprint arXiv:2404.00971.
[34] Zhong-Zhi Li, Ming-Liang Zhang, Fei Yin, Zhi-Long Ji, Jin-Feng Bai, Zhen-Ru Pan, Fan-Hu Zeng, Jian Xu, Jia-Xin Zhang, and Cheng-Lin Liu. 2024b. Cmmath: A chinese multi-modal math skill evaluation benchmark for foundation models. arXiv preprint arXiv:2407.12023.
[35] Baolong Bi, Shenghua Liu, Yiwei Wang, Lingrui Mei, and Xueqi Cheng. 2024b. Lpnl: Scalable link prediction with large language models. In Findings of the Association for Computational Linguistics ACL 2024, pages 3615–3625.
[36] Linmei Hu, Zeyi Liu, Ziwang Zhao, Lei Hou, Liqiang Nie, and Juanzi Li. 2023. A survey of knowledge enhanced pre-trained language models. IEEE Transactions on Knowledge and Data Engineering.
[37] Shirui Pan, Linhao Luo, Yufei Wang, Chen Chen, Jiapu Wang, and Xindong Wu. 2024. Unifying large language models and knowledge graphs: A roadmap. IEEE Transactions on Knowledge and Data Engineering.
[38] Yuheng Huang, Jiayang Song, Zhijie Wang, Huaming Chen, and Lei Ma. 2023. Look before you leap: An exploratory study of uncertainty measurement for large language models. arXiv preprint arXiv:2307.10236.
[39] Zhangyin Feng, Xiaocheng Feng, Dezhi Zhao, Maojin Yang, and Bing Qin. 2024. Retrieval-generation synergy augmented large language models. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 11661–11665. IEEE.
[40] Rui Yang, Haoran Liu, Qingcheng Zeng, Yu He Ke, Wanxin Li, Lechao Cheng, Qingyu Chen, James Caverlee, Yutaka Matsuo, and Irene Li. 2024. Kg-rank: Enhancing large language models for medical qa with knowledge graphs and ranking techniques. arXiv preprint arXiv:2403.05881.
[41] Yunzhi Yao, Peng Wang, Bozhong Tian, Siyuan Cheng, Zhoubo Li, Shumin Deng, Huajun Chen, and Ningyu Zhang. 2023b. Editing large language models: Problems, methods, and opportunities. arXiv preprint arXiv:2305.13172.
[42] Sophie Forgan. 2005. Building the museum: Knowledge, conflict, and the power of place. Isis, 96(4):572–585.
[43] Rongwu Xu, Zehan Qi, Zhijiang Guo, Cunxiang Wang, Hongru Wang, Yue Zhang, and Wei Xu. 2024. Knowledge conflicts for llms: A survey. arXiv preprint arXiv:2403.08319.
[44] Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, Xu Jiang, Karl Cobbe, Tyna Eloundou, Gretchen Krueger, Kevin Button, Matthew Knight, Benjamin Chess, and John Schulman. 2022. Webgpt: Browser-assisted question-answering with human feedback. Preprint, arXiv:2112.09332.
[45] Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2023a. React: Synergizing reasoning and acting in language models. Preprint, arXiv:2210.03629.
[46] Gautier Izacard and Edouard Grave. 2021. Leveraging passage retrieval with generative models for open domain question answering. Preprint, arXiv:2007.01282.
[47] Zexuan Zhong, Tao Lei, and Danqi Chen. 2022. Training language models with memory augmentation. Preprint, arXiv:2205.12674.
[48] Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in Neural Information Processing Systems, 33:9459–9474.
[49] Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, and Haofen Wang. 2023. Retrieval-augmented generation for large language models: A survey. arXiv preprint arXiv:2312.10997.
[50] Zhengbao Jiang, Frank F Xu, Luyu Gao, Zhiqing Sun, Qian Liu, Jane Dwivedi-Yu, Yiming Yang, Jamie Callan, and Graham Neubig. 2023. Active retrieval augmented generation. arXiv preprint arXiv:2305.06983.
[51] Hung-Ting Chen, Michael JQ Zhang, and Eunsol Choi. 2022. Rich knowledge sources bring complex knowledge conflicts: Recalibrating models to reflect conflicting evidence. arXiv preprint arXiv:2210.13701.
[52] Yuepei Li, Kang Zhou, Qiao Qiao, Bach Nguyen, Qing Wang, and Qi Li. 2024a. Investigating context-faithfulness in large language models: The roles of memory strength and evidence style. arXiv preprint arXiv:2409.10955.
[53] Chen Zhu, Ankit Singh Rawat, Manzil Zaheer, Srinadh Bhojanapalli, Daliang Li, Felix Yu, and Sanjiv Kumar. 2020. Modifying memories in transformer models. arXiv preprint arXiv:2012.00363.
[54] Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022a. Locating and editing factual associations in gpt. Advances in Neural Information Processing Systems, 35:17359–17372.
[55] Kevin Meng, Arnab Sen Sharma, Alex Andonian, Yonatan Belinkov, and David Bau. 2022b. Mass-editing memory in a transformer. arXiv preprint arXiv:2210.07229.
[56] Baixiang Huang, Canyu Chen, Xiongxiao Xu, Ali Payani, and Kai Shu. 2024. Can knowledge editing really correct hallucinations? arXiv preprint arXiv:2410.16251.
[57] Aman Madaan, Niket Tandon, Peter Clark, and Yiming Yang. 2022. Memory-assisted prompt editing to improve gpt-3 after deployment. arXiv preprint arXiv:2201.06009.
[58] Roi Cohen, Eden Biran, Ori Yoran, Amir Globerson, and Mor Geva. 2024. Evaluating the ripple effects of knowledge editing in language models. Transactions of the Association for Computational Linguistics, 12:283–298.
[59] Baolong Bi, Shenghua Liu, Lingrui Mei, Yiwei Wang, Pengliang Ji, and Xueqi Cheng. 2024a. Decoding by contrasting knowledge: Enhancing llms' confidence on edited facts. Preprint, arXiv:2405.11613.
[60] Baolong Bi, Shenghua Liu, Yiwei Wang, Lingrui Mei, Hongcheng Gao, Junfeng Fang, and Xueqi Cheng. 2024d. Struedit: Structured outputs enable the fast and accurate knowledge editing for large language models. arXiv preprint arXiv:2409.10132.















