Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?

Zorik GekhmanGal YonaRoee AharoniMatan EyalAmir FederRoi ReichartJonathan Herzig

article2024EMNLP318 citations

Reveals that fine-tuning language models on unfamiliar factual data causes them to fit new information much slower than known facts while linearly increasing their rate of hallucinations, indicating that fine-tuning should focus on formatting existing knowledge rather than teaching new facts.

Listen

When tailoring large language models to specific tasks through supervised fine-tuning, training datasets frequently contain factual information that was never part of the model's original pre-training. An ongoing concern across artificial intelligence development is that exposing models to unfamiliar facts during fine-tuning might teach them to fabricate information—commonly known as hallucination—rather than grounding their outputs in established knowledge. Addressing this issue is critical today as organizations increasingly rely on fine-tuning to deploy models in high-stakes operational environments.

The article evaluates how supervised fine-tuning on new factual knowledge influences a model's ability to utilize its pre-existing knowledge and estimates the extent to which it drives hallucinations. Specifically, the authors assess whether models can genuinely acquire new factual information during alignment or whether this practice primarily degrades factual accuracy.

To investigate this dynamic, the authors established a controlled closed-book question-answering framework using the PaLM 2-S language model and a standardized knowledge base of entity relations. They developed a sampling-based categorization method to partition data into known and unknown examples based on model generation probabilities before fine-tuning, further stratifying known facts by confidence levels. The researchers then conducted systematic fine-tuning experiments across various durations, tracking training dynamics and performance across both in-distribution and out-of-distribution test sets.

The study yielded several key findings regarding model behavior. First, language models fit unknown examples significantly slower during training than known examples, indicating that models struggle to integrate new factual knowledge during fine-tuning. Second, fitting unknown training examples exhibits a strong negative linear correlation with test accuracy, reducing accuracy by roughly eight percentage points in-distribution, whereas fitting known examples improves accuracy by roughly seven percentage points. Third, introducing unknown facts actively induces hallucinations regarding facts the model previously knew, an effect that generalizes even to unrelated, out-of-distribution factual domains. Fourth, fine-tuning exclusively on moderately known examples yielded the best overall test performance (43.6% exact match), outperforming datasets composed solely of highly known facts (40.5%) by enabling better utilization of uncertain pre-existing knowledge.

These findings demonstrate that supervised fine-tuning is an ineffective mechanism for injecting new knowledge into large language models and poses measurable risks to output factuality. For practitioners, attempting to teach new facts during fine-tuning increases operational risk, causes severe overfitting, and damages the model's reliability on knowledge it already possessed. Instead, fine-tuning primarily acts as an alignment tool that surfaces and organizes pre-existing parametric knowledge acquired during pre-training.

To minimize hallucinations and preserve performance, technical leaders should align fine-tuning datasets with a model's existing knowledge profile. Practitioners should implement early stopping on validation sets or proactively filter out unknown factual examples from training pipelines. Alternatively, teams can teach models to abstain by relabeling unknown training instances with uncertainty expressions such as "I don't know," which preserved a 61.8% accuracy on answered questions across training epochs. For substantial knowledge updates, organizations should rely on continuous pre-training or external retrieval mechanisms rather than supervised fine-tuning.

These conclusions are supported with high confidence within the evaluated closed-book question-answering benchmark on PaLM 2-S. However, leaders should note that the study focused on a single model architecture and short-form factual triplets. Further validation is needed to determine how these dynamics scale across diverse model families, parameter-efficient adaptation methods like low-rank adaptation, and long-form open-ended text generation tasks.

arXiv: 2405.05904
Cover for Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?

Abstract

When large language models are aligned via supervised fine-tuning, they may encounter new factual information that was not acquired through pre-training. It is often conjectured that this can teach the model the behavior of hallucinating factually incorrect responses, as the model is trained to generate facts that are not grounded in its pre-existing knowledge. In this work, we study the impact of such exposure to new knowledge on the capability of the fine-tuned model to utilize its pre-existing knowledge. To this end, we design a controlled setup, focused on closed-book QA, where we vary the proportion of the fine-tuning examples that introduce new knowledge. We demonstrate that large language models struggle to acquire new factual knowledge through fine-tuning, as fine-tuning examples that introduce new knowledge are learned significantly slower than those consistent with the model's knowledge. However, we also find that as the examples with new knowledge are eventually learned, they linearly increase the model's tendency to hallucinate. Taken together, our results highlight the risk in introducing new factual knowledge through fine-tuning, and support the view that large language models mostly acquire factual knowledge through pre-training, whereas fine-tuning teaches them to use it more efficiently.

Table of Contents

  • 1 Introduction
  • 2 Study Setup
  • 3 Quantifying Knowledge in LLMs
  • 4 How Harmful are Unknown Examples?
  • 4.1 Higher Unknown Ratio is Proportional to Performance Degradation
  • 4.2 Unknown Examples: Harmful or Neutral?
  • 4.3 Unknown Examples are Fitted Slower than Known Examples
  • 4.4 The Influence of Unknown vs Known on Accuracy: A Linear Model Perspective
  • 4.5 Generalization to New Relations
  • 5 Understanding Knowledge Types: Their Value and Impact
  • 6 SliCK Knowledge Categories Analysis
  • 7 Fine-tuning to Abstain on Unknown Examples
  • 8 Discussion
  • 9 Related Work
  • 10 Conclusion
  • 11 Limitations
  • 12 Acknowledgments
  • References
  • A Data Preprocessing
  • B Hallucinations in the Context of our Study
  • B.1 Hallucinations w.r.t. the Pre-existing Knowledge.
  • B.2 Test performance as Proxy for Hallucinations.
  • C P Correct Approximation
  • D Data Annotation
  • E Fine-tuning Details
  • F The Case for Avoiding Fake Facts
  • G Train Accuracy on Different Known Categories
  • H Linear Model
  • I Out-of-distribution (OOD) Evaluation
  • J Statistic Significance Tests
  • K The P(True) Case Study

Knowls

  1. Knowl 1 — SliCK categorizes facts by sampling-based model knowledge

    model/method

    SliCK (Sampling-based Categorization of Knowledge) classifies a question–answer pair (q,a)(q,a) according to how often a pretrained language model MM produces the gold answer aa when prompted with randomly selected few-shot examples. The continuous quantity Pcorrect(q,a;M,T)P_{\mathrm{correct}}(q,a;M,T) is estimated as the fraction of correct predictions at decoding temperature TT, using Nex=10N_{\mathrm{ex}}=10 distinct four-shot prompts for the same relation as qq. For each prompt, the study obtains one greedy prediction with T=0T=0 and 16 sampled predictions with T=0.5T=0.5; sampling uses the top 40 tokens.

    The four SliCK categories are defined from these empirical probabilities:

    • HighlyKnown: Pcorrect(q,a;M,0)=1P_{\mathrm{correct}}(q,a;M,0)=1; every greedy prediction is correct.
    • MaybeKnown: 0<Pcorrect(q,a;M,0)<10<P_{\mathrm{correct}}(q,a;M,0)<1; some but not all greedy predictions are correct.
    • WeaklyKnown: Pcorrect(q,a;M,0)=0P_{\mathrm{correct}}(q,a;M,0)=0 but Pcorrect(q,a;M,T>0)>0P_{\mathrm{correct}}(q,a;M,T>0)>0; greedy decoding never succeeds, but temperature sampling sometimes does.
    • Unknown: Pcorrect(q,a;M,T)=0P_{\mathrm{correct}}(q,a;M,T)=0 for both greedy and sampled predictions; the model never produces the gold answer in the evaluation samples.

    Thus, Known examples comprise HighlyKnown, MaybeKnown, and WeaklyKnown examples, while Unknown examples are those for which the base model never generates the gold answer. Exact match against the gold answer is used to decide whether a prediction is correct.

  2. Knowl 2 — Controlled closed-book fine-tuning study

    experimental setup

    The study represents factual knowledge as subject–relation–object triples converted into closed-book question–answer pairs, D={(qi,ai)}i=1ND=\{(q_i,a_i)\}_{i=1}^{N}, where qiq_i is a knowledge-seeking question and aia_i is its gold answer. The base model is the PaLM 2-S base model, and MDM_D denotes the model obtained by full-parameter fine-tuning of MM on dataset DD. Exact-match accuracy is the primary evaluation metric.

    The main experiments use ENTITYQUESTIONS, retaining 12 randomly selected relations for training, development, and in-distribution testing. Seven additional relations, with overlapping or closely related relations removed, form an out-of-distribution test set. Subjects and objects are made disjoint between training and test sets, and examples with multiple correct answers are removed.

    To isolate the effect of new factual knowledge, the researchers construct fixed-size fine-tuning datasets with varying proportions of SliCK Unknown examples and the complementary proportion of Known examples, while preserving the relation distribution. The single-category experiments use exactly ∣D∣=6142|D|=6142 examples for every dataset variant, allowing variants containing only HighlyKnown, MaybeKnown, WeaklyKnown, Unknown, or the natural category mixture. Models are trained for up to 50 epochs with learning rate 10−510^{-5}, batch size 128, and dropout 0.05. EARLY_STOP selects the epoch with the highest development accuracy, usually after 5–10 epochs; CONVERGENCE denotes 50 epochs, by which time the model fits the training set completely.

  3. Knowl 3 — Operational definition of hallucination relative to pre-existing knowledge

    definition

    In this closed-book question-answering setting, a hallucination with respect to pre-existing knowledge is a pair (q,a)(q,a) that the pretrained model MM knows according to SliCK, but for which the fine-tuned model MDM_D fails to produce the correct answer. This definition includes both cases in which fine-tuning causes the model to retain the fact but generate an incorrect answer and cases in which the relevant fact is forgotten; the experiments do not distinguish between these mechanisms.

    Because the training and test sets are disjoint, test questions whose answers are Unknown to MM remain unavailable to both compared fine-tuned models. Consequently, differences in test accuracy between fine-tuning conditions are interpreted primarily as differences in failures on test questions whose answers were already known to MM, and therefore as differences in hallucination rates relative to pre-existing knowledge.

  4. Knowl 4 — Unknown examples become harmful through overfitting

    empirical result

    Increasing the proportion of SliCK Unknown examples in a fixed-size fine-tuning dataset consistently lowers test accuracy, regardless of the training duration. Early stopping usually gives the highest accuracy, whereas training to 50-epoch convergence gives the lowest accuracy; the performance gap between early stopping and longer training grows monotonically as the Unknown proportion increases.

    To separate the effect of Unknown examples from the effect of having fewer Known examples, the study compares each dataset DD with DKnownD_{\mathrm{Known}}, formed by deleting all Unknown examples from DD. At EARLY_STOP, DD and DKnownD_{\mathrm{Known}} have nearly identical performance, so most Unknown examples have not yet caused measurable damage. At CONVERGENCE, DD performs worse than DKnownD_{\mathrm{Known}}, and the gap increases with the Unknown proportion. Filtering Unknown examples therefore substantially reduces overfitting without sacrificing early-stopping performance, although it also removes the opportunity to learn from those examples.

  5. Knowl 5 — Unknown facts are learned more slowly than Known facts

    empirical result

    During supervised fine-tuning, the model fits SliCK Known examples substantially faster than SliCK Unknown examples. At the development-optimal EARLY_STOP point, the model has learned most of the Known training subset but only a small fraction of the Unknown subset. This pattern holds across the different mixtures of Known and Unknown examples.

    The slower fitting of Unknown examples means that their negative effect appears mainly in later training stages: early training primarily teaches the model to expose or utilize information already present in its parameters, while continued training eventually fits examples that introduce facts absent from the pretrained model. The authors interpret this as evidence that the model struggles to integrate new factual knowledge through ordinary supervised fine-tuning, rather than acquiring such knowledge as readily as it uses pre-existing knowledge.

  6. Knowl 6 — Fitting Unknown examples has an approximately linear negative effect on accuracy

    theoretical result

    The study models test accuracy as a linear function of the fractions of Known and Unknown fine-tuning examples that the model has already fitted. Let DD be a fine-tuning dataset, NKnownN_{\mathrm{Known}} the number of its Known examples that the model answers correctly, NUnknownN_{\mathrm{Unknown}} the number of its Unknown examples that it answers correctly, and ∣D∣|D| the dataset size. The fitted model is

    Accuracy=β0+βKnownNKnown∣D∣+βUnknownNUnknown∣D∣.\mathrm{Accuracy}=\beta_0+\beta_{\mathrm{Known}}\frac{N_{\mathrm{Known}}}{|D|}+\beta_{\mathrm{Unknown}}\frac{N_{\mathrm{Unknown}}}{|D|}.

    The regression is estimated from accuracy and training-fit measurements collected after every epoch across all dataset variants; it is valid in the bounded region 0≤NKnown,NUnknown≤∣D∣0\leq N_{\mathrm{Known}},N_{\mathrm{Unknown}}\leq |D|. For in-distribution test accuracy, the estimated coefficients are β0=36.9\beta_0=36.9, βKnown=7.3\beta_{\mathrm{Known}}=7.3, and βUnknown=−8.3\beta_{\mathrm{Unknown}}=-8.3, with R2=0.86R^2=0.86. For out-of-distribution accuracy, they are β0=36.2\beta_0=36.2, βKnown=3.2\beta_{\mathrm{Known}}=3.2, and βUnknown=−3.0\beta_{\mathrm{Unknown}}=-3.0, with R2=0.95R^2=0.95. Thus, fitting Known examples is positively associated with accuracy, fitting Unknown examples is negatively associated with accuracy, and the magnitudes of the two effects are approximately equal.

  7. Knowl 7 — The hallucination effect transfers to unseen relations

    empirical result

    The relationship between Unknown fine-tuning examples and hallucination is not restricted to questions using the same relations as the fine-tuning data. On a test set containing seven relations absent from fine-tuning, increasing the Unknown proportion lowers out-of-distribution accuracy, and Unknown examples become harmful mainly after the model has fitted them. The same linear model yields a positive coefficient for fitted Known examples, a negative coefficient for fitted Unknown examples, and approximately equal absolute coefficient magnitudes, with R2=0.95R^2=0.95.

    The out-of-distribution accuracy drop is smaller than the in-distribution drop—up to roughly 6 percentage points versus up to roughly 14 percentage points in-distribution—but its direction is the same. For example, fine-tuning on Unknown questions about where entities are located can increase errors on seemingly unrelated unseen questions about who founded an entity. This cross-relation transfer is consistent with the interpretation that fine-tuning teaches a broader behavior of answering without grounding in pre-existing knowledge, rather than merely damaging individual factual associations.

  8. Knowl 8 — MaybeKnown examples give the best utilization of pre-existing knowledge

    data/table

    The study compares fine-tuning datasets containing only one SliCK category, plus a dataset with the natural category mixture. The dataset size is fixed at 6142 examples. Values are exact-match accuracies in percent; “Full” is the complete test set, and Hkn, Mkn, Wkn, and Unk are test subsets categorized respectively as HighlyKnown, MaybeKnown, WeaklyKnown, and Unknown by the pretrained model.

    Could not parse LaTeX table

    The MaybeKnown-only dataset gives the highest overall early-stopping accuracy, 43.6%, and the highest accuracy on MaybeKnown test questions, 69.9%, while maintaining almost the same HighlyKnown accuracy as the HighlyKnown-only dataset. HighlyKnown-only fine-tuning is therefore not optimal despite its very high performance on HighlyKnown questions. WeaklyKnown and Unknown mixtures are substantially more vulnerable to convergence-time degradation, particularly on HighlyKnown and MaybeKnown test questions. The natural mixture performs almost as well as MaybeKnown-only at early stopping but degrades more after convergence, indicating that WeaklyKnown and Unknown examples increase overfitting risk.

  9. Knowl 9 — Relabeling Unknown examples as abstentions mitigates overfitting

    empirical result

    The study tests a mitigation in which every Unknown fine-tuning example is relabeled with the response “I don’t know” instead of its factual answer. The comparison uses a dataset with 50% Known and 50% Unknown examples. Accuracy is computed only on test questions that the fine-tuned model answers rather than abstains from.

    Could not parse LaTeX table

    Relabeling improves accuracy conditional on answering and removes the convergence-time accuracy decline observed with the original labels. The cost is reduced coverage: the model answers only 58.7% of test questions at early stopping and 55.6% after convergence. This result suggests that teaching the model to abstain on examples that introduce new facts can preserve factuality more effectively than training it to emit those facts.

  10. Knowl 10 — SliCK identifies genuinely difficult facts better than a single confidence score

    empirical result

    The paper validates SliCK by examining test questions classified as Unknown and comparing it with a P(True)P(\mathrm{True}) approach that asks the model whether a proposed gold answer is true. After fine-tuning, SliCK-Unknown test questions have extremely low accuracy—at most 3.2% across the reported fine-tuning conditions—indicating that most were genuinely unavailable to the pretrained model.

    For a comparable fraction of test examples labeled Unknown, thresholding P(True)P(\mathrm{True}) produces substantially higher post-fine-tuning accuracy than SliCK. The study also finds that estimating PcorrectP_{\mathrm{correct}} from fewer than 10 distinct four-shot prompts yields higher accuracy on the resulting Unknown subset; using multiple prompt samples is therefore important because the selected exemplars materially affect the model’s predictions. The authors treat the P(True)P(\mathrm{True}) comparison as a preliminary case study rather than a comprehensive benchmark.

  11. Knowl 11 — Scope is limited to one model and closed-book factual QA

    limitation

    The empirical conclusions are based on a single model, PaLM 2-S, so their stability across model families, scales, and fine-tuning methods is unknown. The experiments use closed-book entity-centric question answering, where exact-match evaluation and disjoint train/test facts make hallucinations measurable; applying the filtering or SliCK procedure to long-form generation would require new methods for comparing sampled answers with gold facts and more difficult evaluation.

    The study also does not test mixtures containing additional diverse instruction-following tasks, because such data could introduce new factual knowledge in an uncontrolled way. Finally, the SliCK annotation is computationally expensive: each example requires 170 model inference steps, leading to more than 15 million inference steps for the full dataset. The results therefore do not establish whether the same effects hold under parameter-efficient fine-tuning, continual pre-training, or substantially broader instruction-tuning mixtures.

Coverage note — Detailed relation-by-relation dataset counts, preprocessing exclusions, fake-fact alternatives, exact statistical-test procedures, and auxiliary learning-rate/error analyses were omitted because they support the main method and findings without constituting separate load-bearing contributions.

References

  1. 1.Badr AlKhamissi, Millicent Li, Asli Celikyilmaz, Mona Diab, and Marjan Ghazvininejad. 2022. A review on language models as knowledge bases. arXiv preprint arXiv:2204.06031.
  2. 2.Rohan Anil, Andrew M Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al. 2023. Palm 2 technical report. arXiv preprint arXiv:2305.10403.
  3. 3.Amos Azaria and Tom Mitchell. 2023. The internal state of an llm knows when its lying. arXiv preprint arXiv:2304.13734.
  4. 4.Dan Biderman, Jacob Portes, Jose Javier Gonzalez Ortiz, Mansheej Paul, Philip Greengard, Connor Jennings, Daniel King, Sam Havens, Vitaliy Chiley, Jonathan Frankle, Cody Blakeney, and John Patrick Cunningham. 2024. LoRA learns less and forgets less. Transactions on Machine Learning Research. Featured Certification.
  5. 5.Collin Burns, Haotian Ye, Dan Klein, and Jacob Steinhardt. 2022. Discovering latent knowledge in language models without supervision. arXiv preprint arXiv:2212.03827.
  6. 6.I Chern, Steffi Chern, Shiqi Chen, Weizhe Yuan, Kehua Feng, Chunting Zhou, Junxian He, Graham Neubig, Pengfei Liu, et al. 2023. Factool: Factuality detection in generative ai–a tool augmented framework for multi-task and multi-domain scenarios. arXiv preprint arXiv:2307.13528.
  7. 7.Roi Cohen, Mor Geva, Jonathan Berant, and Amir Globerson. 2023. Crawling the internal knowledge-base of language models. In Findings of the Association for Computational Linguistics: EACL 2023, pages 1856–1869, Dubrovnik, Croatia. Association for Computational Linguistics.
  8. 8.Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783.
  9. 9.Leo Gao. 2021. Behavior cloning is miscalibrated. AI Alignment Forum.
  10. 10.Zorik Gekhman, Jonathan Herzig, Roee Aharoni, Chen Elkind, and Idan Szpektor. 2023. TrueTeacher: Learning factual consistency evaluation with large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 2053–2070, Singapore. Association for Computational Linguistics.
  11. 11.Gaurav Rohit Ghosal, Tatsunori Hashimoto, and Aditi Raghunathan. 2024. Understanding finetuning for factual knowledge extraction. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024. OpenReview.net.
  12. 12.Yoav Goldberg. 2023. Reinforcement learning for language models.
  13. 13.Arnav Gudibande, Eric Wallace, Charlie Snell, Xinyang Geng, Hao Liu, Pieter Abbeel, Sergey Levine, and Dawn Song. 2023. The false promise of imitating proprietary llms. arXiv preprint arXiv:2305.15717.
  14. 14.Zeyu Han, Chao Gao, Jinyang Liu, Sai Qian Zhang, et al. 2024. Parameter-efficient fine-tuning for large models: A comprehensive survey. arXiv preprint arXiv:2403.14608.
  15. 15.Or Honovich, Roee Aharoni, Jonathan Herzig, Hagai Taitelbaum, Doron Kukliansy, Vered Cohen, Thomas Scialom, Idan Szpektor, Avinatan Hassidim, and Yossi Matias. 2022. TRUE: Re-evaluating factual consistency evaluation. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3905–3920, Seattle, United States. Association for Computational Linguistics.
  16. 16.Or Honovich, Leshem Choshen, Roee Aharoni, Ella Neeman, Idan Szpektor, and Omri Abend. 2021. q2\text{q}^2: Evaluating factual consistency in knowledge-grounded dialogues via question generation and question answering. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7856–7870, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  17. 17.Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. Lora: Low-rank adaptation of large language models. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net.
  18. 18.Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, et al. 2023. A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. arXiv preprint arXiv:2311.05232.
  19. 19.Adam Ibrahim, Benjamin Thérien, Kshitij Gupta, Mats Leon Richter, Quentin Gregory Anthony, Eugene Belilovsky, Timothée Lesort, and Irina Rish. 2024. Simple and scalable strategies to continually pre-train large language models. Transactions on Machine Learning Research.
  20. 20.Zhengbao Jiang, Zhiqing Sun, Weijia Shi, Pedro Rodriguez, Chunting Zhou, Graham Neubig, Xi Lin, Wen-tau Yih, and Srini Iyer. 2024. Instruction-tuned language models are better knowledge learners. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5421–5434, Bangkok, Thailand. Association for Computational Linguistics.
  21. 21.Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, et al. 2022. Language models (mostly) know what they know. arXiv preprint arXiv:2207.05221.
  22. 22.Ehsan Kamalloo, Nouha Dziri, Charles L. A. Clarke, and Davood Rafiei. 2023. Evaluating open-domain question answering in the era of large language models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023, pages 5591–5606. Association for Computational Linguistics.
  23. 23.Katie Kang, Eric Wallace, Claire Tomlin, Aviral Kumar, and Sergey Levine. 2024. Unfamiliar finetuning examples control how language models hallucinate. arXiv preprint arXiv:2403.05612.
  24. 24.Wojciech Kryscinski, Bryan McCann, Caiming Xiong, and Richard Socher. 2020. Evaluating the factual consistency of abstractive text summarization. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9332–9346, Online. Association for Computational Linguistics.
  25. 25.Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar. 2023. Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation. arXiv preprint arXiv:2302.09664.
  26. 26.Philippe Laban, Tobias Schnabel, Paul N. Bennett, and Marti A. Hearst. 2022. SummaC: Re-visiting NLI-based models for inconsistency detection in summarization. Transactions of the Association for Computational Linguistics, 10:163–177.
  27. 27.Yoonsang Lee, Pranav Atreya, Xi Ye, and Eunsol Choi. 2023. Crafting in-context examples according to lms’ parametric knowledge. arXiv preprint arXiv:2311.09579.
  28. 28.Bill Yuchen Lin, Abhilasha Ravichander, Ximing Lu, Nouha Dziri, Melanie Sclar, Khyathi Chandu, Chandra Bhagavatula, and Yejin Choi. 2023. The unlocking spell on base llms: Rethinking alignment via in-context learning. ArXiv preprint.
  29. 29.Sheng-Chieh Lin, Luyu Gao, Barlas Oguz, Wenhan Xiong, Jimmy Lin, Wen-tau Yih, and Xilun Chen. 2024. Flame: Factuality-aware alignment for large language models. arXiv preprint arXiv:2405.01525.
  30. 30.Qing Lyu, Kumar Shridhar, Chaitanya Malaviya, Li Zhang, Yanai Elazar, Niket Tandon, Marianna Apidianaki, Mrinmaya Sachan, and Chris Callison-Burch. 2024. Calibrating large language models with sample consistency. arXiv preprint arXiv:2402.13904.
  31. 31.Potsawee Manakul, Adian Liusie, and Mark JF Gales. 2023. Selfcheckgpt: Zero-resource black-box hallucination detection for generative large language models. arXiv preprint arXiv:2303.08896.
  32. 32.Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022. Locating and editing factual associations in GPT. In Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 - December 9, 2022.
  33. 33.Kevin Meng, Arnab Sen Sharma, Alex J. Andonian, Yonatan Belinkov, and David Bau. 2023. Mass-editing memory in a transformer. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net.
  34. 34.Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2023. FActScore: Fine-grained atomic evaluation of factual precision in long form text generation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 12076–12100, Singapore. Association for Computational Linguistics.
  35. 35.Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi. 2022. Cross-task generalization via natural language crowdsourcing instructions. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3470–3487, Dublin, Ireland. Association for Computational Linguistics.
  36. 36.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 - December 9, 2022.
  37. 37.Jupinder Parmar, Sanjev Satheesh, Mostofa Patwary, Mohammad Shoeybi, and Bryan Catanzaro. 2024. Reuse, don’t retrain: A recipe for continued pretraining of language models. arXiv preprint arXiv:2407.07263.
  38. 38.Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019. Language models as knowledge bases? In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2463–2473, Hong Kong, China. Association for Computational Linguistics.
  39. 39.Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2024. Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems, 36.
  40. 40.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. SQuAD: 100,000+ questions for machine comprehension of text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 2383–2392, Austin, Texas. Association for Computational Linguistics.
  41. 41.Mengjie Ren, Boxi Cao, Hongyu Lin, Cao Liu, Xianpei Han, Ke Zeng, Wan Guanglu, Xunliang Cai, and Le Sun. 2024. Learning or self-aligning? rethinking instruction fine-tuning. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 6090–6105, Bangkok, Thailand. Association for Computational Linguistics.
  42. 42.Ohad Rubin, Jonathan Herzig, and Jonathan Berant. 2022. Learning to retrieve prompts for in-context learning. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2655–2671, Seattle, United States. Association for Computational Linguistics.
  43. 43.John Schulman. 2023. Reinforcement learning from human feedback: Progress and challenges.
  44. 44.Thomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski, Jacopo Staiano, Alex Wang, and Patrick Gallinari. 2021. QuestEval: Summarization asks for fact-based evaluation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6594–6604, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  45. 45.Christopher Sciavolino, Zexuan Zhong, Jinhyuk Lee, and Danqi Chen. 2021. Simple entity-centric questions challenge dense retrievers. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, pages 6138–6148. Association for Computational Linguistics.
  46. 46.Katherine Tian, Eric Mitchell, Huaxiu Yao, Christopher D Manning, and Chelsea Finn. 2023a. Fine-tuning language models for factuality. arXiv preprint arXiv:2311.08401.
  47. 47.Katherine Tian, Eric Mitchell, Allan Zhou, Archit Sharma, Rafael Rafailov, Huaxiu Yao, Chelsea Finn, and Christopher D Manning. 2023b. Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback. arXiv preprint arXiv:2305.14975.
  48. 48.Pranav Narayanan Venkit, Tatiana Chakravorti, Vipul Gupta, Heidi Biggs, Mukund Srinath, Koustava Goswami, Sarah Rajtmajer, and Shomir Wilson. 2024. " confidently nonsensical?”: A critical survey on the perspectives and challenges of’hallucinations’ in nlp. arXiv preprint arXiv:2404.07461.
  49. 49.Denny Vrandečić and Markus Krötzsch. 2014. Wikidata: a free collaborative knowledgebase. Commun. ACM, 57(10):78–85.
  50. 50.Cunxiang Wang, Sirui Cheng, Qipeng Guo, Yuanhao Yue, Bowen Ding, Zhikun Xu, Yidong Wang, Xiangkun Hu, Zheng Zhang, and Yue Zhang. 2023. Evaluating open-qa evaluation. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023.
  51. 51.Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le. 2022. Finetuned language models are zero-shot learners. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net.
  52. 52.Yuqing Yang, Ethan Chern, Xipeng Qiu, Graham Neubig, and Pengfei Liu. 2023. Alignment for honesty. arXiv preprint arXiv:2312.07000.
  53. 53.Xunjian Yin, Baizhou Huang, and Xiaojun Wan. 2023. ALCUNA: Large language models meet new knowledge. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 1397–1414, Singapore. Association for Computational Linguistics.
  54. 54.Gal Yona, Roee Aharoni, and Mor Geva. 2024. Narrowing the knowledge evaluation gap: Open-domain question answering with multi-granularity answers. arXiv preprint arXiv:2401.04695.
  55. 55.Hanning Zhang, Shizhe Diao, Yong Lin, Yi R Fung, Qing Lian, Xingyao Wang, Yangyi Chen, Heng Ji, and Tong Zhang. 2023. R-tuning: Teaching large language models to refuse unknown questions. arXiv preprint arXiv:2311.09677.
  56. 56.Zexuan Zhong, Zhengxuan Wu, Christopher D. Manning, Christopher Potts, and Danqi Chen. 2023. Mquake: Assessing knowledge editing in language models via multi-hop questions. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023, pages 15686–15702. Association for Computational Linguistics.
  57. 57.Chunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, Susan Zhang, Gargi Ghosh, Mike Lewis, Luke Zettlemoyer, and Omer Levy. 2023. LIMA: less is more for alignment. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023.
  58. 58.Chen Zhu, Ankit Singh Rawat, Manzil Zaheer, Srinadh Bhojanapalli, Daliang Li, Felix Yu, and Sanjiv Kumar. 2020. Modifying memories in transformer models. arXiv preprint arXiv:2012.00363.

Citation

MLA
Gekhman, Z., et al. “Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024, pp. 7765–84, https://doi.org/10.18653/v1/2024.emnlp-main.444.
APA
Gekhman, Z., Yona, G., Aharoni, R., Eyal, M., Feder, A., Reichart, R., & Herzig, J. (2024). Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 7765–7784. https://doi.org/10.18653/v1/2024.emnlp-main.444
Chicago
Gekhman, Z., G. Yona, R. Aharoni, et al. 2024. “Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 7765–84. https://doi.org/10.18653/v1/2024.emnlp-main.444.
Harvard
Gekhman, Z. et al. (2024) “Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?”, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 7765–7784. Available at: https://doi.org/10.18653/v1/2024.emnlp-main.444.
Vancouver
1. Gekhman Z, Yona G, Aharoni R, Eyal M, Feder A, Reichart R, Herzig J (2024) Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 7765–7784

BibTeX

@inproceedings{gekhman-etal-2024-fine,
    title = "Does Fine-Tuning {LLM}s on New Knowledge Encourage Hallucinations?",
    author = "Gekhman, Zorik  and
      Yona, Gal  and
      Aharoni, Roee  and
      Eyal, Matan  and
      Feder, Amir  and
      Reichart, Roi  and
      Herzig, Jonathan",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.444/",
    doi = "10.18653/v1/2024.emnlp-main.444",
    pages = "7765--7784"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/