Unfamiliar Finetuning Examples Control How Language Models Hallucinate

Katie KangEric WallaceClaire J. TomlinAviral KumarSergey Levine

article2025NAACL107 citations

Reveals that language models default to mirroring their unfamiliar finetuning data when hallucinating, enabling a practical method to mitigate factual errors in long-form generation by supervising unfamiliar examples with conservative reward models.

Listen

Large language models frequently generate plausible but factually incorrect statements, commonly known as hallucinations. This failure mode poses significant operational and reputational risks, particularly when models are queried about topics that extend beyond their underlying training data. The article addresses the root mechanisms driving these errors during model adaptation and aims to demonstrate how practitioners can systematically control and reduce hallucinations in both short-form answering and complex, long-form text generation.

To establish these principles, the authors conducted controlled experiments across multiple fine-tuning paradigms, including supervised fine-tuning, reinforcement learning, and reward model training using established benchmarks like MMLU and TriviaQA. They evaluated open-source base models (primarily Llama 2 7B and Mistral 7B) by categorizing test queries by unfamiliarity based on baseline model performance. The investigation then extended to long-form generation tasks—specifically biography generation on WikiBios and narrative plot summaries on WikiPlots—evaluating factual accuracy using automated fact-checking pipelines.

The findings reveal that when models face unfamiliar queries, their responses consistently default toward the distribution of responses present in their unfamiliar fine-tuning data rather than making arbitrary errors. In supervised fine-tuning, if unfamiliar training examples are labeled with abstentions like "I don't know," the model reliably learns to abstain on unfamiliar test prompts while answering familiar ones accurately. In reinforcement learning, scoring models (reward models) often hallucinate by overestimating rewards on unfamiliar concepts, which inadvertently teaches language models to fabricate information. Finally, the authors show that training a "conservative reward model" on responses generated by the model's own base distribution forces the reward model to assign low scores to unknown concepts; fine-tuning with this conservative reward model significantly reduced false facts across all unfamiliarity levels while maintaining or increasing true factual statements.

These insights demonstrate that language model hallucinations are manageable rather than entirely random. In practice, uncurated fine-tuning data and standard reward models actively incentivize models to produce confident, incorrect outputs. By structuring unfamiliar training data and reward systems to favor cautious predictions, organizations can reduce error rates and improve output reliability in high-stakes generative applications without requiring prohibitively expensive real-time verification at every generation step.

Engineering teams should audit their fine-tuning datasets to ensure unfamiliar examples explicitly model uncertainty or abstention. For reinforcement learning workflows, practitioners should adopt conservative reward modeling protocols—using the target model's own generations to populate the reward dataset—to prevent reward inflation on unfamiliar queries. However, leaders should note that these experiments primarily addressed specific generation tasks and clear distinctions between known and unknown data. Because many real-world enterprise queries inhabit a spectrum of partial familiarity, additional testing and pilot evaluations across broader, generalized domains are recommended before critical production deployment.

Cover for Unfamiliar Finetuning Examples Control How Language Models Hallucinate

Abstract

Large language models are known to hallucinate, but the underlying mechanism that govern how models hallucinate are not yet fully understood. In this work, we find that unfamiliar examples in the models’ finetuning data – those that introduce concepts beyond the base model’s scope of knowledge – are crucial in shaping these errors. In particular, we find that an LLM’s hallucinated predictions tend to mirror the responses associated with its unfamiliar finetuning examples. This suggests that by modifying how unfamiliar finetuning examples are supervised, we can influence a model’s responses to unfamiliar queries (e.g., say “I don’t know”). We empirically validate this observation in a series of controlled experiments involving SFT, RL, and reward model finetuning on TriviaQA and MMLU. Our work further investigates RL finetuning strategies for improving the factuality of long-form model generations. We find that, while hallucinations from the reward model can significantly undermine the effectiveness of RL factuality finetuning, strategically controlling how reward models hallucinate can minimize these negative effects. Leveraging our previous observations on controlling hallucinations, we propose an approach for learning more reliable reward models, and show that they improve the efficacy of RL factuality finetuning in long-form biography and book/movie plot generation tasks.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Problem Setting
  • 4 Understanding How LLMs Hallucinate
  • 4.1 Main Hypothesis
  • 4.2 Experiments
  • 5 Controlling Hallucinations in Long-Form Generations
  • 5.1 RL Factuality Finetuning with Conservative Reward Models
  • 5.2 Experiments
  • 6 Conclusion
  • 7 Limitations
  • References
  • A Additional Results
  • B Compute and Other Details
  • C MMLUTraining Details
  • C.1 SFT Models
  • C.2 RL Models
  • C.3 Reward Models
  • D TriviaQA Training Details
  • D.1 SFT Models
  • D.2 RL Models
  • D.3 Reward Models
  • E Long-form Tasks Training Details
  • E.1 Data
  • E.2 Reward Models
  • E.3 SFT Models
  • E.4 RL Models

Knowls

  1. Knowl 1 — Unfamiliar queries elicit an input-independent finetuning default

    model/method

    The paper’s conceptual model treats an input as unfamiliar when the pretrained language model lacks useful query–response associations for it. For a finetuning dataset Dunf={(xi,si)}D_{\mathrm{unf}}=\{(x_i,s_i)\} of unfamiliar query–supervision pairs, define Y\mathcal{Y} as the response space, QQ as a probability distribution over Y\mathcal{Y}, and L(Q,si)L(Q,s_i) as the training loss for supervision sis_i. The model’s proposed default response distribution is an aggregate-loss minimizer:

    Punf∈arg⁡min⁡Q∈Δ(Y)∑(xi,si)∈DunfL(Q,si),P_{\mathrm{unf}} \in \arg\min_{Q\in\Delta(\mathcal{Y})}\sum_{(x_i,s_i)\in D_{\mathrm{unf}}} L(Q,s_i),

    where Δ(Y)\Delta(\mathcal{Y}) is the set of probability distributions over responses. Because this default is optimized across the unfamiliar examples rather than for a particular query, it is input-independent. The paper hypothesizes that a finetuned model approaches this distribution on unfamiliar test inputs, so its errors or abstentions reflect how unfamiliar training examples were supervised.

  2. Knowl 2 — Conservative reward-model training controls reward overestimation

    model/method

    A conservative reward model is trained so that its unfamiliar finetuning examples tend to have low rewards, making low reward predictions the model’s default on unfamiliar inputs. The proposed procedure is: (1) use the pretrained model to be used for reward modeling to perform the target task, either after supervised finetuning or with few-shot prompting; (2) sample responses to task prompts from that model; (3) assign the responses ground-truth factuality rewards; and (4) train a reward model on the labeled query–response pairs. Since the response-generating model and reward model share a pretrained knowledge base, responses that are incorrect on unfamiliar queries are also likely to be unfamiliar to the reward model. The resulting low-reward training examples are intended to limit high, unjustified reward predictions on unfamiliar responses. This method still requires ground-truth reward labels, but the paper argues that it requires fewer such labels than using ground-truth rewards for all responses during online RL.

  3. Knowl 3 — Conservative-reward RL reduces false facts in long-form generations

    empirical result

    In biography and book/movie-plot generation, the authors compared standard supervised finetuning (SFT), PPO-based RL with a standard reward model, and PPO-based RL with a conservative reward model. As query unfamiliarity increased, standard SFT produced fewer true facts and more false facts. Compared with SFT, RL using the conservative reward model produced the same or more true facts and significantly fewer false facts across the evaluated unfamiliarity levels. Compared with RL using the standard reward model, conservative-reward RL produced around the same number of true facts but substantially fewer false facts across those levels. These results support the claim that controlling reward-model hallucinations can improve the factuality of long-form generations.

  4. Knowl 4 — SFT predictions approach the response distribution used for unfamiliar examples

    empirical result

    Controlled SFT experiments on MMLU and TriviaQA tested whether changing only the supervision for unfamiliar finetuning examples changes predictions on unfamiliar test queries. On MMLU, one dataset assigned target answer letters uniformly across A–D for both familiar and unfamiliar examples; another kept familiar targets uniform but assigned unfamiliar targets 50% B and 50% C. As test queries became more unfamiliar, the first model’s answer probabilities approached a uniform distribution, while the second model’s probabilities approached the 50% B/50% C distribution. On TriviaQA, the authors compared (i) ground-truth answers for all training examples, (ii) ground-truth answers for familiar examples and “I don’t know” for unfamiliar examples, and (iii) the same overall proportion of “I don’t know” labels as (ii), but assigned to a random subset rather than specifically to unfamiliar examples. The first model increasingly gave incorrect answers on unfamiliar queries; the second increasingly abstained; the third did not show the same unfamiliarity-dependent shift. The comparison indicates that the placement of abstention supervision on unfamiliar examples matters, not merely its overall frequency.

  5. Knowl 5 — RL predictions approach the action distribution favored by unfamiliar-example rewards

    empirical result

    The authors finetuned models with PPO on MMLU and TriviaQA using two reward functions. Both rewarded a correct answer with +2+2 and an incorrect answer with −3-3; one also gave abstention −3-3, while the other gave abstention 00. For MMLU, abstention was option E; for TriviaQA, it was “I don’t know.” Under the first reward function, random guessing has higher average reward than abstaining on unfamiliar queries, so the aggregate-reward default is guessing. Under the second, abstaining has higher average reward, so the default is abstention. Models favored correct answers on familiar queries, but as queries became more unfamiliar, models trained with the first reward function increasingly guessed and models trained with the second increasingly abstained. This is evidence that the reward function governing unfamiliar examples shapes RL-finetuned predictions.

  6. Knowl 6 — Reward-model predictions reflect reward labels on unfamiliar training examples

    empirical result

    The paper trained factuality reward models to classify a query–response pair as correct (reward 1) or incorrect (reward 0). In both MMLU and TriviaQA experiments, familiar training examples were balanced between the two reward labels. The authors compared one training distribution in which all unfamiliar examples had reward 1 with another in which all unfamiliar examples had reward 0. As test queries became more unfamiliar, the first model increasingly predicted reward 1 and the second increasingly predicted reward 0. Thus, reward models can hallucinate in different directions depending on the reward labels associated with their unfamiliar finetuning examples, potentially creating different biases in subsequent RL.

  7. Knowl 7 — Overestimated rewards are hypothesized to be more harmful than underestimated rewards

    assumption

    The authors hypothesize that overestimating factuality is more damaging to RL factuality finetuning than underestimating it. Under a reward that gives positive value to correct facts and negative value to incorrect facts, a response containing an incorrect fact should score lower than a comparable response that omits that fact. If a reward model instead marks the incorrect fact as true and favors the response containing it, RL can be driven to generate more false information. This motivates designing reward models that default to low rewards on unfamiliar inputs; the paper presents this asymmetry as a rationale for the method, rather than as a separately proven result.

  8. Knowl 8 — Long-form evaluation compares reward models and generation methods under controlled conditions

    experimental setup

    The long-form experiments used biography prompts from WikiBios (“Write a biography for [name].”) and plot prompts from WikiPlots (“What is the premise of [title]?”); the first sentence of each Wikipedia entry served as the SFT target. The biography data contained 104,539 finetuning examples and 5,000 evaluation examples; the plot data contained 10,000 finetuning examples and 4,795 evaluation examples. All generation systems used Llama2 7B as the pretrained model. The standard reward model was trained on GPT-3.5-generated responses, while the conservative reward model used responses generated by Llama2 7B from the same prompts; FActScore supplied factuality labels, assigning +2+2 per correct fact and −3-3 per incorrect fact. In the implementation, reward modeling first decomposed responses into atomic facts and then predicted each fact’s factuality, from which response rewards were calculated; 10,000 labeled examples were used to train each reward-model approach for each task. The SFT model and the two PPO RL models used the same finetuning prompts, and the RL comparisons differed in reward model. FActScore measured generated true and false facts. Query unfamiliarity was assessed by few-shot prompting Llama2 7B, sampling two responses, and averaging their number of incorrect facts.

  9. Knowl 9 — Unfamiliarity is estimated from pretrained-model performance

    definition

    An unfamiliar input is a query whose answer requires concepts beyond the pretrained model’s knowledge, even though the query is stylistically similar to finetuning examples. The authors evaluated held-out queries from the same task distribution as finetuning examples to isolate shifts relative to pretraining, and treated higher task-specific error as greater unfamiliarity. For MMLU, unfamiliarity was measured by the negative log-likelihood of the correct answer under the few-shot pretrained model; for TriviaQA, it was the number of incorrect responses among 12 sampled few-shot answers. In the reported SFT training setup, MMLU examples with score above 0.36 and TriviaQA examples with more than 6 incorrect samples were classified as unfamiliar. For long-form evaluation, unfamiliarity was measured as the average number of incorrect facts in two responses sampled from few-shot Llama2 7B.

  10. Knowl 10 — The analysis leaves partial familiarity and broader task generalization open

    limitation

    The conceptual account focuses on entirely unfamiliar examples and does not provide a comparably detailed characterization of queries with partial familiarity, which the authors identify as common in real-world use. The factuality experiments also focus on models finetuned for particular applications, such as biography generation; the effectiveness of extending the approach to more general prompted-generation tasks remains unestablished.

Coverage note — Appendix optimizer and hardware settings, illustrative generations, and the additional Mistral 7B replication are omitted as implementation details or corroborative examples rather than distinct load-bearing contributions.

References

  1. 1.Ayush Agrawal, Lester Mackey, and Adam Tauman Kalai. 2023. Do language models know when they’re hallucinating references? arXiv preprint arXiv:2305.18248.
  2. 2.Amos Azaria and Tom Mitchell. 2023. The internal state of an LLM knows when its lying. arXiv preprint arXiv:2304.13734.
  3. 3.Jon Bell. 2017. Wikiplots.
  4. 4.Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al. 2023. Sparks of artificial general intelligence: Early experiments with GPT-4. arXiv preprint arXiv:2303.12712.
  5. 5.Yung-Sung Chuang, Yujia Xie, Hongyin Luo, Yoon Kim, James Glass, and Pengcheng He. 2023. DoLa: Decoding by contrasting layers improves factuality in large language models. arXiv preprint arXiv:2309.03883.
  6. 6.Luyu Gao, Zhuyun Dai, Panupong Pasupat, Anthony Chen, Arun Tejasvi Chaganty, Yicheng Fan, Vincent Zhao, Ni Lao, Hongrae Lee, Da-Cheng Juan, et al. 2023. RARR: Researching and revising what language models say, using language models. In ACL.
  7. 7.Yoav Goldberg. 2023. Reinforcement learning for language models.
  8. 8.Alexander Havrilla, Maksym Zhuravinskyi, Duy Phung, Aman Tiwari, Jonathan Tow, Stella Biderman, Quentin Anthony, and Louis Castricato. 2023. trlX: A framework for large scale reinforcement learning from human feedback. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 8578–8595, Singapore. Association for Computational Linguistics.
  9. 9.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2020. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300.
  10. 10.Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2023. Mistral 7b. Preprint, arXiv:2310.06825.
  11. 11.Liqiang Jing, Ruosen Li, Yunmo Chen, Mengzhao Jia, and Xinya Du. 2023. FAITHSCORE: Evaluating hallucinations in large vision-language models. arXiv preprint arXiv:2311.01477.
  12. 12.Mandar Joshi, Eunsol Choi, Daniel S Weld, and Luke Zettlemoyer. 2017. Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension. arXiv preprint arXiv:1705.03551.
  13. 13.Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, et al. 2022. Language models (mostly) know what they know. arXiv preprint arXiv:2207.05221.
  14. 14.Adam Tauman Kalai and Santosh S Vempala. 2023. Calibrated language models must hallucinate. arXiv preprint arXiv:2311.14648.
  15. 15.Nikhil Kandpal, Haikang Deng, Adam Roberts, Eric Wallace, and Colin Raffel. 2023. Large language models struggle to learn long-tail knowledge. In International Conference on Machine Learning.
  16. 16.Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar. 2023. Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation. arXiv preprint arXiv:2302.09664.
  17. 17.Nayeon Lee, Wei Ping, Peng Xu, Mostofa Patwary, Pascale N Fung, Mohammad Shoeybi, and Bryan C Catanzaro. 2022. Factuality enhanced language models for open-ended text generation. Advances in Neural Information Processing Systems.
  18. 18.Kenneth Li, Oam Patel, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg. 2023. Inference-time intervention: Eliciting truthful answers from a language model. arXiv preprint arXiv:2306.03341.
  19. 19.Stephanie Lin, Jacob Hilton, and Owain Evans. 2022. Teaching models to express their uncertainty in words. arXiv preprint arXiv:2205.14334.
  20. 20.Kevin Liu, Stephen Casper, Dylan Hadfield-Menell, and Jacob Andreas. 2023. Cognitive dissonance: Why do language model outputs disagree with internal representations of truthfulness? arXiv preprint arXiv:2312.03729.
  21. 21.Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Daniel Khashabi, and Hannaneh Hajishirzi. 2023. When not to trust language models: Investigating effectiveness of parametric and non-parametric memories. In ACL.
  22. 22.Potsawee Manakul, Adian Liusie, and Mark JF Gales. 2023. Selfcheckgpt: Zero-resource black-box hallucination detection for generative large language models. arXiv preprint arXiv:2303.08896.
  23. 23.Mohsen Mesgar, Edwin Simpson, and Iryna Gurevych. 2020. Improving factual consistency between a response and persona facts. arXiv preprint arXiv:2005.00036.
  24. 24.Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Wei Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2023. FActScore: Fine-grained atomic evaluation of factual precision in long form text generation. arXiv preprint arXiv:2305.14251.
  25. 25.Niels Mündler, Jingxuan He, Slobodan Jenko, and Martin Vechev. 2023. Self-contradictory hallucinations of large language models: Evaluation, detection and mitigation. arXiv preprint arXiv:2305.15852.
  26. 26.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744.
  27. 27.Baolin Peng, Michel Galley, Pengcheng He, Hao Cheng, Yujia Xie, Yu Hu, Qiuyuan Huang, Lars Liden, Zhou Yu, Weizhu Chen, et al. 2023. Check your facts and try again: Improving large language models with external knowledge and automated feedback. arXiv preprint arXiv:2302.12813.
  28. 28.Paul Roit, Johan Ferret, Lior Shani, Roee Aharoni, Geoffrey Cideron, Robert Dadashi, Matthieu Geist, Sertan Girgin, Léonard Hussenot, Orgad Keller, et al. 2023. Factually consistent summarization via reinforcement learning with textual entailment feedback. arXiv preprint arXiv:2306.00186.
  29. 29.John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347.
  30. 30.John Shulman. 2023. Reinforcement learning from human feedback: Progress and challenges.
  31. 31.Kurt Shuster, Spencer Poff, Moya Chen, Douwe Kiela, and Jason Weston. 2021. Retrieval augmentation reduces hallucination in conversation. arXiv preprint arXiv:2104.07567.
  32. 32.Chenglei Si, Zhe Gan, Zhengyuan Yang, Shuohang Wang, Jianfeng Wang, Jordan Boyd-Graber, and Lijuan Wang. 2022. Prompting GPT-3 to be reliable. arXiv preprint arXiv:2210.09150.
  33. 33.Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano. 2020. Learning to summarize with human feedback. Advances in Neural Information Processing Systems, 33:3008–3021.
  34. 34.Marco Antonio Stranisci, Rossana Damiano, Enrico Mensa, Viviana Patti, Daniele Radicioni, and Tommaso Caselli. 2023. Wikibio: a semantic resource for the intersectional analysis of biographical events. arXiv preprint arXiv:2306.09505.
  35. 35.Zhiqing Sun, Sheng Shen, Shengcao Cao, Haotian Liu, Chunyuan Li, Yikang Shen, Chuang Gan, Liang-Yan Gui, Yu-Xiong Wang, Yiming Yang, et al. 2023. Aligning large multimodal models with factually augmented RLHF. arXiv preprint arXiv:2309.14525.
  36. 36.Katherine Tian, Eric Mitchell, Huaxiu Yao, Christopher D Manning, and Chelsea Finn. 2023a. Finetuning language models for factuality. arXiv preprint arXiv:2311.08401.
  37. 37.Katherine Tian, Eric Mitchell, Allan Zhou, Archit Sharma, Rafael Rafailov, Huaxiu Yao, Chelsea Finn, and Christopher D Manning. 2023b. Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback. arXiv preprint arXiv:2305.14975.
  38. 38.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  39. 39.Logesh Kumar Umapathi, Ankit Pal, and Malaikannan Sankarasubbu. 2023. Med-HALT: Medical domain hallucination test for large language models. arXiv preprint arXiv:2307.15343.
  40. 40.Neeraj Varshney, Wenlin Yao, Hongming Zhang, Jianshu Chen, and Dong Yu. 2023. A stitch in time saves nine: Detecting and mitigating hallucinations of LLMs by validating low-confidence generation. arXiv preprint arXiv:2307.03987.
  41. 41.Weijia Xu, Sweta Agrawal, Eleftheria Briakou, Marianna J Martindale, and Marine Carpuat. 2023. Understanding and detecting hallucinations in neural machine translation via model introspection. TACL.
  42. 42.Yuqing Yang, Ethan Chern, Xipeng Qiu, Graham Neubig, and Pengfei Liu. 2023. Alignment for honesty. arXiv preprint arXiv:2312.07000.
  43. 43.Sina J Semnani Violet Z Yao, Heidi C Zhang, and Monica S Lam. 2023. WikiChat: Combating hallucination of large language models by few-shot grounding on wikipedia. arXiv preprint arXiv:2305.14292.
  44. 44.Hanning Zhang, Shizhe Diao, Yong Lin, Yi R Fung, Qing Lian, Xingyao Wang, Yangyi Chen, Heng Ji, and Tong Zhang. 2023a. R-tuning: Teaching large language models to refuse unknown questions. arXiv preprint arXiv:2311.09677.
  45. 45.Yue Zhang, Leyang Cui, Wei Bi, and Shuming Shi. 2023b. Alleviating hallucinations of large language models through induced hallucinations. arXiv preprint arXiv:2312.15710.
  46. 46.Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021. Calibrate before use: Improving few-shot performance of language models. In International Conference on Machine Learning.

Citation

MLA
Kang, K., et al. “Unfamiliar Finetuning Examples Control How Language Models Hallucinate”. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2025, pp. 3600–12, https://doi.org/10.18653/v1/2025.naacl-long.183.
APA
Kang, K., Wallace, E., Tomlin, C., Kumar, A., & Levine, S. (2025). Unfamiliar Finetuning Examples Control How Language Models Hallucinate. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 3600–3612. https://doi.org/10.18653/v1/2025.naacl-long.183
Chicago
Kang, K., E. Wallace, C. Tomlin, A. Kumar, and S. Levine. 2025. “Unfamiliar Finetuning Examples Control How Language Models Hallucinate”. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 3600–3612. https://doi.org/10.18653/v1/2025.naacl-long.183.
Harvard
Kang, K. et al. (2025) “Unfamiliar Finetuning Examples Control How Language Models Hallucinate”, Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp. 3600–3612. Available at: https://doi.org/10.18653/v1/2025.naacl-long.183.
Vancouver
1. Kang K, Wallace E, Tomlin C, Kumar A, Levine S (2025) Unfamiliar Finetuning Examples Control How Language Models Hallucinate. In: Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp 3600–3612

BibTeX

@inproceedings{kang-etal-2025-unfamiliar,
    title = "Unfamiliar Finetuning Examples Control How Language Models Hallucinate",
    author = "Kang, Katie  and
      Wallace, Eric  and
      Tomlin, Claire  and
      Kumar, Aviral  and
      Levine, Sergey",
    editor = "Chiruzzo, Luis  and
      Ritter, Alan  and
      Wang, Lu",
    booktitle = "Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
    month = apr,
    year = "2025",
    address = "Albuquerque, New Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.naacl-long.183/",
    doi = "10.18653/v1/2025.naacl-long.183",
    pages = "3600--3612",
    ISBN = "979-8-89176-189-6"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/