Fine-Tuning Language Models for Factuality

Katherine TianEric MitchellHuaxiu YaoChristopher D. ManningChelsea Finn

article2024ICLR296 citations

Demonstrates that training language models with direct preference optimization on automatically generated factuality rankings reduces open-ended generation errors by up to 58% without requiring human annotation.

Listen

Large language models often generate fluent and convincing text that contains factual errors, known as hallucinations. This tendency poses significant operational, reputational, and safety risks, especially as organizations increasingly rely on automated text generation for research, customer service, and specialized domains. Traditional methods to address this issue rely heavily on human fact-checking, which is expensive and slow, costing thousands of dollars to evaluate even modest datasets. Consequently, scalable and cost-effective methods are required to improve model reliability without human annotation.

The article demonstrates an automated fine-tuning pipeline designed to systematically reduce factual errors in open-ended text generation without requiring human intervention. It evaluates how effectively preference-based reinforcement learning can optimize language models for factuality across complex domains like biographical writing and medical question-answering.

To achieve this, the authors developed two automated methods to evaluate candidate model responses and generate preference rankings without human annotators. The first method uses reference-based fact-checking, where individual claims are extracted and verified against Wikipedia data. The second method is entirely reference-free, measuring the model's internal confidence by converting factual claims into targeted questions, resampling answers, and computing semantic consistency. These automated rankings were then used to train 7-billion-parameter open-source models (Llama-1 and Llama-2) using Direct Preference Optimization, an efficient preference-learning algorithm.

The findings show that automated factuality tuning substantially enhances accuracy across tasks. Fine-tuning with reference-based preferences reduced factual error rates in Llama-2 by 58% on biography generation and by 40% on medical question-answering compared to standard instruction-tuned chat models. This approach was the only method evaluated that achieved a strict improvement, simultaneously increasing the total number of correct statements while reducing incorrect ones. The reference-free approach also outperformed standard reinforcement learning baselines, reducing biography errors by over 50% and medical errors by 20% to 30% without using external knowledge bases. Additionally, human evaluations and advanced automated reviews validated these improvements, confirming that accuracy gains were genuine.

These results demonstrate that organizations can significantly mitigate the risk of generating inaccurate information while avoiding the high costs of human feedback pipelines. The findings also reveal that factuality tuning alters text style toward more direct and concise responses rather than conversational storytelling. Crucially, the technique works complementarily with inference-time decoding interventions, allowing multiple safety and accuracy methods to be layered together.

Organizations developing or deploying language models should adopt automated factuality preference pipelines before releasing systems for knowledge-critical applications. For general domains with strong reference data, reference-based tuning offers the highest accuracy gains, while reference-free confidence scoring provides a viable alternative for proprietary or niche domains lacking reference corpora. Future work should pilot these techniques on larger language models and explore deeper integration between factuality objectives and broader conversational capabilities.

Confidence in these findings is strong for 7-billion-parameter models across the tested biographical and medical tasks. However, decision-makers should exercise caution when extrapolating results to larger model architectures or different tasks, as real-world production environments and complex reasoning domains may present additional unmeasured failure modes.

arXiv: 2311.08401
Cover for Fine-Tuning Language Models for Factuality

Abstract

The fluency and creativity of large pre-trained language models (LLMs) have led to their widespread use, sometimes even as a replacement for traditional search engines. Yet language models are prone to making convincing but factually inaccurate claims, often referred to as 'hallucinations.' These errors can inadvertently spread misinformation or harmfully perpetuate misconceptions. Further, manual fact-checking of model responses is a time-consuming process, making human factuality labels expensive to acquire. In this work, we fine-tune language models to be more factual, without human labeling and targeting more open-ended generation settings than past work. We leverage two key recent innovations in NLP to do so. First, several recent works have proposed methods for judging the factuality of open-ended text by measuring consistency with an external knowledge base or simply a large model's confidence scores. Second, the direct preference optimization algorithm enables straightforward fine-tuning of language models on objectives other than supervised imitation, using a preference ranking over possible model responses. We show that learning from automatically generated factuality preference rankings, generated either through existing retrieval systems or our novel retrieval-free approach, significantly improves the factuality (percent of generated claims that are correct) of Llama-2 on held-out topics compared with RLHF or decoding strategies targeted at factuality. At 7B scale, compared to Llama-2-chat, we observe 58% and 40% reduction in factual error rate when generating biographies and answering medical questions, respectively.

Table of Contents

  • 1 Introduction
  • 2 Preliminaries
  • 3 Constructing Preferences Encouraging Factuality in Long-Form Text
  • 3.1 Reference-Based Truthfulness Estimation
  • 3.2 Reference-Free Confidence-Based Truthfulness Estimation
  • 3.3 Factuality Tuning: Putting it all Together
  • 4 Experiments
  • 4.1 Fine-Tuning for Factuality Across Domains
  • 4.2 Fine-tuning Chat Models for Factuality
  • 4.3 Complementary Benefits of Factuality Tuning and Decoding-Time Factuality Interventions
  • 4.4 Impact of Design Decisions of Open-Ended Model Confidence Scoring
  • 4.5 Validating Metrics for Factuality
  • 5 Related Work
  • 6 Conclusion
  • References
  • A Appendix
  • A.1 Prompts
  • A.2 Sample Model Generations

Knowls

  1. Knowl 1 — Automated Factuality Fine-Tuning Pipeline via Direct Preference Optimization

    model/method

    The factuality fine-tuning pipeline automatically improves the factual accuracy of language models on open-ended generation tasks without human labeling. The process consists of four main stages:

    1. Candidate Generation: Given a prompt dataset Dp\mathcal{D}_p of open-ended queries xx (e.g., "Write a short biography of Mary Wollstonecraft"), nn candidate responses are sampled from a reference language model πref\pi_{\text{ref}} using temperature sampling at τ=1.0\tau = 1.0.

    2. Factuality Scoring: Each response is assigned a scalar truthfulness score measuring factual precision. This score is calculated via either a reference-based method (measuring support from an external knowledge base like Wikipedia) or a reference-free method (measuring the model's own calibrated uncertainty across extracted atomic claims).

    3. Preference Dataset Construction: For each prompt xx, all (n2)\binom{n}{2} pairs of candidate completions (yw,yl)(y_w, y_l) are ranked such that yw≻yly_w \succ y_l if the truthfulness score of ywy_w is strictly greater than that of yly_l. Pairs with identical scores are discarded. For mm prompts, this produces m(n2)−km \binom{n}{2} - k preference pairs, where kk is the number of tied pairs.

    4. Direct Preference Optimization (DPO): The model is fine-tuned on this preference dataset using the DPO classification loss. The factuality evaluation and claim extraction are required only during offline preference generation; test-time generation follows standard autoregressive sampling.

  2. Knowl 2 — Reference-Free Confidence-Based Factuality Estimation (FactTune-MC)

    model/method

    Reference-free factuality estimation uses a language model's own epistemic uncertainty over individual claims as a proxy for factual correctness, eliminating the requirement for external reference retrieval. The procedure operates as follows:

    1. Atomic Claim Extraction: An LLM (e.g., GPT-3.5) decomposes an open-ended text completion into a set of distinct atomic factual assertions.

    2. Unambiguous Question Generation: Each atomic claim is converted by GPT-3.5 into a specific, minimally ambiguous question designed to test that exact fact. Ambiguity minimization prevents false uncertainty caused by underspecified question phrasing.

    3. Response Resampling: The model being evaluated (or fine-tuned) generates K=20K = 20 independent answer completions for each generated question using few-shot prompting.

    4. Semantic Equivalence Binning: The KK sampled answers are clustered into discrete semantic equivalence classes. In the primary implementation, equivalence is determined via heuristic exact string matching after removing stop words.

    5. Confidence Scoring: For each atomic fact, the truthfulness score is defined as the maximum bin fraction (the number of answers falling into the largest semantic equivalence bin divided by KK).

    6. Aggregation: The overall response-level truthfulness score is the arithmetic mean of the maximum confidence scores across all extracted atomic facts in the response.

  3. Knowl 3 — Reference-Based Factuality Estimation via FactScore (FactTune-FS)

    model/method

    Reference-based factuality estimation scores the factual precision of long-form generation by evaluating atomic statements against an authoritative knowledge base (such as Wikipedia):

    1. Atomic Fact Decomposition: GPT-3.5 extracts a list of all atomic claims present in the generated passage.

    2. Natural Language Inference Checking: For each atomic claim, a fact-checking model (a Llama-1-7B model fine-tuned for natural language inference) determines whether the claim is entailed/supported by the corresponding reference document (e.g., the Wikipedia article for the named entity or medical condition).

    3. Precision Scoring: The truthfulness score assigned to the passage is the fraction of extracted atomic claims that are classified as supported by the reference text.

  4. Knowl 4 — Direct Preference Optimization Loss for Factuality Alignment

    equation

    The direct preference optimization (DPO) objective optimizes a parameterized policy πθ\pi_\theta directly from preference pairs without explicitly fitting a reward model or sampling inside the training loop:

    LDPO(πθ;πref)=−E(x,yw,yl)∼D[log⁡σ(βlog⁡πθ(yw∣x)πref(yw∣x)−βlog⁡πθ(yl∣x)πref(yl∣x))]\mathcal{L}_{\text{DPO}}(\pi_\theta; \pi_{\text{ref}}) = -\mathbb{E}_{(x, y_w, y_l) \sim \mathcal{D}} \left[ \log \sigma \left( \beta \log \frac{\pi_\theta(y_w \mid x)}{\pi_{\text{ref}}(y_w \mid x)} - \beta \log \frac{\pi_\theta(y_l \mid x)}{\pi_{\text{ref}}(y_l \mid x)} \right) \right]

    where:

    • xx is the input query prompt.
    • ywy_w is the winning (more factual) candidate response.
    • yly_l is the losing (less factual) candidate response.
    • D\mathcal{D} is the dataset of preference tuples (x,yw,yl)(x, y_w, y_l).
    • πθ(y∣x)\pi_\theta(y \mid x) is the probability of response yy under the policy model.
    • πref(y∣x)\pi_{\text{ref}}(y \mid x) is the probability of response yy under the reference/initialization model (obtained via supervised fine-tuning).
    • β>0\beta > 0 is a scalar regularizer controlling the Kullback-Leibler (KL) divergence penalty between πθ\pi_\theta and πref\pi_{\text{ref}}.
    • σ(z)=11+e−z\sigma(z) = \frac{1}{1 + e^{-z}} is the logistic sigmoid function, derived from the Bradley-Terry preference model p(yw≻yl∣x)=σ(r(x,yw)−r(x,yl))p(y_w \succ y_l \mid x) = \sigma(r(x, y_w) - r(x, y_l)).
  5. Knowl 5 — Empirical Performance of Factuality Fine-Tuning across Domains

    data/table

    Factuality fine-tuning was evaluated on Llama-1-7B and Llama-2-7B across two tasks: biography generation (355 entities) and medical question-answering (200 entities, 6 questions each). FactTune-FS uses FactScore reference-based preferences; FactTune-MC uses reference-free model confidence preferences. Baseline methods include Supervised Fine-Tuning (SFT), Inference-Time Intervention (ITI), Decoding by Contrasting Layers (DOLA), and RLHF (Llama-2-Chat).

    Biographies Medical QA
    Base Model Method # Correct # Incorrect % Correct # Correct # Incorrect % Correct
    Llama-1 ITI 11.67 6.69 0.669 8.91 5.16 0.633
    DOLA 11.75 3.84 0.754 8.03 5.91 0.576
    SFT 13.78 12.16 0.568 10.75 6.31 0.630
    FactTune-FS 14.81 3.75 0.812 10.88 4.50 0.707
    FactTune-MC 10.59 2.94 0.783 12.31 6.88 0.642
    Llama-2 ITI 18.50 5.75 0.760 10.97 4.06 0.730
    DOLA 13.41 5.84 0.696 9.72 4.38 0.690
    Chat (RLHF) 19.03 6.41 0.748 9.63 5.50 0.636
    SFT 12.19 5.19 0.701 11.75 6.75 0.635
    FactTune-FS 17.06 2.00 0.895 12.53 3.47 0.783
    FactTune-MC 11.31 2.06 0.846 11.41 4.80 0.704

    FactTune-FS achieves strict improvements over SFT across both domains by simultaneously increasing the total number of correct facts and drastically reducing incorrect facts. On Llama-2, FactTune-FS decreases the error count on biographies from 5.19 to 2.00 and on Medical QA from 6.75 to 3.47. FactTune-MC provides a substantial factual accuracy improvement over SFT and RLHF baselines without accessing any external reference text.

  6. Knowl 6 — Factuality Fine-Tuning of Dialogue Models

    data/table

    Factuality fine-tuning can be applied directly to models already aligned with RLHF for conversational dialogue (Llama-2-7B-Chat), improving factuality without degrading overall factual content volume.

    Biographies Medical QA
    Base Model Method # Correct # Incorrect % Correct # Correct # Incorrect % Correct
    Llama-2-Chat Unmodified 19.03 6.41 0.748 9.63 5.50 0.636
    DOLA 21.00 5.19 0.802 11.50 8.25 0.582
    FactTune-FS 19.94 4.06 0.831 9.38 5.25 0.682
    FactTune-MC 20.91 4.84 0.812 10.34 5.69 0.645

    Applying FactTune-FS to Llama-2-Chat raises biographical precision from 74.8% to 83.1% and Medical QA precision from 63.6% to 68.2%, reducing error counts from 6.41 to 4.06 in biographies. FactTune-MC also outperforms unmodified Llama-2-Chat and DOLA decoding on factuality percentage across both tasks.

  7. Knowl 7 — Composability of Factuality Fine-Tuning and Inference-Time Decoding Interventions

    data/table

    Factuality fine-tuning (FactTune-FS) and inference-time decoding interventions (DOLA — Decoding by Contrasting Layers) operate through complementary mechanisms and can be combined additively at generation time.

    Biographies Medical QA
    Base Model Method # Correct # Incorrect % Correct # Correct # Incorrect % Correct
    Llama-1 FactTune-FS 14.81 3.75 0.812 10.88 4.50 0.707
    FactTune-FS + DOLA 12.44 2.00 0.864 11.47 3.75 0.767
    Llama-2 FactTune-FS 17.06 2.00 0.895 12.53 3.47 0.783
    FactTune-FS + DOLA 16.22 2.65 0.865 12.56 3.44 0.794

    Applying DOLA on top of FactTune-FS models increases the proportion of correct facts in three out of four settings (e.g., Llama-1 Biographies improves from 81.2% to 86.4%, Llama-1 Medical QA from 70.7% to 76.7%, and Llama-2 Medical QA from 78.3% to 79.4%).

  8. Knowl 8 — Ablation of Design Choices in Reference-Free Confidence Scoring

    data/table

    An empirical ablation was performed on 12 samples per prompt to assess the components of reference-free truthfulness estimation: extraction method (Named Entity / Noun Chunk vs. GPT-3.5 Atomic Question), semantic equivalence matching (String Heuristic vs. GPT-3.5 LLM), and confidence metric (Semantic Entropy vs. Maximum Confidence).

    Biographies Medical QA
    Fact Ext. Equiv. Metric # Correct # Incorrect % Correct # Correct # Incorrect % Correct
    Entity Heuristic Entropy 13.8 6.31 0.693 9.5 5.47 0.660
    Entity Heuristic Max Conf 12.7 6.31 0.693 9.5 4.78 0.673
    Atomic Heuristic Entropy 10.6 2.88 0.810 12.6 5.25 0.711
    Atomic Heuristic Max Conf 12.2 2.56 0.840 10.2 5.19 0.673
    Atomic LLM Entropy 11.0 3.22 0.778 11.9 6.16 0.661
    Atomic LLM Max Conf 13.7 4.16 0.794 11.7 6.00 0.668

    Key takeaways:

    1. Extraction: Converting facts to atomic questions consistently outperforms entity-level extraction (e.g., 84.0% vs. 69.3% precision on biographies).
    2. Equivalence Matching: Heuristic string matching outperforms GPT-3.5 LLM equivalence checking across settings. The string heuristic underestimates semantic entropy uniformly, whereas LLM matching exhibits bidirectional classification noise that degrades preference pair ordering.
    3. Confidence Metric: Maximum confidence provides higher accuracy than semantic entropy on biographies under atomic extraction (84.0% vs. 81.0%).
  9. Knowl 9 — Validation of Factuality Metrics Against Reward Overoptimization

    empirical result

    To confirm that the performance gains of FactTune models reflect genuine factual accuracy rather than metric overoptimization (Goodhart's Law on FactScore), model outputs were validated against human annotations on Prolific and independent GPT-4 error assessments.

    1. Human Evaluation: Evaluators verified Llama-1-7B responses against reference Wikipedia pages. Human-annotated accuracy was:

      • Biographies: SFT achieved 0.582, whereas FactTune-FS achieved 0.846 (matching FactScore: 0.669 vs. 0.921).
      • Medical QA: SFT achieved 0.662, whereas FactTune-FS achieved 0.838 (matching FactScore: 0.534 vs. 0.806).
    2. GPT-4 Error Correlation: FactScore error counts and GPT-4 error counts demonstrate a strong positive linear correlation across SFT, FactTune-MC, and FactTune-FS across both datasets, verifying that the fine-tuning gains generalize across diverse evaluation protocols.

  10. Knowl 10 — Stylistic and Structural Shifts Induced by Factuality Fine-Tuning

    empirical result

    Factuality fine-tuning causes systematic qualitative modifications to generated text compared to supervised fine-tuning (SFT):

    1. Tone and Directness: Generations from FactTune-FS and FactTune-MC adopt a more objective, terse, and direct tone, omitting casual filler phrases, narrative flourishes, and conversational framing. GPT-4 judged FactTune-FS generations to be less conversational than SFT in 77.5% (n=40n = 40) of Llama-1 cases and 65.6% (n=32n = 32) of Llama-2 cases.
    2. Chronological Structure: In open-ended biographies, factuality-tuned models produce accurate individual atomic facts but occasionally present them outside their natural chronological sequence.

Coverage note — None was omitted; all key methodologies (reference-based and reference-free preference generation), core equations (DPO), main empirical results, dialogue tuning results, decoding composability, scoring ablations, metric validations, and qualitative findings were captured.

References

  1. 1.Ayush Agrawal, Mirac Suzgun, Lester Mackey, and Adam Tauman Kalai. Do language models know when they’re hallucinating references?, 2023. arXiv preprint arxiv:2305.18248. 1
  2. 2.Amos Azaria and Tom Mitchell. The internal state of an LLM knows when its lying, 2023. 9
  3. 3.Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, Ben Mann, and Jared Kaplan. Training a helpful and harmless assistant with reinforcement learning from human feedback, 2022. 3
  4. 4.Ralph Allan Bradley and Milton E Terry. Rank analysis of incomplete block designs: I. the method of paired comparisons. Biometrika, 39(3/4):324–345, 1952. 3
  5. 5.Sebastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, Harsha Nori, Hamid Palangi, Marco Tulio Ribeiro, and Yi Zhang. Sparks of artificial general intelligence: Early experiments with gpt-4, 2023. 9
  6. 6.Meng Cao, Yue Dong, and Jackie Cheung. Hallucinated but factual! inspecting the factuality of hallucinations in abstractive summarization. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 3340–3354, Dublin, Ireland, May 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.acl-long.236. URL https://aclanthology.org/2022.acl-long.236. 9
  7. 7.Hung-Ting Chen, Michael Zhang, and Eunsol Choi. Rich knowledge sources bring complex knowledge conflicts: Recalibrating models to reflect conflicting evidence. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pp. 2292–2307, Abu Dhabi, United Arab Emirates, December 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.emnlp-main.146. URL https://aclanthology.org/2022.emnlp-main.146. 9
  8. 8.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde, Jared Kaplan, Harrison Edwards, Yura Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, David W. Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert-Voss, William H. Guss, Alex Nichol, Igor Babuschkin, S. Arun Balaji, Shantanu Jain, Andrew Carr, Jan Leike, Joshua Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew M. Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, and Wojciech Zaremba. Evaluating large language models trained on code. ArXiv, abs/2107.03374, 2021. URL https://api.semanticscholar.org/CorpusID:235755472. 1
  9. 9.I-Chun Chern, Steffi Chern, Shiqi Chen, Weizhe Yuan, Kehua Feng, Chunting Zhou, Junxian He, Graham Neubig, and Pengfei Liu. Factool: Factuality detection in generative ai – a tool augmented framework for multi-task and multi-domain scenarios, 2023. 2, 3, 9
  10. 10.Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. Deep reinforcement learning from human preferences. In I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett (eds.), Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc., 2017. URL https://proceedings.neurips.cc/paper_files/paper/2017/file/d5e2c0adad503c91f91df240d0cd4e49-Paper.pdf. 1
  11. 11.Yung-Sung Chuang, Yujia Xie, Hongyin Luo, Yoon Kim, James Glass, and Pengcheng He. Dola: Decoding by contrasting layers improves factuality in large language models, 2023. 6, 9
  12. 12.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Alex Castro-Ros, Marie Pellat, Kevin Robinson, Dasha Valter, Sharan Narang, Gaurav Mishra, Adams Yu, Vincent Zhao, Yanping Huang, Andrew Dai, Hongkun Yu, Slav Petrov, Ed H. Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, and Jason Wei. Scaling instruction-finetuned language models, 2022. 1
  13. 13.Shehzaad Dhuliawala, Mojtaba Komeili, Jing Xu, Roberta Raileanu, Xian Li, Asli Celikyilmaz, and Jason Weston. Chain-of-verification reduces hallucination in large language models, 2023. 9
  14. 14.Leo Gao, John Schulman, and Jacob Hilton. Scaling laws for reward model overoptimization, 2022. 3, 9
  15. 15.Luyu Gao, Zhuyun Dai, Panupong Pasupat, Anthony Chen, Arun Tejasvi Chaganty, Yicheng Fan, Vincent Y. Zhao, Ni Lao, Hongrae Lee, Da-Cheng Juan, and Kelvin Guu. Rarr: Researching and revising what language models say, using language models, 2023. 9
  16. 16.Matthew Honnibal and Ines Montani. spaCy 2: Natural language understanding with Bloom embeddings, convolutional neural networks and incremental parsing. To appear, 2017. 5
  17. 17.Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, Scott Johnston, Sheer El-Showk, Andy Jones, Nelson Elhage, Tristan Hume, Anna Chen, Yuntao Bai, Sam Bowman, Stanislav Fort, Deep Ganguli, Danny Hernandez, Josh Jacobson, Jackson Kernion, Shauna Kravec, Liane Lovitt, Kamal Ndousse, Catherine Olsson, Sam Ringer, Dario Amodei, Tom Brown, Jack Clark, Nicholas Joseph, Ben Mann, Sam McCandlish, Chris Olah, and Jared Kaplan. Language models (mostly) know what they know, 2022. URL http://arxiv.org/abs/2207.05221. Arxiv arxiv:2207.05221. 1, 4, 9, 10
  18. 18.Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar. Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation, 2023. 2, 5, 9
  19. 19.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Kuttler, Mike Lewis, Wen-tau Yih, Tim Rocktaschel, Sebastian Riedel, and Douwe Kiela. Retrievalaugmented generation for knowledge-intensive NLP tasks. In H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin (eds.), Advances in Neural Information Processing Systems, volume 33, pp. 9459–9474. Curran Associates, Inc., 2020. URL https://proceedings.neurips.cc/paper_files/paper/2020/file/6b493230205f780e1bc26945df7481e5-Paper.pdf. 9
  20. 20.Kenneth Li, Oam Patel, Fernanda Viegas, Hanspeter Pfister, and Martin Wattenberg. Inference-time intervention: Eliciting truthful answers from a language model, 2023. 6, 9
  21. 21.Shayne Longpre, Kartik Perisetla, Anthony Chen, Nikhil Ramesh, Chris DuBois, and Sameer Singh. Entitybased knowledge conflicts in question answering, 2022. 9
  22. 22.Bill MacCartney and Christopher D. Manning. Modeling semantic containment and exclusion in natural language inference. In Proceedings of the 22nd International Conference on Computational Linguistics (Coling 2008), pp. 521–528, Manchester, UK, August 2008. Coling 2008 Organizing Committee. URL http://www.aclweb.org/anthology/C08-1066. 4
  23. 23.Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Daniel Khashabi, and Hannaneh Hajishirzi. When not to trust language models: Investigating effectiveness of parametric and non-parametric memories. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 9802–9822, Toronto, Canada, July 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.acl-long.546. URL https://aclanthology.org/2023.acl-long.546. 9
  24. 24.Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. On faithfulness and factuality in abstractive summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 1906–1919, Online, July 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.acl-main.173. URL https://aclanthology.org/2020.acl-main.173. 9
  25. 25.Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen tau Yih, Pang Wei Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. Factscore: Fine-grained atomic evaluation of factual precision in long form text generation, 2023. 2, 3, 4, 8, 9
  26. 26.Niels Mundler, Jingxuan He, Slobodan Jenko, and Martin Vechev. Self-contradictory hallucinations of large language models: Evaluation, detection and mitigation, 2023. 9
  27. 27.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. Training language models to follow instructions with human feedback, 2022. 1, 3
  28. 28.Baolin Peng, Michel Galley, Pengcheng He, Hao Cheng, Yujia Xie, Yu Hu, Qiuyuan Huang, Lars Liden, Zhou Yu, Weizhu Chen, and Jianfeng Gao. Check your facts and try again: Improving large language models with external knowledge and automated feedback, 2023. 9
  29. 29.Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, and Chelsea Finn. Direct preference optimization: Your language model is secretly a reward model, 2023. 2, 3
  30. 30.Rajkumar Ramamurthy, Prithviraj Ammanabrolu, Kiante Brantley, Jack Hessel, Rafet Sifa, Christian Bauckhage, Hannaneh Hajishirzi, and Yejin Choi. Is reinforcement learning (not) for natural language processing?: Benchmarks, baselines, and building blocks for natural language policy optimization. 2022. URL https://arxiv.org/abs/2210.01241. 3
  31. 31.Anna Rohrbach, Lisa Anne Hendricks, Kaylee Burns, Trevor Darrell, and Kate Saenko. Object hallucination in image captioning. In Ellen Riloff, David Chiang, Julia Hockenmaier, and Jun’ichi Tsujii (eds.), Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pp. 4035–4045, Brussels, Belgium, October-November 2018. Association for Computational Linguistics. doi: 10.18653/v1/D18-1437. URL https://aclanthology.org/D18-1437. 9
  32. 32.John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms, 2017. 3
  33. 33.Chenglei Si, Zhe Gan, Zhengyuan Yang, Shuohang Wang, Jianfeng Wang, Jordan Boyd-Graber, and Lijuan Wang. Prompting gpt-3 to be reliable, 2023. 9
  34. 34.Nisan Stiennon, Long Ouyang, Jeff Wu, Daniel M. Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul Christiano. Learning to summarize from human feedback. Neural Information Processing Systems, 18, 2020. 3
  35. 35.Katherine Tian, Eric Mitchell, Allan Zhou, Archit Sharma, Rafael Rafailov, Huaxiu Yao, Chelsea Finn, and Christopher D. Manning. Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback, 2023. 1, 4
  36. 36.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothee Lacroix, Baptiste Roziere, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. Llama: Open and efficient foundation language models, 2023a. 1, 4, 9
  37. 37.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. Llama 2: Open foundation and fine-tuned chat models, 2023b. 1, 6
  38. 38.Weijia Xu, Sweta Agrawal, Eleftheria Briakou, Marianna J. Martindale, and Marine Carpuat. Understanding and Detecting Hallucinations in Neural Machine Translation via Model Introspection. Transactions of the Association for Computational Linguistics, 11:546–564, 06 2023. ISSN 2307-387X. doi: 10.1162/tacl_a_00563. URL https://doi.org/10.1162/tacl_a_00563. 9
  39. 39.Muru Zhang, Ofir Press, William Merrill, Alisa Liu, and Noah A Smith. How language model hallucinations can snowball. arXiv preprint arXiv:2305.13534, 2023. 9
  40. 40.Yuhao Zhang, Derek Merck, Emily Tsai, Christopher D Manning, and Curtis Langlotz. Optimizing the factual correctness of a summary: A study of summarizing radiology reports. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL), 2020. URL https://arxiv.org/pdf/1911.02541.pdf. 9
  41. 41.Rui Zheng, Shihan Dou, Songyang Gao, Yuan Hua, Wei Shen, Binghai Wang, Yan Liu, Senjie Jin, Qin Liu, Yuhao Zhou, Limao Xiong, Lu Chen, Zhiheng Xi, Nuo Xu, Wenbin Lai, Minghao Zhu, Cheng Chang, Zhangyue Yin, Rongxiang Weng, Wensen Cheng, Haoran Huang, Tianxiang Sun, Hang Yan, Tao Gui, Qi Zhang, Xipeng Qiu, and Xuanjing Huang. Secrets of RLHF in large language models part I: PPO, 2023. 3
  42. 42.Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving. Fine-tuning language models from human preferences, 2020. 1

Citation

MLA
Tian, K., et al. “Fine-tuning Language Models for Factuality”. arXiv, 2023, http://arxiv.org/abs/2311.08401v1.
APA
Tian, K., Mitchell, E., Yao, H., Manning, C. D., & Finn, C. (2023). Fine-tuning Language Models for Factuality. arXiv. http://arxiv.org/abs/2311.08401v1
Chicago
Tian, K., E. Mitchell, H. Yao, C. D. Manning, and C. Finn. 2023. “Fine-tuning Language Models for Factuality”. arXiv. http://arxiv.org/abs/2311.08401v1.
Harvard
Tian, K. et al. (2023) “Fine-tuning Language Models for Factuality”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2311.08401v1.
Vancouver
1. Tian K, Mitchell E, Yao H, Manning CD, Finn C (2023) Fine-tuning Language Models for Factuality. arXiv

BibTeX

@article{tian2023fine,
  title = {Fine-tuning Language Models for Factuality},
  author = {Tian, Katherine and Mitchell, Eric and Yao, Huaxiu and Manning, Christopher D. and Finn, Chelsea},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2311.08401v1},
  eprint = {2311.08401}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors