Training Language Models to Generate Text with Citations via Fine-grained Rewards
Chengyu HuangZeqiu WuYushi HuWenya Wang
Proposes a reinforcement learning and rejection sampling framework using sentence-level fine-grained rewards for citation precision, recall, and answer correctness, enabling smaller open-source models like LLaMA-2-7B to outperform GPT-3.5-turbo in generating factually accurate, well-cited text.
Large language models frequently generate incorrect claims or hallucinate facts, which undermines user trust and limits their deployment in high-stakes environments. While grounding responses in external knowledge through citations enables straightforward verification, standard prompting techniques yield unreliable citations and poor factual consistency, especially in smaller, open-source models.
The article establishes and evaluates a training framework that uses fine-grained rewards to train language models to produce accurate responses supported by precise in-text citations. The researchers investigate whether decomposing reward signals into localized, specific feedback mechanisms outperforms traditional holistic training strategies.
To conduct the study, the authors initialized a standard open-source model (LLaMA-2-7B) using weak supervision distilled from ChatGPT. They then trained the model using three distinct reward signals targeting specific objectives: answer correctness, sentence-level citation recall (whether cited passages support the sentence), and citation precision (whether cited sources are necessary and not redundant). These granular rewards were applied through two primary training algorithms: sentence-level rejection sampling—a structured candidate ranking method—and token-level reinforcement learning using Proximal Policy Optimization. The framework was evaluated on three diverse question-answering benchmarks (ASQA, QAMPARI, and ELI5) comprising roughly 3,000 test examples, and tested for generalizability on the expert-level EXPERTQA dataset.
The experimental findings demonstrate that training with fine-grained rewards significantly outperforms holistic reward systems across all evaluated metrics and datasets. Combining fine-grained rejection sampling with reinforcement learning produced the strongest results, enabling the 7-billion-parameter open-source model to surpass ChatGPT (GPT-3.5-turbo) across the evaluated benchmarks—achieving average relative improvements of 4.0% on ASQA, 0.9% on QAMPARI, and 10.6% on ELI5. Rejection sampling alone proved more effective and computationally efficient than reinforcement learning alone, though layering reinforcement learning on top of rejection sampling yielded peak performance. Furthermore, the model exhibited strong transferability, maintaining high factual precision exceeding 80% and achieving superior citation support scores on domain-specific questions in EXPERTQA.
These results indicate that organizations do not need to rely solely on massive, proprietary commercial models to achieve reliable, evidence-backed text generation. Smaller, open models trained with fine-grained feedback can deliver superior attribution and verifiable outputs at potentially lower operational costs and with greater deployment control. However, an analysis of remaining errors revealed that models still misinterpret complex source texts (accounting for roughly 63% of observed citation errors) and include redundant citations (roughly 32% of errors), highlighting that source text comprehension remains a critical risk factor.
Organizations implementing attributable generation systems should adopt fine-grained rejection sampling as an efficient, high-impact baseline training strategy before investing in complex reinforcement learning pipelines. Prior to full-scale deployment in production, teams must improve the underlying document retrieval systems, as the model's ability to provide correct answers remains fundamentally constrained by the coverage of retrieved context. Future development should explore iterative multi-round training and methods to bootstrap training data without requiring proprietary model distillation.
These findings carry high confidence across standard question-answering formats, supported by consistent gains across multiple benchmarks and evaluation metrics. However, readers should account for certain limitations: the initial distillation step currently requires access to a proprietary model, and long-form answer evaluations rely in part on automated natural language inference models whose claim extractions may occasionally be incomplete or misaligned with human nuances.
- Paper: Enabling Large Language Models to Generate Text with Citations, Tianyu Gao et al. (2023). ALCE establishes the citation-recall and citation-precision benchmark framework that this paper adapts to measure the effects of training for citation quality.
- Paper: WebGPT: Browser-assisted question-answering with human feedback, Reiichiro Nakano et al. (2021). WebGPT provides an earlier precedent for training long-form, source-citing answers with human feedback and rejection sampling, methods that help contextualize this paper’s training approach.
- Paper: FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation, Sewon Min et al. (2023). FActScore’s atomic evaluation of factual claims provides useful groundwork for understanding the source’s fine-grained correctness reward.
No sufficiently relevant recommendations were found.
