Training Language Models to Generate Text with Citations via Fine-grained Rewards

Chengyu HuangZeqiu WuYushi HuWenya Wang

article2024ACL43 citations

Proposes a reinforcement learning and rejection sampling framework using sentence-level fine-grained rewards for citation precision, recall, and answer correctness, enabling smaller open-source models like LLaMA-2-7B to outperform GPT-3.5-turbo in generating factually accurate, well-cited text.

Listen

Large language models frequently generate incorrect claims or hallucinate facts, which undermines user trust and limits their deployment in high-stakes environments. While grounding responses in external knowledge through citations enables straightforward verification, standard prompting techniques yield unreliable citations and poor factual consistency, especially in smaller, open-source models.

The article establishes and evaluates a training framework that uses fine-grained rewards to train language models to produce accurate responses supported by precise in-text citations. The researchers investigate whether decomposing reward signals into localized, specific feedback mechanisms outperforms traditional holistic training strategies.

To conduct the study, the authors initialized a standard open-source model (LLaMA-2-7B) using weak supervision distilled from ChatGPT. They then trained the model using three distinct reward signals targeting specific objectives: answer correctness, sentence-level citation recall (whether cited passages support the sentence), and citation precision (whether cited sources are necessary and not redundant). These granular rewards were applied through two primary training algorithms: sentence-level rejection sampling—a structured candidate ranking method—and token-level reinforcement learning using Proximal Policy Optimization. The framework was evaluated on three diverse question-answering benchmarks (ASQA, QAMPARI, and ELI5) comprising roughly 3,000 test examples, and tested for generalizability on the expert-level EXPERTQA dataset.

The experimental findings demonstrate that training with fine-grained rewards significantly outperforms holistic reward systems across all evaluated metrics and datasets. Combining fine-grained rejection sampling with reinforcement learning produced the strongest results, enabling the 7-billion-parameter open-source model to surpass ChatGPT (GPT-3.5-turbo) across the evaluated benchmarks—achieving average relative improvements of 4.0% on ASQA, 0.9% on QAMPARI, and 10.6% on ELI5. Rejection sampling alone proved more effective and computationally efficient than reinforcement learning alone, though layering reinforcement learning on top of rejection sampling yielded peak performance. Furthermore, the model exhibited strong transferability, maintaining high factual precision exceeding 80% and achieving superior citation support scores on domain-specific questions in EXPERTQA.

These results indicate that organizations do not need to rely solely on massive, proprietary commercial models to achieve reliable, evidence-backed text generation. Smaller, open models trained with fine-grained feedback can deliver superior attribution and verifiable outputs at potentially lower operational costs and with greater deployment control. However, an analysis of remaining errors revealed that models still misinterpret complex source texts (accounting for roughly 63% of observed citation errors) and include redundant citations (roughly 32% of errors), highlighting that source text comprehension remains a critical risk factor.

Organizations implementing attributable generation systems should adopt fine-grained rejection sampling as an efficient, high-impact baseline training strategy before investing in complex reinforcement learning pipelines. Prior to full-scale deployment in production, teams must improve the underlying document retrieval systems, as the model's ability to provide correct answers remains fundamentally constrained by the coverage of retrieved context. Future development should explore iterative multi-round training and methods to bootstrap training data without requiring proprietary model distillation.

These findings carry high confidence across standard question-answering formats, supported by consistent gains across multiple benchmarks and evaluation metrics. However, readers should account for certain limitations: the initial distillation step currently requires access to a proprietary model, and long-form answer evaluations rely in part on automated natural language inference models whose claim extractions may occasionally be incomplete or misaligned with human nuances.

No sufficiently relevant recommendations were found.

Cover for Training Language Models to Generate Text with Citations via Fine-grained Rewards

Abstract

While recent Large Language Models (LLMs) have proven useful in answering user queries, they are prone to hallucination, and their responses often lack credibility due to missing references to reliable sources. An intuitive solution to these issues would be to include in-text citations referring to external documents as evidence. While previous works have directly prompted LLMs to generate in-text citations, their performances are far from satisfactory, especially when it comes to smaller LLMs. In this work, we propose an effective training framework using fine-grained rewards to teach LLMs to generate highly supportive and relevant citations, while ensuring the correctness of their responses. We also conduct a systematic analysis of applying these fine-grained rewards to common LLM training strategies, demonstrating its advantage over conventional practices. We conduct extensive experiments on Question Answering (QA) datasets taken from the ALCE benchmark and validate the model's generalizability using EXPERTQA. On LLaMA-2-7B, the incorporation of fine-grained rewards achieves the best performance among the baselines, even surpassing that of GPT-3.5-turbo.1

Table of Contents

  • 1 Introduction
  • 2 Problem Definition and Methods
  • 2.1 Distillation from ChatGPT
  • 2.2 Training with Fine-grained Rewards
  • 2.2.1 Rejection Sampling (RS)
  • 2.2.2 Reinforcement Learning (RL)
  • 2.2.3 Combining RS and RL
  • 2.3 Training with Holistic Rewards
  • 3 Experiment Setup
  • 3.1 Datasets
  • 3.2 Evaluation Metrics
  • 3.3 Training Details
  • 3.4 Baselines
  • 4 Results and Analysis
  • 4.1 Main Results
  • 4.2 Ablation of Reward Models
  • 4.3 Retrieval Analysis
  • 4.4 Citation Error Analysis
  • 4.5 Analysis on Generalizability
  • 5 Related Work
  • 6 Conclusion and Future Directions
  • References
  • A Details on Fine-grained Rewards
  • A.1 Correctness Recall Reward
  • A.2 Citation Recall Reward
  • A.3 Citation Precision Reward
  • B Datasets and Metrics
  • B.1 Datasets
  • B.2 Retrieval
  • B.3 Metrics
  • C Additional Training Details
  • D Complete Main Experiments
  • E Retrieval Analysis
  • F Training Curves
  • G Citation Error Analysis
  • H Additional Experiment on EXPERTQA
  • I Prompts
  • J Examples

Knowls

  1. Knowl 1 — Three localized rewards define attributable-generation training

    model/method

    The paper trains a language model to produce answers that are both correct and supported by appropriately chosen citations. It uses three separate rewards, each evaluated at the granularity of the content it measures:

    • Answer correctness, R1R_1: Let tt be the number of gold answer units and hh the number captured by the response. ASQA and QAMPARI use exact-string matches to identify captured units; ELI5 and EXPERTQA use three subclaims inferred from the reference answer and an NLI entailment test. Each captured unit earns +w1+w_1 and each missed unit earns −w1-w_1, except that QAMPARI does not penalize missing answers once the response has captured five. Thus, the response-level reward is R1(x,y)=w1(2h−t)R_1(x,y)=w_1(2h-t) outside QAMPARI, and R1(x,y)=w1[h−max⁡(min⁡(t,5)−h,0)]R_1(x,y)=w_1[h-\max(\min(t,5)-h,0)] on QAMPARI, where xx is the question and yy is the generated response.
    • Citation recall, R2R_2: For each answer sentence sis_i, concatenate its cited passages CiC_i and use an NLI model to test whether they entail sis_i. The sentence receives +w2+w_2 if entailed and −w2-w_2 otherwise. On QAMPARI, each comma-separated answer item is treated as a sentence.
    • Citation precision, R3R_3: Score each citation separately. A citation to passage cj∈Cic_j\in C_i earns +w3+w_3 only if all cited passages together entail sentence sis_i and either cjc_j itself entails sis_i, or removing cjc_j makes the remaining cited passages fail to entail it. Otherwise it earns −w3-w_3.

    The weights w1,w2,w3w_1,w_2,w_3 set the reward scale. This decomposition provides distinct signals for answer content, sentence support, and citation usefulness rather than judging only the response as a whole.

  2. Knowl 2 — Sentence-level fine-grained rejection sampling

    algorithm

    Fine-grained rejection sampling (f.g. RS) uses sentence-level beam search to select training responses for supervised fine-tuning. Given a question and its retrieved passages, the algorithm constructs candidate responses incrementally, scoring partial generations with the three rewards for correctness, citation recall, and citation precision.

    Input: Prompt x; reward functions R1, R2, R3; beam width B; continuations per beam sequence K; maximum sentence steps H
    Output: Highest-scoring completed response for supervised fine-tuning
    Initialize the beam with the empty response
    For each sentence step from 1 to H:
        For each response prefix in the beam:
            Generate K candidate continuations from the language model
        Score every resulting candidate prefix with the sum of its available R1, R2, and R3 segment rewards
        Keep the B highest-scoring candidate prefixes as the next beam
        If all beam sequences have ended generation, stop
    Return the highest-scoring sequence in the beam

    The score for a partial response ywy^w is R(x,yw)=∑u=13∑k=1luRuk(x,yw)R(x,y^w)=\sum_{u=1}^{3}\sum_{k=1}^{l_u}R_u^k(x,y^w), where uu indexes the three reward types, lul_u is the number of segments scored so far for reward RuR_u, and RukR_u^k is that reward for segment kk. Correctness is scored on the response as one segment, while citation recall is scored sentence by sentence and citation precision citation by citation. The selected response becomes the fine-tuning target. In the LLaMA-2-7B experiments, B=8B=8, K=2K=2, and H=5H=5 for ASQA and ELI5 or H=10H=10 for QAMPARI.

  3. Knowl 3 — Fine-grained PPO assigns rewards to their corresponding tokens

    equation

    The paper also trains the language-model policy with proximal policy optimization (PPO) using sparse, localized task rewards and a per-token KL penalty. For a prompt-response pair (x,y)(x,y), let ata_t be the generated token at time tt, gtg_t its generation prefix, Pθ(at∣gt)P_\theta(a_t\mid g_t) the probability of that token under the current policy, and Pθinit(at∣gt)P_{\theta_{\mathrm{init}}}(a_t\mid g_t) its probability under the initial policy. Let T11T^1_1 be the end-of-sequence token position, Ti2T^2_i the last-token position of sentence ii, and Tj3T^3_j the right-bracket position of citation jj. If l1,l2,l3l_1,l_2,l_3 are respectively the number of response-level correctness segments, sentences, and citations, the reward at token tt is

    rt=∑u=13∑k=1lu1(t=Tku)Ruk(x,y)−βlog⁡Pθ(at∣gt)Pθinit(at∣gt).r_t=\sum_{u=1}^{3}\sum_{k=1}^{l_u}\mathbf{1}(t=T^u_k)R_u^k(x,y)-\beta\log\frac{P_\theta(a_t\mid g_t)}{P_{\theta_{\mathrm{init}}}(a_t\mid g_t)}.

    Here Ruk(x,y)R_u^k(x,y) is the reward for segment kk of reward type uu, and 1\mathbf{1} is the indicator function. Consequently, correctness reward is assigned to the response EOS token, citation-recall reward to each sentence’s final token, and citation-precision reward to each citation’s closing bracket; rewards assigned to the same token are summed. The coefficient β\beta controls the KL penalty. The authors optimize the policy and value models with PPO.

  4. Knowl 4 — ChatGPT distillation initializes the staged training framework

    model/method

    Because open-weight models initially produced citations for few sentences and citation-annotated training answers are costly to collect, the authors first create a supervised starting point by distilling from ChatGPT. They prompt GPT-3.5-turbo-0301 with in-context demonstrations to answer training questions using retrieved passages and citations, then fine-tune LLaMA-2-7B on those generated answers. The model input combines the task instruction, question, and retrieved passages; the resulting model is denoted MdistM_{\mathrm{dist}}.

    The fine-grained rejection-sampling and PPO methods are applied after this initialization. In the combined setting, fine-grained RS is applied first and fine-grained RL further trains the RS model. For comparison, the authors also train with a holistic reward: standard rejection sampling ranks complete responses by the sum of the three rewards, while holistic RL assigns that total reward only to the response’s final token and gives other tokens zero task reward. The staged framework therefore compares both the granularity of feedback and the training strategy.

  5. Knowl 5 — Evaluation and training use three QA datasets with five retrieved passages

    experimental setup

    The main experiments fine-tune LLaMA-2-7B on a mixture of three ALCE attributable-generation datasets: ASQA, a long-form dataset with ambiguous questions; QAMPARI, a factoid dataset whose answers are lists; and ELI5, a long-form question-answering dataset. The mixed training and development sets contain roughly 1,000 and 334 examples from each dataset, respectively, and the test evaluation uses about 1,000 examples per dataset. Each question is paired with the top five retrieved passages: GTR retrieves from the 2018-12-20 Wikipedia snapshot for ASQA and QAMPARI, while BM25 retrieves from Sphere for ELI5.

    For f.g. RS, the LLaMA-2-7B experiments use beam width 8, two sampled continuations per beam sequence, and maximum search depth 5 on ASQA and ELI5 or 10 on QAMPARI. Holistic RS samples 16 responses. The three reward weights are each set to 0.20.2. The separately evaluated out-of-training-domain dataset EXPERTQA contains 2,169 examples, with top-five passages retrieved from Sphere using BM25. The main metrics are correctness recall, citation recall (the fraction of response sentences supported by their cited passages), and citation precision (the fraction of citations that help support their sentences); QAMPARI additionally reports correctness recall capped at five answer hits, denoted Rec.-5.

  6. Knowl 6 — Fine-grained rewards improve citation metrics over holistic rewards

    data/table

    The table reports test-set scores for the combined-training setting on ASQA, QAMPARI, and ELI5. Each metric is a percentage: ASQA uses exact-match correctness recall (EM Rec.); QAMPARI uses correctness Rec.-5; ELI5 uses NLI-based claim recall (Claim Rec.). For each dataset, citation recall (Rec.) and citation precision (Prec.) measure support and citation relevance. The table compares in-context baselines, the distilled model, and holistic (h.) versus fine-grained (f.g.) reward training.

    ASQA QAMPARI ELI5
    Model EM Rec. Cite Rec. Cite Prec. Rec.-5 Cite Rec. Cite Prec. Claim Rec. Cite Rec. Cite Prec.
    ICL ChatGPT 39.96 74.72 70.97 18.34 18.57 17.65 13.47 50.94 47.58
    ICL LLaMA-2-7B 34.15 14.12 15.26 8.24 9.23 7.51 7.83 14.44 8.92
    Distilled MmathrmdistM_{\\mathrm{dist}} 35.56 74.80 67.99 17.26 16.18 18.69 12.03 49.69 45.71
    Holistic RL 34.33 75.77 70.12 17.30 16.44 16.39 11.52 51.77 49.32
    Fine-grained RL 35.99 76.30 72.38 18.39 18.81 17.82 11.60 51.29 51.09
    Holistic RS 37.96 74.86 68.48 14.62 15.21 16.71 11.60 54.10 48.95
    Fine-grained RS 40.07 76.71 74.35 16.14 18.95 18.56 11.67 58.75 55.03
    Holistic RS+RL 37.33 74.86 69.37 15.02 15.67 16.82 11.21 55.62 50.58
    Fine-grained RS+RL 40.05 77.83 76.33 16.65 19.54 19.50 11.54 60.86 60.23

    Fine-grained RS+RL attains the highest citation-recall and citation-precision scores on all three datasets, and the highest ASQA correctness recall among the listed systems. Compared with holistic RS+RL, it improves citation recall/precision from 74.86/69.37 to 77.83/76.33 on ASQA, 15.67/16.82 to 19.54/19.50 on QAMPARI, and 55.62/50.58 to 60.86/60.23 on ELI5. Correctness does not improve on every dataset: its QAMPARI Rec.-5 and ELI5 Claim Rec. remain below the ChatGPT in-context scores.

  7. Knowl 7 — Reward ablations expose a correctness–citation trade-off

    empirical result

    The authors ablate fine-grained rejection sampling in separate, single-dataset training runs. Removing both citation rewards (R2R_2 and R3R_3) favors correctness on some measures but weakens citation quality; removing correctness reward (R1R_1) tends to improve citation scores on long-form datasets while reducing correctness. The reported tuples below are correctness recall, citation recall, citation precision, and response length in tokens; QAMPARI correctness recall is Rec.-5.

    • ASQA: Full fine-grained RS scores 40.24,77.65,74.96,51.3440.24, 77.65, 74.96, 51.34. Without R2R_2 and R3R_3, it scores 41.29,49.51,67.54,56.3441.29, 49.51, 67.54, 56.34; without R1R_1, it scores 39.79,79.42,75.69,55.8939.79, 79.42, 75.69, 55.89.
    • QAMPARI: Full fine-grained RS scores 17.48,20.67,20.62,21.65,11.2417.48, 20.67, 20.62, 21.65, 11.24, where the first three values are correctness Rec.-5, citation recall, and citation precision, followed by the reported answer-length and additional length-related entry in the ablation table. Without R2R_2 and R3R_3, correctness Rec.-5 rises to 23.1223.12, while citation recall and precision fall to 16.7316.73 and 15.2915.29. Without R1R_1, correctness Rec.-5 falls to 12.6212.62, while citation recall and precision are 20.3120.31 and 21.4521.45.
    • ELI5: Full fine-grained RS scores 11.87,61.27,56.45,83.0111.87, 61.27, 56.45, 83.01. Without R2R_2 and R3R_3, the scores are 11.60,41.06,43.38,88.3811.60, 41.06, 43.38, 88.38; without R1R_1, they are 11.17,62.92,58.51,84.2611.17, 62.92, 58.51, 84.26.

    The answer-correctness-only model also tends to produce longer responses, apparently seeking more opportunities to include gold answers. The citation-only setting shows the complementary pattern: citation quality improves on the long-form datasets while correctness recall drops. These ablations support using all three reward signals rather than optimizing a single objective.

  8. Knowl 8 — Fine-grained training transfers citation gains to EXPERTQA

    empirical result

    On EXPERTQA, a held-out dataset of 2,169 questions requiring domain knowledge, the authors evaluate AutoAIS (the percentage of sentences supported by citations), FActscore (factuality), and the total number of generated sentences. Fine-grained training improves citation support over the distilled model and both in-context baselines, while FActscore remains above 83% for all trained variants.

    SystemAutoAISFActscoreTotal sentences
    ICL ChatGPT56.9885.838,145
    ICL LLaMA-2-7B19.6382.7710,053
    Distilled model51.3383.467,293
    Holistic RL53.6483.828,139
    Fine-grained RL56.1583.897,322
    Holistic RS57.8883.478,436
    Fine-grained RS63.4983.857,450
    Holistic RS+RL59.4283.718,012
    Fine-grained RS+RL66.1283.786,256

    Fine-grained RS+RL achieves the highest AutoAIS score, 66.12, while its FActscore is 83.78. The result indicates transfer of citation-attribution performance beyond the three training datasets; the factuality metric is not itself used as a training reward.

  9. Knowl 9 — Observed citation errors mainly involve passage interpretation

    empirical result

    In a manual inspection of responses from the fine-grained RS+RL model, the authors identify three kinds of citation errors. Reported proportions are 5.26% for mixing up passage identifiers, 31.58% for redundant citations, and 63.16% for misinterpreting cited passages. Identifier mix-ups occur when content from one passage is attributed to another. Redundancy occurs when irrelevant passages are cited, including cases where the answer comes from the model’s parametric knowledge because the retrieved passages do not support it.

    The largest category, misinterpretation, includes distortion of a fact from one passage (52.63%) and incorrect synthesis across multiple passages (10.53%), such as asserting a relationship between entities that the passages do not establish. The authors associate single-passage misinterpretation particularly with QAMPARI’s multi-hop questions. These error types show that improving citation relevance alone does not ensure that cited evidence has been read and synthesized correctly.

  10. Knowl 10 — Reward quality and proprietary distillation constrain the method

    limitation

    The correctness reward for ELI5 depends on subclaims generated from reference answers by text-davinci-003. Although the paper reports that more than 90% of these subclaims are faithful to their source answers, the subclaims may be incomplete. As a result, the reward can fail to represent all aspects of the reference and may not align fully with evaluation metrics or actual correctness. Checking completeness would require substantial human annotation.

    The training framework also requires an initial distillation stage using ChatGPT-generated answers, which can limit access where a capable proprietary model is unavailable. The authors suggest bootstrapping high-quality responses with in-context learning and beam-search sampling, then applying behavioral cloning and reinforcement learning, as a possible direction rather than a demonstrated solution. They additionally note that large retrieval corpora can introduce noisy or biased content into model responses.

Coverage note — The separate T5-large experiments, detailed prompt examples, and auxiliary fluency analyses are omitted because they are secondary to the paper’s main LLaMA-2-7B methods and attributable-generation results.

References

  1. 1.Akari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil, and Hannaneh Hajishirzi. 2023. Self-rag: Learning to retrieve, generate, and critique through self-reflection. arXiv preprint arXiv:2310.11511, 2023.
  2. 2.Bernd Bohnet, Vinh Q. Tran, Pat Verga, Roee Aharoni, Daniel Andor, Livio Baldini Soares, Massimiliano Ciaramita, Jacob Eisenstein, Kuzman Ganchev, Jonathan Herzig, Kai Hui, Tom Kwiatkowski, Ji Ma, Jianmo Ni, Lierni Sestorain Saralegui, Tal Schuster, William W. Cohen, Michael Collins, Dipanjan Das, Donald Metzler, Slav Petrov, and Kellie Webster. 2023. Attributed question answering: Evaluation and modeling for attributed large language models. arXiv preprint arXiv:2212.08037.
  3. 3.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Daniel Ziegler Aditya Ramesh, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems.
  4. 4.Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2023. Qlora: Efficient finetuning of quantized llms. arXiv preprint arXiv:2305.14314.
  5. 5.Angela Fan, Yacine Jernite, Ethan Perez, David Grangier, Jason Weston, and Michael Auli. 2019. Eli5: Long form question answering. In Association for Computational Linguistics (ACL), pages 3558—-3567.
  6. 6.Luyu Gao, Zhuyun Dai, Panupong Pasupat, Anthony Chen, Arun Tejasvi Chaganty, Yicheng Fan, Vincent Zhao, Ni Lao, Hongrae Lee, Da-Cheng Juan, and Kelvin Guu. 2023a. Rarr: Researching and revising what language models say, using language models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), page 16477–16508.
  7. 7.Tianyu Gao, Howard Yen, Jiatong Yu, and Danqi Chen. 2023b. Enabling large language models to generate text with citations. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 6465—-6488.
  8. 8.Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang. 2020. Realm: Retrieval augmented language model pre-training. In Proceedings of International Conference on Machine Learning, 2020, pages 33–40.
  9. 9.Hangfeng He, Hongming Zhang, and Dan Roth. 2022. Rethinking with retrieval: Faithful large language model inference. arXiv preprint arXiv:2301.00303, 2022.
  10. 10.Or Honovich, Roee Aharoni, Jonathan Herzig, Hagai Taitelbaum, Doron Kukliansy, Vered Cohen, Thomas Scialom, Idan Szpektor, Avinatan Hassidim, and Yossi Matias. 2022. True: Re-evaluating factual consistency evaluation. In North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT), pages 3905–3920.
  11. 11.Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685, 2021.
  12. 12.Jie Huang and Kevin Chen-Chuan Chang. 2023. Citation: A key to building responsible and accountable large language models. arXiv preprint arXiv:2307.02185.
  13. 13.Gautier Izacard, Patrick Lewis, Maria Lomeli, Lucas Hosseini, Fabio Petroni, Timo Schick, Jane Dwivedi-Yu, Armand Joulin, Sebastian Riedel, and Edouard Grave. 2022. Atlas: Few-shot learning with retrieval augmented language models. arXiv preprint arXiv:2208.03299, 2022.
  14. 14.Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2022. Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12):1––38.
  15. 15.Zhengbao Jiang, Frank F Xu, Luyu Gao, Zhiqing Sun, Qian Liu, Jane Dwivedi-Yu, and Yiming Yang. 2023. Active retrieval augmented generation. arXiv preprint arXiv:2305.06983, 2023.
  16. 16.Ehsan Kamalloo, Aref Jafari, Xinyu Zhang, Nandan Thakur, and Jimmy Lin. 2023. HAGRID: A human-llm collaborative dataset for generative information-seeking with attribution. arXiv:2307.16883.
  17. 17.Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis. 2020. Generalization through memorization: Nearest neighbor language models. In International Conference on Learning Representations.
  18. 18.Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. 2023. Efficient memory management for large language model serving with pagedattention. In Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles.
  19. 19.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Kuttler, Mike Lewis, Wen tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. In Advances in Neural Information Processing Systems.
  20. 20.Dongfang Li, Zetian Sun, Xinshuo Hu, Zhenyu Liu, Ziyang Chen, Baotian Hu, Aiguo Wu, and Min Zhang. 2023a. A survey of large language models attribution. arXiv:2311.03731.
  21. 21.Xinze Li, Yixin Cao, Liangming Pan, Yubo Ma, and Aixin Sun. 2023b. Towards verifiable generation: A benchmark for knowledge-aware language model attribution. arXiv preprint arXiv:2310.05634.
  22. 22.Xi Victoria Lin, Xilun Chen, Mingda Chen, Weijia Shi, Maria Lomeli, Rich James, Pedro Rodriguez, Jacob Kahn, Gergely Szilvasy, Mike Lewis, Luke Zettlemoyer, and Scott Yih. 2022. Ra-dit: Retrievalaugmented dual instruction tuning. arXiv preprint arXiv:2310.01352, 2022.
  23. 23.Nelson F. Liu, Tianyi Zhang, and Percy Liang. 2023. Evaluating verifiability in generative search engines. arXiv preprint arXiv:2304.09848.
  24. 24.Hongyin Luo, Yung-Sung Chuang, Yuan Gong, Tianhua Zhang, Yoon Kim, Xixin Wu, Danny Fox, Helen Meng, and James Glass. 2023. Sail: Search-augmented instruction learning. arXiv preprint arXiv:2305.15225, 2023.
  25. 25.Chaitanya Malaviya, Subin Lee, Sihao Chen, Elizabeth Sieber, Mark Yatskar, and Dan Roth. 2023. Expertqa: Expert-curated questions and attributed answers. arXiv preprint arXiv:2309.07852, 2023.
  26. 26.Jacob Menick, Maja Trebacz, Vladimir Mikulik, John Aslanides, Francis Song, Martin Chadwick, Mia Glaese, Susannah Young, Lucy Campbell-Gillingam, Geoffrey Irving, and Nat McAleese. 2022. Teaching language models to support answers with verified quotes. arXiv preprint arXiv:2203.11147, 2022.
  27. 27.Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen tau Yih, Pang Wei Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2023. Factscore: Fine-grained atomic evaluation of factual precision in long form text generation. arXiv preprint arXiv:2305.14251v1.
  28. 28.Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, Xu Jiang, Karl Cobbe, Tyna Eloundou, Gretchen Krueger, Kevin Button, Matthew Knight, Benjamin Chess, and John Schulman. 2021. Webgpt: Browser-assisted question answering with human feedback. arXiv preprint arXiv:2112.09332, 2021.
  29. 29.Jianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai, Gustavo Hernandez Abrego, Ji Ma, Vincent Zhao, Yi Luan, Keith Hall, Ming-Wei Chang, and Yinfei Yang. 2022. Large dual encoders are generalizable retrievers. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 9844—-9855.
  30. 30.Aaron Parisi, Yao Zhao, and Noah Fiedel. 2022. Talm: Tool augmented language models. arXiv preprint arXiv:2205.12255, 2022.
  31. 31.Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Dmytro Okhonko, Samuel Broscheit, Gautier Izacard, Patrick Lewis, Barlas Oguz, Edouard Grave, Wen ˘ tau Yih, and et al. 2021. The web is your oyster - knowledge-intensive nlp against a very large web corpus. arXiv preprint arXiv:2112.09924, 2021.
  32. 32.Krishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun, Sean Welleck, Yejin Choi, and Zaid Harchaoui. 2021. Mauve: Measuring the gap between neural text and human text using divergence frontiers. In Advances in Neural Information Processing Systems.
  33. 33.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research (JMLR), 21(140).
  34. 34.Ori Ram, Yoav Levine, Itay Dalmedigos, Dor Muhlgay, Amnon Shashua, Kevin Leyton-Brown, and Yoav Shoham. 2023. In-context retrieval-augmented language models. In Transactions of the Association for Computational Linguistics, 2023.
  35. 35.Samuel Joseph Amouyal Ohad Rubin, Ori Yoran, Tomer Wolfson, Jonathan Herzig, and Jonathan Berant. 2022. Qampari: An open-domain question answering benchmark for questions with many answers from multiple paragraphs. arXiv preprint arXiv:2205.12665, 2022.
  36. 36.Michael Santacroce, Yadong Lu, Han Yu, Yuanzhi Li, and Yelong Shen. 2023. Efficient rlhf: Reducing the memory usage of ppo. arXiv preprint arXiv:2309.00754, 2023.
  37. 37.Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2023. Toolformer: Language models can teach themselves to use tools. arXiv preprint arXiv:2302.04761, 2023.
  38. 38.John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017.
  39. 39.Ivan Stelmakh, Yi Luan, Bhuwan Dhingra, and Ming-Wei Chang. 2022. Asqa: Factoid questions meet long-form answers. arXiv preprint arXiv:2204.06092, 2022.
  40. 40.Haitian Sun, William Cohen, and Ruslan Salakhutdinov. 2022. Conditionalqa: A complex reading comprehension dataset with conditional answers. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics, page 3627–3637.
  41. 41.Hao Sun, Hengyi Cai, Bo Wang, Yingyan Hou, Xiaochi Wei, Shuaiqiang Wang, Yan Zhang, and Dawei Yin. 2023. Towards verifiable text generation with evolving memory and self-reflection. arXiv preprint arXiv:2312.09075.
  42. 42.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, and et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023.
  43. 43.Zeqiu Wu, Yushi Hu, Weijia Shi, Nouha Dziri, Alane Suhr, Prithviraj Ammanabrolu, Noah A Smith, Mari Ostendorf, and Hannaneh Hajishirzi. 2023. Fine-grained human feedback gives better rewards for language model training. In Advances in Neural Information Processing Systems.
  44. 44.Xi Ye, Ruoxi Sun, Sercan Ö. Arik, and Tomas Pfister. 2023. Effective large language model adaptation for improved grounding. arXiv preprint arXiv:2311.09533.
  45. 45.Ori Yoran, Tomer Wolfson, Ori Ram, and Jonathan Berant. 2023. Making retrieval-augmented language models robust to irrelevant context. arXiv preprint arXiv:2310.01558, 2023.
  46. 46.Xiang Yue, Boshi Wang, Kai Zhang, Ziru Chen, Yu Su, and Huan Sun. 2023. Automatic evaluation of attribution by large language models. arXiv preprint arXiv:2305.06311.
  47. 47.Zexuan Zhong, Tao Lei, and Danqi Chen. 2022. Training language models with memory augmentation. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 5657–5673.
  48. 48.Andy Zhou, Kai Yan, Michal Shlapentokh-Rothman, Haohan Wang, and Yu-Xiong Wang. 2023. Language agent tree search unifies reasoning acting and planning in language models. arXiv preprint arXiv:2310.04406, 2023.

Citation

MLA
Huang, C., et al. “Training Language Models to Generate Text with Citations via Fine-grained Rewards”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 2926–49, https://doi.org/10.18653/v1/2024.acl-long.161.
APA
Huang, C., Wu, Z., Hu, Y., & Wang, W. (2024). Training Language Models to Generate Text with Citations via Fine-grained Rewards. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2926–2949. https://doi.org/10.18653/v1/2024.acl-long.161
Chicago
Huang, C., Z. Wu, Y. Hu, and W. Wang. 2024. “Training Language Models to Generate Text with Citations via Fine-grained Rewards”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2926–49. https://doi.org/10.18653/v1/2024.acl-long.161.
Harvard
Huang, C. et al. (2024) “Training Language Models to Generate Text with Citations via Fine-grained Rewards”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 2926–2949. Available at: https://doi.org/10.18653/v1/2024.acl-long.161.
Vancouver
1. Huang C, Wu Z, Hu Y, Wang W (2024) Training Language Models to Generate Text with Citations via Fine-grained Rewards. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 2926–2949

BibTeX

@inproceedings{huang-etal-2024-training,
    title = "Training Language Models to Generate Text with Citations via Fine-grained Rewards",
    author = "Huang, Chengyu  and
      Wu, Zeqiu  and
      Hu, Yushi  and
      Wang, Wenya",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.161/",
    doi = "10.18653/v1/2024.acl-long.161",
    pages = "2926--2949"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/