Learning to summarize from human feedback

Nisan StiennonLong OuyangJeff WuDaniel M. ZieglerRyan J. LoweChelsea VossAlec RadfordDario AmodeiPaul Christiano

article2020NeurIPS3,475 citations

Demonstrates that fine-tuning language models with reinforcement learning on human preference comparisons produces summaries preferred by human evaluators over those generated by standard supervised learning or human reference authors.

Listen

The paper describes a method for training language models to produce higher-quality summaries by directly optimizing for human preferences rather than using proxy metrics. Current approaches to summarization rely on supervised fine-tuning to match human-written reference summaries and automatic evaluation with metrics such as ROUGE. These proxies often diverge from what people actually judge as good summaries, leading to factual errors, poor coverage, and other problems that are hard to fix with maximum-likelihood training alone.

The work set out to close that gap by collecting human judgments of summary quality, learning a model that predicts those judgments, and then using the resulting model as a reward signal for reinforcement learning. Researchers gathered more than 64,000 pairwise comparisons of summaries on a filtered version of the Reddit TL;DR dataset. They trained reward models ranging from 1.3 billion to 6.7 billion parameters to predict which summary a human would prefer. They then fine-tuned GPT-style policy models with proximal policy optimization, adding a KL penalty to keep outputs close to the supervised baseline. The process was repeated in batches, and the same reward models were later used to evaluate transfer to the CNN/Daily Mail news domain without further fine-tuning.

Human feedback policies produced summaries that labelers preferred to both the original reference summaries and to much larger models trained only with supervised learning. A 1.3 billion parameter feedback model outperformed a 10-times-larger supervised model, and the 6.7 billion parameter feedback model was rated highest overall. On CNN/Daily Mail the same Reddit-trained feedback models generated fluent, high-coverage summaries that nearly matched the quality of models fine-tuned directly on news data. The learned reward models agreed with human raters at rates comparable to inter-rater agreement and outperformed ROUGE at predicting preferences. Scaling both model size and the amount of comparison data improved reward-model accuracy, while excessive optimization against a fixed reward model eventually produced worse summaries.

These results show that reward modeling from human feedback can measurably improve output quality on a subjective generation task and can generalize across domains. The approach offers a practical route to aligning model behavior with nuanced human criteria that are difficult to capture in hand-crafted loss functions. It also demonstrates that such alignment remains feasible at the scale of current large language models.

Further work is needed to reduce the cost of data collection, to test whether similar gains appear on longer or more open-ended tasks, and to develop safeguards against reward hacking. The released dataset of comparisons and the inference code for the 1.3 billion parameter models provide a starting point for that research. The main limitations are the expense of high-quality human labels and the risk that over-optimization against an imperfect reward model can degrade performance once a certain threshold is crossed.

  • Paper: Deep reinforcement learning from human preferences, Paul F. Christiano et al. (2017). Reading Christiano et al. (2017) first is essential because it introduces the foundational framework of training deep reinforcement learning agents directly from human preference comparisons rather than fixed programmatic rewards.
  • Paper: Fine-Tuning Language Models from Human Preferences, Daniel M. Ziegler et al. (2019). Ziegler et al. (2019) directly precedes this paper by first demonstrating how to apply reinforcement learning from human preferences to fine-tune language models for summarization.
  • Paper: ROUGE: A Package for Automatic Evaluation of Summaries, Chin-Yew Lin (2004). Familiarity with the ROUGE evaluation package is required to understand the metric limitations that motivated shifting from automatic proxies to human preference optimization.
Cover for Learning to summarize from human feedback

Abstract

As language models become more powerful, training and evaluation are increasingly bottlenecked by the data and metrics used for a particular task. For example, summarization models are often trained to predict human reference summaries and evaluated using ROUGE, but both of these metrics are rough proxies for what we really care about -- summary quality. In this work, we show that it is possible to significantly improve summary quality by training a model to optimize for human preferences. We collect a large, high-quality dataset of human comparisons between summaries, train a model to predict the human-preferred summary, and use that model as a reward function to fine-tune a summarization policy using reinforcement learning. We apply our method to a version of the TL;DR dataset of Reddit posts and find that our models significantly outperform both human reference summaries and much larger models fine-tuned with supervised learning alone. Our models also transfer to CNN/DM news articles, producing summaries nearly as good as the human reference without any news-specific fine-tuning. We conduct extensive analyses to understand our human feedback dataset and fine-tuned models We establish that our reward model generalizes to new datasets, and that optimizing our reward model results in better summaries than optimizing ROUGE according to humans. We hope the evidence from our paper motivates machine learning researchers to pay closer attention to how their training loss affects the model behavior they actually want.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 Method and experiment details
  • 3.1 High-level methodology
  • 3.2 Datasets and task
  • 3.3 Collecting human feedback
  • 3.4 Models
  • 4 Results
  • 4.1 Summarizing Reddit posts from human feedback
  • 4.2 Transfer to summarizing news articles
  • 4.3 Understanding the reward model
  • 4.4 Analyzing automatic metrics for summarization
  • 5 Discussion
  • Acknowledgements
  • References
  • Table of Contents
  • A TL;DR dataset details
  • B Further model training details
  • B.1 Hyperparameters
  • B.2 Input format
  • C Human data collection details
  • C.1 Process for ensuring high-quality human data
  • C.2 Assessing human feedback quality
  • C.3 Labeler demographics
  • C.4 Labeler website
  • C.5 Instructions for labelers
  • C.6 Composition of the labeled dataset
  • **Coherence**
  • **Rubric:**
  • **Accuracy**
  • **Rubric:**
  • **Coverage**
  • **Rubric:**
  • **Overall quality**
  • **Rubric:**
  • C.7 Example comparison tasks
  • D Choice of baselines
  • E CNN/DM lead-3 vs reference summaries
  • F Controlling for summary length
  • G Additional results
  • G.1 Value function ablation
  • G.2 Evaluating policies along axes of quality
  • G.3 Studying best-of-N optimization
  • G.4 ROUGE scores
  • G.5 Bigram overlap statistics
  • G.6 Reward model validation sets
  • G.7 Measuring agreement between different evaluation metrics
  • H Samples
  • H.1 Random samples
  • H.2 Overoptimized samples

Knowls

  1. Knowl 1 — Three-Phase RLHF Workflow for Abstractive Text Summarization

    model/method

    The reinforcement learning from human feedback (RLHF) framework for text summarization consists of three sequential, iterative phases using autoregressive Transformer decoder language models:

    1. Data Collection from Multiple Policies: For each input post xx, candidate summaries y0,y1y_0, y_1 are generated by various sources (e.g., supervised baselines, earlier policy iterations, or reference summaries). Human labelers evaluate each pair (y0,y1)(y_0, y_1) to select the higher-quality summary yiy_i.

    2. Reward Model Training: A supervised reward model rθ(x,y)r_\theta(x, y) is initialized from a supervised fine-tuned checkpoint by appending a randomly initialized scalar linear head. It is optimized to predict human pairwise preferences by minimizing binary cross-entropy loss over comparison pairs.

    3. Policy Optimization via Reinforcement Learning: A policy πϕRL\pi_\phi^{\text{RL}} (initialized from a supervised fine-tuned model πSFT\pi^{\text{SFT}}) generates summaries token-by-token. The policy is updated using Proximal Policy Optimization (PPO) to maximize the scalar score assigned by rθ(x,y)r_\theta(x, y) to the completed summary, subject to a per-token Kullback-Leibler (KL) divergence penalty against πSFT\pi^{\text{SFT}} to prevent policy collapse and distributional drift away from the reward model's training distribution.

  2. Knowl 2 — Reward Model Objective for Pairwise Human Preferences

    equation

    Given a prompt xx and a pair of candidate summaries (y0,y1)(y_0, y_1) where a human evaluator prefers summary yiy_i (i{0,1}i \in \{0, 1\}), the reward model parameters θ\theta are trained by minimizing the cross-entropy loss:

    L(rθ)=E(x,y0,y1,i)D[logσ(rθ(x,yi)rθ(x,y1i))]\mathcal{L}(r_\theta) = -\mathbb{E}_{(x, y_0, y_1, i) \sim \mathcal{D}}\left[\log \sigma\left(r_\theta(x, y_i) - r_\theta(x, y_{1-i})\right)\right]

    where:

    • D\mathcal{D} denotes the dataset of human pairwise comparisons,
    • rθ(x,y)Rr_\theta(x, y) \in \mathbb{R} is the scalar output of the reward model for prompt xx and candidate summary yy,
    • σ(z)=11+ez\sigma(z) = \frac{1}{1 + e^{-z}} represents the standard logistic sigmoid function.

    At the end of training, reward values are shifted by an additive constant such that human reference summaries in the dataset achieve an empirical mean reward score of 00.

  3. Knowl 3 — Policy Reinforcement Learning Objective with KL Divergence Penalty

    equation

    When optimizing a summarization policy πϕRL\pi_\phi^{\text{RL}} with parameters ϕ\phi using Proximal Policy Optimization (PPO), the scalar reward R(x,y)R(x, y) assigned to a full generated summary yy for input text xx is defined as:

    R(x,y)=rθ(x,y)βlog[πϕRL(yx)πSFT(yx)]R(x, y) = r_\theta(x, y) - \beta \log \left[ \frac{\pi_\phi^{\text{RL}}(y \mid x)}{\pi^{\text{SFT}}(y \mid x)} \right]

    where:

    • rθ(x,y)r_\theta(x, y) is the scalar score produced by the learned reward model,
    • πSFT(yx)\pi^{\text{SFT}}(y \mid x) is the conditional probability of summary yy given xx under the initial supervised fine-tuned model,
    • πϕRL(yx)\pi_\phi^{\text{RL}}(y \mid x) is the probability under the current RL policy,
    • β0\beta \ge 0 is the KL penalty coefficient (set to β=0.05\beta = 0.05 in main 1.3B and 6.7B parameter training runs).

    In the episodic formulation, each generation step corresponds to predicting a Byte-Pair Encoding (BPE) token, intermediate rewards are 00, the episode terminates on the end-of-sequence (EOS) token, and the discount factor is γ=1\gamma = 1.

  4. Knowl 4 — Human Preference Superiority of RLHF Policies on Reddit TL;DR

    empirical result

    Policies trained with human feedback on the Reddit TL;DR dataset outperform supervised learning baselines:

    • A 1.3 billion parameter RLHF policy achieves a 61% human preference rate against ground-truth reference summaries, significantly outperforming a 10-times larger supervised baseline with 13 billion parameters (43% preference against reference summaries).
    • A 6.7 billion parameter RLHF policy achieves an unadjusted raw preference rate of ~70% over reference summaries. When controlling for summary length differences using logistic regression, the 6.7B RLHF model remains preferred over human reference summaries ~65% of the time.
    • On a 7-point Likert scale evaluation of overall summary quality, the 6.7B RLHF model receives a perfect score of 7/7 in 45% of cases, compared to 20% for the 6.7B supervised baseline and 23% for human-written reference summaries. The largest improvements over baselines occur along the coverage dimension.
  5. Knowl 5 — Zero-Shot Domain Transfer of Summarization Models to CNN/DailyMail

    empirical result

    Models trained via RLHF solely on the Reddit TL;DR dataset generalize zero-shot to summarize CNN/DailyMail (CNN/DM) news articles without any news-specific fine-tuning:

    • On a 1–7 Likert overall quality scale, the 6.7B parameter TL;DR RLHF model attains an average score of 5.25±0.875.25 \pm 0.87, compared to 4.41±0.384.41 \pm 0.38 for the 6.7B supervised TL;DR transfer baseline and 4.60±0.304.60 \pm 0.30 for a 6.7B zero-shot pretrained model.
    • The zero-shot 6.7B RLHF model nearly matches the overall quality of human reference summaries (5.545.54) and an in-domain 6.7B model fine-tuned supervised on CNN/DM (5.405.40), despite generating summaries that are substantially shorter (mean length of 175 characters for RLHF vs. 300–316 characters for supervised CNN/DM models and T5).
  6. Knowl 6 — Scaling Trends of Reward Model Accuracy with Dataset and Model Size

    empirical result

    In controlled ablations training 7 reward models spanning 160 million to 13 billion parameters across dataset sizes from 8,000 to 64,832 human comparison pairs:

    • Doubling the quantity of human comparison training data leads to an average increase of approximately 1.1%1.1\% in reward model pairwise validation accuracy.
    • Doubling the parameter size of the reward model leads to an average increase of approximately 1.8%1.8\% in validation accuracy.
    • The 6.7 billion parameter reward model trained on the full 64k dataset achieves a validation accuracy that approaches the inter-annotator agreement rate of a single human labeler.
  7. Knowl 7 — Reward Model Over-Optimization and True Preference Degradation

    empirical result

    When optimizing a policy against a fixed reward model, continuing optimization past an optimal point leads to policy over-optimization (Goodhart's Law):

    • As the KL divergence between the RL policy and the initial supervised baseline πSFT\pi^{\text{SFT}} increases from KL=0\text{KL} = 0 to KL1020\text{KL} \approx 10\text{--}20, predicted reward and human-assessed summary quality both improve.
    • As optimization continues further (e.g., KL>50\text{KL} > 50 obtained by reducing the KL penalty coefficient β\beta), predicted reward continues to increase while actual human preference peaks and declines sharply, eventually becoming anti-correlated with true human preferences.
    • In best-of-NN rejection sampling experiments, optimizing against ROUGE degrades human-evaluated quality significantly earlier and peaks at a much lower preference rate than optimizing against learned reward models.
  8. Knowl 8 — Decoupled Value Function and Policy Networks in Transformer PPO

    model/method

    In Transformer-based PPO for text generation, using a separate, independent network architecture for the value function rather than sharing parameters with the policy network prevents policy degradation:

    • Architecture Separation: The policy πϕRL\pi_\phi^{\text{RL}} is initialized from the supervised fine-tuned model πSFT\pi^{\text{SFT}}, whereas the value function VψV_\psi is initialized from the learned reward model rθr_\theta.
    • Preventing Representation Destruction: In shared parameter setups, aggressive value function gradient updates early in reinforcement learning training degrade and partially destroy pretrained policy representations, leading to lower asymptotic total reward.
  9. Knowl 9 — Breakdown of ROUGE and Supervised Log-Likelihood as Proxies for Generation Quality

    empirical result

    Automatic metrics such as ROUGE and supervised model log-likelihood fail to correlate with human preference when comparing outputs from strong generative policies:

    • When evaluating pairwise comparisons between samples from a 1.3B supervised baseline, ROUGE agrees with human labelers 57.7%±1.3%57.7\% \pm 1.3\% of the time. When evaluating comparisons between samples from a 6.7B RLHF policy, ROUGE agreement drops to chance level (49.9%±2.1%49.9\% \pm 2.1\%).
    • Similarly, agreement between supervised model log-likelihood and human judgments falls to 48.0%\le 48.0\% on RLHF policy samples, and scaling supervised model size from 1.3B to 6.7B does not improve agreement.
    • In contrast, a 6.7B learned reward model achieves 62.3%±2.1%62.3\% \pm 2.1\% agreement with individual labelers and 75.0%±9.8%75.0\% \pm 9.8\% agreement with a 3-labeler ensemble on RLHF policy comparisons.
  10. Knowl 10 — TL;DR Summarization Dataset Filtering and Length Control

    experimental setup

    To train summarization models without length confounding and noisy demonstrations, the raw Reddit TL;DR dataset (~3 million posts) is curated as follows:

    1. Exact duplicate post bodies (~20,000 instances) and posts from subreddits outside a curated whitelist of general-audience subreddits (dominated by relationship and advice topics) are removed.
    2. Posts whose bodies exceed 512 BPE tokens or contain flagged sensitive topics are removed, yielding 287,790 prompt posts used for RL training.
    3. Reference summaries starting with edit markers (e.g., 'Edit', 'Update', 'P.S.') or high profanity are filtered out.
    4. Reference summary length is strictly constrained to between 24 and 48 tokens. This eliminates low-quality short summaries (<16 tokens) and ensures length distribution overlap with RL policies for length-controlled evaluations, resulting in a final supervised dataset of 123,169 post-summary pairs.
  11. Knowl 11 — Protocol for Human Feedback Collection and Quality Control

    experimental setup

    Human feedback data is collected using an offline, batch-based protocol with extensive labeler calibration:

    1. Naïve Interpretations: Before viewing the source Reddit post, labelers write their interpretation of candidate summaries. This exposes ambiguities in the summary that would otherwise be masked by reading the original context first.
    2. Confidence-Weighted Comparisons: Labelers evaluate summary pairs on a 9-point confidence scale (ranging from high confidence that summary A is better to high confidence that summary B is better).
    3. Quality and Calibration Monitoring: Approximately 10%–20% of comparison questions are sampled from a shared pool to track inter-labeler agreement. Continuous communication via office hours and shared chat ensures calibration.
    4. Agreement Rates: On supervised model comparisons, human labelers achieve an agreement rate of 77%±2%77\% \pm 2\% with researchers, matching the researcher-researcher inter-annotator agreement rate (73%±4%73\% \pm 4\%).

Coverage note — No substantial contributed material was omitted. All key algorithmic components (PPO formulation with KL penalty, reward modeling loss, decoupled value network), dataset filtering procedures, human validation protocols, and primary empirical findings (TL;DR preference gains, CNN/DM transfer, scaling laws, over-optimization, and metric agreement breakdowns) are covered.

References

  1. 1.D. Bahdanau, P. Brakel, K. Xu, A. Goyal, R. Lowe, J. Pineau, A. Courville, and Y. Bengio. An actor-critic algorithm for sequence prediction. arXiv preprint arXiv:1607.07086, 2016.
  2. 2.B. T. Bartell, G. W. Cottrell, and R. K. Belew. Automatic combination of multiple ranked retrieval systems. In SIGIR’94, pages 173–181. Springer, 1994.
  3. 3.F. Böhm, Y. Gao, C. M. Meyer, O. Shapira, I. Dagan, and I. Gurevych. Better rewards yield better summaries: Learning to summarise without references. arXiv preprint arXiv:1909.01214, 2019.
  4. 4.T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei. Language models are few-shot learners. 2020.
  5. 5.S. Cabi, S. Gómez Colmenarejo, A. Novikov, K. Konyushkova, S. Reed, R. Jeong, K. Zolna, Y. Aytar, D. Budden, M. Vecerik, et al. Scaling data-driven robotics with reward sketching and batch reinforcement learning. arXiv, pages arXiv–1909, 2019.
  6. 6.A. T. Chaganty, S. Mussman, and P. Liang. The price of debiasing automatic metrics in natural language evaluation. arXiv preprint arXiv:1807.02202, 2018.
  7. 7.W. S. Cho, P. Zhang, Y. Zhang, X. Li, M. Galley, C. Brockett, M. Wang, and J. Gao. Towards coherent and cohesive long-form text generation. arXiv preprint arXiv:1811.00511, 2018.
  8. 8.S. Chopra, M. Auli, and A. M. Rush. Abstractive sentence summarization with attentive recurrent neural networks. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 93–98, 2016.
  9. 9.P. Christiano, B. Shlegeris, and D. Amodei. Supervising strong learners by amplifying weak experts. arXiv preprint arXiv:1810.08575, 2018.
  10. 10.P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei. Deep reinforcement learning from human preferences. In Advances in Neural Information Processing Systems, pages 4299–4307, 2017.
  11. 11.P. Covington, J. Adams, and E. Sargin. Deep neural networks for youtube recommendations. In Proceedings of the 10th ACM conference on recommender systems, pages 191–198, 2016.
  12. 12.A. M. Dai and Q. V. Le. Semi-supervised sequence learning. In Advances in neural information processing systems, pages 3079–3087, 2015.
  13. 13.J. Dodge, G. Ilharco, R. Schwartz, A. Farhadi, H. Hajishirzi, and N. Smith. Fine-tuning pretrained language models: Weight initializations, data orders, and early stopping. arXiv preprint arXiv:2002.06305, 2020.
  14. 14.L. Dong, N. Yang, W. Wang, F. Wei, X. Liu, Y. Wang, J. Gao, M. Zhou, and H.-W. Hon. Unified language model pre-training for natural language understanding and generation. In Advances in Neural Information Processing Systems, 2019.
  15. 15.Y. Dong, Y. Shen, E. Crawford, H. van Hoof, and J. C. K. Cheung. Banditsum: Extractive summarization as a contextual bandit. arXiv preprint arXiv:1809.09672, 2018.
  16. 16.B. Dorr, D. Zajic, and R. Schwartz. Hedge trimmer: A parse-and-trim approach to headline generation. In Proceedings of the HLT-NAACL 03 on Text summarization workshop-Volume 5, pages 1–8. Association for Computational Linguistics, 2003.
  17. 17.S. Fidler et al. Teaching machines to describe images with natural language feedback. In Advances in Neural Information Processing Systems, pages 5068–5078, 2017.
  18. 18.N. Fuhr. Optimum polynomial retrieval functions based on the probability ranking principle. ACM Transactions on Information Systems (TOIS), 7(3):183–204, 1989.
  19. 19.Y. Gao, C. M. Meyer, M. Mesgar, and I. Gurevych. Reward learning for efficient reinforcement learning in extractive document summarisation. arXiv preprint arXiv:1907.12894, 2019.
  20. 20.X. Glorot and Y. Bengio. Understanding the difficulty of training deep feedforward neural networks. In Proceedings of the thirteenth international conference on artificial intelligence and statistics, pages 249–256, 2010.
  21. 21.B. Hancock, A. Bordes, P.-E. Mazare, and J. Weston. Learning from dialogue after deployment: Feed yourself, chatbot! arXiv preprint arXiv:1901.05415, 2019.
  22. 22.K. M. Hermann, T. Kocisky, E. Grefenstette, L. Espeholt, W. Kay, M. Suleyman, and P. Blunsom. Teaching machines to read and comprehend. In Advances in neural information processing systems, pages 1693–1701, 2015.
  23. 23.A. Holtzman, J. Buys, L. Du, M. Forbes, and Y. Choi. The curious case of neural text degeneration. arXiv preprint arXiv:1904.09751, 2019.
  24. 24.B. Ibarz, J. Leike, T. Pohlen, G. Irving, S. Legg, and D. Amodei. Reward learning from human preferences and demonstrations in atari. In Advances in neural information processing systems, pages 8011–8023, 2018.
  25. 25.N. Jaques, A. Ghandeharioun, J. H. Shen, C. Ferguson, A. Lapedriza, N. Jones, S. Gu, and R. Picard. Way off-policy batch deep reinforcement learning of implicit human preferences in dialog. arXiv preprint arXiv:1907.00456, 2019.
  26. 26.N. Jaques, S. Gu, D. Bahdanau, J. M. Hernández-Lobato, R. E. Turner, and D. Eck. Sequence tutor: Conservative fine-tuning of sequence generation models with kl-control. In International Conference on Machine Learning, pages 1645–1654. PMLR, 2017.
  27. 27.N. Jaques, S. Gu, R. E. Turner, and D. Eck. Tuning recurrent neural networks with reinforcement learning. 2017.
  28. 28.H. J. Jeon, S. Milli, and A. D. Dragan. Reward-rational (implicit) choice: A unifying formalism for reward learning. arXiv preprint arXiv:2002.04833, 2020.
  29. 29.T. Joachims. Optimizing search engines using clickthrough data. In Proceedings of the eighth ACM SIGKDD international conference on Knowledge discovery and data mining, pages 133–142, 2002.
  30. 30.T. Joachims, L. Granka, B. Pan, H. Hembrooke, and G. Gay. Accurately interpreting click-through data as implicit feedback. In ACM SIGIR Forum, volume 51, pages 4–11. Acm New York, NY, USA, 2005.
  31. 31.D. P. Kingma and J. Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  32. 32.J. Kreutzer, S. Khadivi, E. Matusov, and S. Riezler. Can neural machine translation be improved with user feedback? arXiv preprint arXiv:1804.05958, 2018.
  33. 33.W. Kryscinski, N. S. Keskar, B. McCann, C. Xiong, and R. Socher. Neural text summarization: A critical evaluation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 540–551, 2019.
  34. 34.C. Lawrence and S. Riezler. Improving a neural semantic parser by counterfactual learning from human bandit feedback. arXiv preprint arXiv:1805.01252, 2018.
  35. 35.J. Leike, D. Krueger, T. Everitt, M. Martic, V. Maini, and S. Legg. Scalable agent alignment via reward modeling: a research direction. arXiv preprint arXiv:1811.07871, 2018.
  36. 36.M. Lewis, Y. Liu, N. Goyal, M. Ghazvininejad, A. Mohamed, O. Levy, V. Stoyanov, and L. Zettlemoyer. Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. arXiv preprint arXiv:1910.13461, 2019.
  37. 37.M. Li, J. Weston, and S. Roller. Acute-eval: Improved dialogue evaluation with optimized questions and multi-turn comparisons. arXiv preprint arXiv:1909.03087, 2019.
  38. 38.R. Likert. A technique for the measurement of attitudes. Archives of psychology, 1932.
  39. 39.C.-Y. Lin and F. J. Och. Automatic evaluation of machine translation quality using longest common subsequence and skip-bigram statistics. In Proceedings of the 42nd Annual Meeting on Association for Computational Linguistics, page 605. Association for Computational Linguistics, 2004.
  40. 40.T.-Y. Liu. Learning to rank for information retrieval. Springer Science & Business Media, 2011.
  41. 41.J. Maynez, S. Narayan, B. Bohnet, and R. McDonald. On faithfulness and factuality in abstractive summarization, 2020.
  42. 42.B. McCann, N. S. Keskar, C. Xiong, and R. Socher. The natural language decathlon: Multitask learning as question answering. arXiv preprint arXiv:1806.08730, 2018.
  43. 43.K. Nguyen, H. Daumé III, and J. Boyd-Graber. Reinforcement learning for bandit neural machine translation with simulated human feedback. arXiv preprint arXiv:1707.07402, 2017.
  44. 44.T. Niu and M. Bansal. Polite dialogue generation without parallel data. Transactions of the Association for Computational Linguistics, 6:373–389, 2018.
  45. 45.R. Paulus, C. Xiong, and R. Socher. A deep reinforced model for abstractive summarization. arXiv preprint arXiv:1705.04304, 2017.
  46. 46.E. Perez, S. Karamcheti, R. Fergus, J. Weston, D. Kiela, and K. Cho. Finding generalizable evidence by learning to convince q&a models. arXiv preprint arXiv:1909.05863, 2019.
  47. 47.A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever. Improving language understanding by generative pre-training. URL https://s3-us-west-2. amazonaws. com/openai-assets/researchcovers/languageunsupervised/language understanding paper. pdf, 2018.
  48. 48.A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever. Language models are unsupervised multitask learners. OpenAI Blog, 1(8):9, 2019.
  49. 49.C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv preprint arXiv:1910.10683, 2019.
  50. 50.M. Ranzato, S. Chopra, M. Auli, and W. Zaremba. Sequence level training with recurrent neural networks. arXiv preprint arXiv:1511.06732, 2015.
  51. 51.D. R. Reddy et al. Speech understanding systems: A summary of results of the five-year research effort. department of computer science, 1977.
  52. 52.S. Ross, G. Gordon, and D. Bagnell. A reduction of imitation learning and structured prediction to no-regret online learning. In Proceedings of the fourteenth international conference on artificial intelligence and statistics, pages 627–635, 2011.
  53. 53.S. Rothe, S. Narayan, and A. Severyn. Leveraging pre-trained checkpoints for sequence generation tasks. Transactions of the Association for Computational Linguistics, 2020.
  54. 54.A. M. Rush, S. Chopra, and J. Weston. A neural attention model for abstractive sentence summarization. arXiv preprint arXiv:1509.00685, 2015.
  55. 55.N. Schluter. The limits of automatic summarisation according to rouge. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers, pages 41–45, 2017.
  56. 56.F. Schmidt. Generalization in generation: A closer look at exposure bias. arXiv preprint arXiv:1910.00292, 2019.
  57. 57.J. Schulman, P. Moritz, S. Levine, M. Jordan, and P. Abbeel. High-dimensional continuous control using generalized advantage estimation. In Proceedings of the International Conference on Learning Representations (ICLR), 2016.
  58. 58.J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017.
  59. 59.A. See, P. J. Liu, and C. D. Manning. Get to the point: Summarization with pointer-generator networks. arXiv preprint arXiv:1704.04368, 2017.
  60. 60.K. Song, X. Tan, T. Qin, J. Lu, and T.-Y. Liu. Mass: Masked sequence to sequence pre-training for language generation. arXiv preprint arXiv:1905.02450, 2019.
  61. 61.P. Tambwekar, M. Dhuliawala, A. Mehta, L. J. Martin, B. Harrison, and M. O. Riedl. Controllable neural story generation via reinforcement learning. arXiv preprint arXiv:1809.10736, 2018.
  62. 62.A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin. Attention is all you need. In Advances in neural information processing systems, pages 5998–6008, 2017.
  63. 63.M. Völske, M. Potthast, S. Syed, and B. Stein. Tl; dr: Mining reddit to learn automatic summarization. In Proceedings of the Workshop on New Frontiers in Summarization, pages 59–63, 2017.
  64. 64.S. Welleck, I. Kulikov, S. Roller, E. Dinan, K. Cho, and J. Weston. Neural text generation with unlikelihood training. arXiv preprint arXiv:1908.04319, 2019.
  65. 65.Y. Wu and B. Hu. Learning to extract coherent summary via deep reinforcement learning. In Thirty-Second AAAI Conference on Artificial Intelligence, 2018.
  66. 66.Y. Wu, M. Schuster, Z. Chen, Q. V. Le, M. Norouzi, W. Macherey, M. Krikun, Y. Cao, Q. Gao, K. Macherey, et al. Google’s neural machine translation system: Bridging the gap between human and machine translation. arXiv preprint arXiv:1609.08144, 2016.
  67. 67.Y. Yan, W. Qi, Y. Gong, D. Liu, N. Duan, J. Chen, R. Zhang, and M. Zhou. Prophetnet: Predicting future n-gram for sequence-to-sequence pre-training. arXiv preprint arXiv:2001.04063, 2020.
  68. 68.S. Yi, R. Goel, C. Khatri, A. Cervone, T. Chung, B. Hedayatnia, A. Venkatesh, R. Gabriel, and D. Hakkani-Tur. Towards coherent and engaging spoken dialog response generation using automatic conversation evaluators. arXiv preprint arXiv:1904.13015, 2019.
  69. 69.H. Zhang, D. Duckworth, D. Ippolito, and A. Neelakantan. Trading off diversity and quality in natural language generation. arXiv preprint arXiv:2004.10450, 2020.
  70. 70.J. Zhang, Y. Zhao, M. Saleh, and P. J. Liu. Pegasus: Pre-training with extracted gap-sentences for abstractive summarization. arXiv preprint arXiv:1912.08777, 2019.
  71. 71.Y. Zhang, D. Li, Y. Wang, Y. Fang, and W. Xiao. Abstract text summarization with a convolutional seq2seq model. Applied Sciences, 9(8):1665, 2019.
  72. 72.W. Zhou and K. Xu. Learning to compare for better training and evaluation of open domain natural language generation models. arXiv preprint arXiv:2002.05058, 2020.
  73. 73.D. M. Ziegler, N. Stiennon, J. Wu, T. B. Brown, A. Radford, D. Amodei, P. Christiano, and G. Irving. Fine-tuning language models from human preferences. arXiv preprint arXiv:1909.08593, 2019.

Citation

MLA
Stiennon, N., et al. “Learning to Summarize from Human Feedback”. arXiv, 2020, http://arxiv.org/abs/2009.01325v3.
APA
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D. M., Lowe, R., Voss, C., Radford, A., Amodei, D., & Christiano, P. (2020). Learning to summarize from human feedback. arXiv. http://arxiv.org/abs/2009.01325v3
Chicago
Stiennon, N., L. Ouyang, J. Wu, et al. 2020. “Learning to Summarize from Human Feedback”. arXiv. http://arxiv.org/abs/2009.01325v3.
Harvard
Stiennon, N. et al. (2020) “Learning to summarize from human feedback”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2009.01325v3.
Vancouver
1. Stiennon N, Ouyang L, Wu J, Ziegler DM, Lowe R, Voss C, Radford A, Amodei D, Christiano P (2020) Learning to summarize from human feedback. arXiv

BibTeX

@article{stiennon2020learning,
  title = {Learning to summarize from human feedback},
  author = {Stiennon, Nisan and Ouyang, Long and Wu, Jeff and Ziegler, Daniel M. and Lowe, Ryan and Voss, Chelsea and Radford, Alec and Amodei, Dario and Christiano, Paul},
  year = {2020},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2009.01325v3},
  eprint = {2009.01325}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: Published with permission