Pretraining Language Models with Human Preferences

Tomasz KorbakKejian ShiAngelica ChenRasika Vinayak BhaleraoChristopher L. BuckleyJason PhangSamuel R. BowmanEthan Perez

article2023ICML308 citations

Demonstrates that integrating human preferences directly into language model pretraining via conditional training reduces undesirable outputs by up to an order of magnitude while preserving downstream capabilities, outperforming the standard pipeline of pretraining followed by post-hoc alignment fine-tuning.

Listen

Modern language models are typically pretrained by imitating vast amounts of uncurated internet text. Consequently, they often internalize and reproduce harmful behaviors, such as generating offensive language, leaking personally identifiable information, and producing flawed code. The standard industry approach attempts to fix these issues after pretraining through safety filters or post-hoc finetuning techniques like reinforcement learning from human feedback. However, because large models strongly resist unlearning their initial training data, these downstream adjustments often fail or require costly interventions. Preemptively filtering datasets also introduces severe trade-offs, such as bottlenecking data scale, reducing model capabilities, and amplifying biases.

The article evaluates whether incorporating human preferences directly into the pretraining phase—a framework termed Pretraining with Human Feedback—can steer models away from generating undesirable text while preserving their overall performance and knowledge. Across a compute-optimal scale of 3.32 billion tokens and 124-million-parameter models, the authors benchmark standard imitation learning against five preference-guided pretraining objectives: conditional training, dataset filtering, unlikelihood loss, reward-weighted regression, and advantage-weighted regression. These objectives were systematically evaluated across three distinct tasks: mitigating toxic speech, preventing personal data leakage, and adhering to the PEP8 Python coding standard.

The investigation produced four central findings. First, conditional training—a technique that prepends segments with control tokens indicating preference scores—emerged as the most effective and Pareto-optimal approach across all tasks. Second, conditional training reduced the generation of undesirable content by up to an order of magnitude (for example, reducing toxicity scores from 0.0141 down to 0.0011) and demonstrated continuous improvement as training progressed without plateauing. Third, this method preserved general capabilities, matching standard imitation pretraining on zero-shot passage understanding and downstream language classification benchmarks. Finally, pretraining with human feedback substantially outperformed the conventional pipeline of standard pretraining followed by post-hoc finetuning, proving two to three times more effective at suppressing undesirable outputs and maintaining superior robustness against automated adversarial red-teaming.

These results demonstrate that language models should be aligned from the beginning of training rather than taught bad behaviors only to attempt unlearning them later. Conditional training enables models to retain broad knowledge from lower-quality or toxic text without imitating it during generation. In practice, this shifts the cost and risk structure of model development: the additional computational overhead of running segment-level reward scoring during data ingestion is minimal compared to the compounding costs and security risks of deploying poorly aligned base models.

Organizations developing or deploying foundation models should consider moving beyond pure imitation pretraining by adopting segment-level conditional training during initial pretraining runs. When planning training budgets, teams should incorporate reward scoring pipelines into early data preparation workflows rather than relying solely on post-hoc alignment stages. However, decision-makers must note that while pretraining with feedback dramatically improves safety and adversarial robustness, it does not guarantee complete immunity against determined adversarial prompts. Further validation at larger model parameter scales (such as tens or hundreds of billions of parameters) is recommended to confirm that these scaling trends fully generalize to enterprise-scale deployments.

arXiv: 2302.08582
Cover for Pretraining Language Models with Human Preferences

Abstract

Language models (LMs) are pretrained to imitate internet text, including content that would violate human preferences if generated by an LM: falsehoods, offensive comments, personally identifiable information, low-quality or buggy code, and more. Here, we explore alternative objectives for pretraining LMs in a way that also guides them to generate text aligned with human preferences. We benchmark five objectives for pretraining with human feedback across three tasks and study how they affect the trade-off between alignment and capabilities of pretrained LMs. We find a Pareto-optimal and simple approach among those we explored: conditional training, or learning distribution over tokens conditional on their human preference scores given by a reward model. Conditional training reduces the rate of undesirable content by up to an order of magnitude, both when generating without a prompt and with an adversarially-chosen prompt. Moreover, conditional training maintains the downstream task performance of standard LM pretraining, both before and after task-specific finetuning. Pretraining with human feedback results in much better preference satisfaction than standard LM pretraining followed by finetuning with feedback, i.e., learning and then unlearning undesirable behavior. Our results suggest that we should move beyond imitation learning when pretraining LMs and incorporate human preferences from the start of training.

Table of Contents

  • 1. Introduction
  • 2. Methods
  • 3. Experimental Setup
  • 3.1. Tasks
  • 3.2. Model Architecture and Hyperparameters
  • 3.3. Training Data
  • 4. Pretraining Experiments
  • 4.1. Capabilities-Alignment Trade-offs
  • 4.2. Robustness to Red-Teaming
  • 4.3. Downstream Benchmarks
  • 4.4. Diversity
  • 5. Finetuning with Human Feedback
  • 6. Related Work
  • 7. Conclusion
  • References
  • A. Acknowledgments
  • B. Hyperparameters and Implementation Details
  • C. Details on the red-teaming procedure
  • D. Details on GLUE evaluation
  • E. Additional results on scores of LM samples
  • F. Additional results for diversity evaluation
  • G. Additional results for finetuning experiments

Knowls

  1. Knowl 1 — Pretraining with human-feedback rewards

    definition

    Pretraining with human feedback (PHF) augments ordinary language-model pretraining with a scalar reward function that estimates how preferable each training segment is. A document is x=(x1,…,x∣x∣)x=(x^1,\ldots,x^{|x|}), where segment xi=(x1i,…,xNii)x^i=(x^i_1,\ldots,x^i_{N_i}) consists of tokens from a fixed vocabulary VV; πθ\pi_\theta is an autoregressive language model with parameters θ\theta; and R(xi)∈RR(x^i)\in\mathbb{R} is a segment-level preference score. Training maximizes an objective L(x)\mathcal{L}(x) over documents, θ∗=arg⁡max⁡θ∑x∈DL(x)\theta^*=\arg\max_\theta\sum_{x\in D}\mathcal{L}(x), while retaining undesirable examples in DD rather than deleting them.

    The standard maximum-likelihood objective is LMLE(x)=log⁡πθ(x)=∑i∑j=1Nilog⁡πθ(xji∣x<j≤i)\mathcal{L}_{\mathrm{MLE}}(x)=\log\pi_\theta(x)=\sum_i\sum_{j=1}^{N_i}\log\pi_\theta(x^i_j\mid x^{\leq i}_{<j}), where x<j≤ix^{\leq i}_{<j} denotes all tokens preceding token xjix^i_j in the document. The paper evaluates rewards for three preference targets: negative Detoxify toxicity probability for toxicity avoidance, negative Scrubadub-detected PII instances per character for privacy, and negative pycodestyle violations per character for PEP8-compliant Python. Text is scored at sentence level for toxicity and PII and at line level for PEP8.

  2. Knowl 2 — Conditional training as the principal PHF method

    model/method

    Conditional training represents the preference score of every document segment with a control token and trains the language model to predict the resulting augmented sequence. For a threshold tt, define ci=⟨ ⁣∣good ⁣∣⟩c_i=\langle\!|\mathrm{good}\!|\rangle when R(xi)≥tR(x^i)\geq t and ci=⟨ ⁣∣bad ⁣∣⟩c_i=\langle\!|\mathrm{bad}\!|\rangle otherwise. The objective is

    LCond(x)=log⁡πθ(c1,x1,…,c∣x∣,x∣x∣).\mathcal{L}_{\mathrm{Cond}}(x)=\log\pi_\theta(c_1,x^1,\ldots,c_{|x|},x^{|x|}).

    Unlike document-level filtering, the control token is attached separately to each sentence or code line. At generation time, the model is prompted with ⟨ ⁣∣good ⁣∣⟩\langle\!|\mathrm{good}\!|\rangle and samples from the conditional distribution intended to represent preferred text. During the experiments, the model was also exposed to 1% of sentences without a control token; this slightly improved capability as measured by KL divergence from GPT-3 while causing only a negligible alignment penalty. For toxicity and PII generation, both control tokens were blocked after the initial good prefix; for PEP8, bad tokens were blocked and generated good tokens were removed in post-processing.

  3. Knowl 3 — Alternative PHF objectives

    model/method

    The paper compares conditional training with four other reward-aware objectives. Let Rˉ(x)=∣x∣−1∑iR(xi)\bar R(x)=|x|^{-1}\sum_iR(x^i) be the document-average reward, let ℓi=log⁡πθ(xi∣x<i)\ell_i=\log\pi_\theta(x^i\mid x^{<i}) be the log probability of segment ii, and let pij=πθ(xji∣x<j≤i)p_{ij}=\pi_\theta(x^i_j\mid x^{\leq i}_{<j}). The alternatives are:

    • Filtering: LFilt(x)=ℓ1+⋯+ℓ∣x∣\mathcal{L}_{\mathrm{Filt}}(x)=\ell_1+\cdots+\ell_{|x|} if Rˉ(x)>t\bar R(x)>t, and 00 otherwise. In practice, low-reward documents are discarded and the remaining data are repeated to preserve the training-token budget.
    • Unlikelihood: high-reward segments use likelihood training, while low-reward segments use token-level unlikelihood:
    LUL(x)=∑i:R(xi)>tℓi+α∑i:R(xi)≤t∑j=1Nilog⁡(1−pij),\mathcal{L}_{\mathrm{UL}}(x)=\sum_{i:R(x^i)>t}\ell_i+\alpha\sum_{i:R(x^i)\leq t}\sum_{j=1}^{N_i}\log(1-p_{ij}),

    where α\alpha controls the strength of penalizing the observed tokens in low-reward segments.

    • Reward-weighted regression:
    LRWR(x)=∑iℓiexp⁡ ⁣(R(xi)β),\mathcal{L}_{\mathrm{RWR}}(x)=\sum_i\ell_i\exp\!\left(\frac{R(x^i)}{\beta}\right),

    where β>0\beta>0 controls reward reweighting.

    • Advantage-weighted regression: a value head sharing all but the final head with the language model estimates Vθ(xji)V_\theta(x^i_j), and the advantage is A(xji)=R(xi)−Vθ(xji)A(x^i_j)=R(x^i)-V_\theta(x^i_j). The joint policy-and-value objective is
    LAWR(x)=α∑i∑j=1Nilog⁡pijexp⁡ ⁣(A(xji)β)−(1−α)∑i∑j=1Ni(Vθ(xji)−R(xi))2,\mathcal{L}_{\mathrm{AWR}}(x)=\alpha\sum_i\sum_{j=1}^{N_i}\log p_{ij}\exp\!\left(\frac{A(x^i_j)}{\beta}\right)-(1-\alpha)\sum_i\sum_{j=1}^{N_i}\bigl(V_\theta(x^i_j)-R(x^i)\bigr)^2,

    with α\alpha trading off policy and value losses and β\beta controlling advantage reweighting.

  4. Knowl 4 — Evaluation design for alignment and capability

    experimental setup

    All models use the 124-million-parameter GPT-2-small architecture and are trained for 3.32 billion tokens. Toxicity and PII models use 1.95 million documents subsampled from The Pile; code models use 1.5 million Python files from a cleaned GitHub corpus. Learning rate and batch size are tuned separately for each task and objective.

    Alignment is measured by the misalignment score, defined as the negative task reward. For each model, 4,096 unconditional samples are generated with temperature 0.70.7, nucleus probability p=0.9p=0.9, and lengths between 10 and 128 tokens; the reported score is the average over these samples. Capability is additionally approximated by the estimated divergence from a reference model, D^KL(p ∥ πθ)=N−1∑n=1Nlog⁡[p(xn)/πθ(xn)]\widehat D_{\mathrm{KL}}(p\,\|\,\pi_\theta)=N^{-1}\sum_{n=1}^{N}\log[p(x_n)/\pi_\theta(x_n)], where pp is GPT-3 for toxicity and PII or a 12-billion-parameter Codex model for PEP8, N=4096N=4096, and xnx_n are reference-model samples of at most 64 tokens. Lower KL and lower misalignment are preferred.

    Downstream capability tests include zero-shot LAMBADA for toxicity and PII models, HumanEval pass@10 and pass@100 for code models, and GLUE after task-specific fine-tuning for the natural-language models. Sample diversity is assessed with unigram and bigram entropy, distinct-token ratios, Self-BLEU-5, and a degeneration ratio.

  5. Knowl 5 — Conditional training gives the best alignment-capability trade-off

    empirical result

    Across toxicity, PII, and PEP8, every PHF objective substantially reduces undesirable content relative to MLE, but conditional training is the only method consistently on the Pareto frontier of alignment versus KL divergence from the reference model. On toxicity, the average misalignment score falls from 0.01410.0141 for MLE to 0.00110.0011 for conditional training, an approximately order-of-magnitude reduction. Conditional training is strictly Pareto-optimal for toxicity and lies on the frontier for PII and PEP8. Its misalignment score continues to decrease throughout the 3.32-billion-token run without a clear plateau.

    Filtering is the strongest competing alignment baseline and obtains the lowest PEP8 misalignment score, but it incurs a large capability penalty on PII and PEP8. Reward-weighted and advantage-weighted regression improve alignment only slightly while substantially increasing KL divergence. Unlikelihood is highly task-dependent: it performs especially well for toxicity but yields only small alignment gains for PII and PEP8. Thus, conditioning on segment-level preference scores preserves more of the original language-model distribution than the other reward-aware objectives while still suppressing undesirable generations.

  6. Knowl 6 — Adversarial red-teaming exposes improved but incomplete robustness

    empirical result

    The paper evaluates robustness with an iterative black-box red-teaming procedure. An InstructGPT red model proposes adversarial prompts, while a target language model generates 512 responses per prompt using temperature 0.70.7, top-p=0.9p=0.9, and response lengths of 10--64 tokens. For a prompt aa, its utility is the average target-response misalignment, u(a)=N−1∑j[−R(xj)]u(a)=N^{-1}\sum_j[-R(x_j)]. At each of ten rounds, four prompts are sampled from the current pool with probability proportional to exp⁡(u(a)/β)\exp(u(a)/\beta), and InstructGPT generates 20 new prompts from those examples. Results are averaged over ten independent trials.

    Conditional training and filtering are the most robust methods overall. After ten red-teaming rounds, conditional training reduces toxicity and PII misalignment by up to an order of magnitude relative to MLE. Unlikelihood is the most robust method for toxicity but the least robust for PII; all other PHF methods are generally more robust than MLE. Nevertheless, every PHF model remains exploitable: successive rounds continue to increase the misalignment of elicited responses, with no clear plateau after ten rounds. PHF therefore improves adversarial robustness without guaranteeing safe behavior under arbitrary prompts.

  7. Knowl 7 — Conditional training largely preserves downstream capabilities

    empirical result

    Conditional training generally retains MLE-level representation quality and zero-shot performance, unlike most alternative PHF objectives. On LAMBADA, conditional training slightly exceeds MLE accuracy for both toxicity and PII models, whereas filtering, reward-weighted regression, and advantage-weighted regression generally reduce accuracy; unlikelihood matches MLE for PII but performs poorly for toxicity.

    After fine-tuning on eight GLUE tasks, the average test scores for toxicity-pretrained models are 72.1±0.7472.1\pm0.74 for MLE, 71.4±0.6071.4\pm0.60 for conditional training, 71.0±0.4771.0\pm0.47 for filtering, 66.0±0.8366.0\pm0.83 for AWR, 58.2±1.5758.2\pm1.57 for RWR, and 68.8±0.3968.8\pm0.39 for unlikelihood. For PII-pretrained models, the corresponding scores are 72.1±0.6672.1\pm0.66, 72.5±0.9172.5\pm0.91, 71.4±0.5571.4\pm0.55, 72.7±0.4172.7\pm0.41, 70.1±2.2970.1\pm2.29, and 72.9±0.6172.9\pm0.61. Code models show a larger capability gap on HumanEval: filtering closes the gap in pass@100, while conditional training is no longer the best PHF method and unlikelihood obtains the lowest scores. These results indicate that preference-aware pretraining can preserve useful representations even though it changes the generation distribution.

  8. Knowl 8 — Pretraining with feedback beats feedback-only fine-tuning

    empirical result

    The paper compares PHF from random initialization with the standard pipeline of MLE pretraining followed by feedback-based fine-tuning. MLE checkpoints trained on either 1.66 billion tokens or 2.97 billion tokens are fine-tuned for another 1.66 billion or 0.30 billion tokens, respectively, using the same five PHF objectives and the same task data.

    Conditional PHF pretraining achieves lower misalignment than every feedback-only fine-tuning run on all three tasks, often by a large margin. For PII, conditional pretraining reaches a misalignment score of 0.00130.0013, compared with 0.00180.0018 after fine-tuning for 1.66 billion additional tokens and 0.00230.0023 after the shorter run whose total training exposure is about 3.3 billion tokens. The advantage grows as less feedback is available for fine-tuning. Conditional-pretrained models are also substantially more robust to red-teaming: on PII, ten red-teaming rounds are needed to reach the misalignment level that a model fine-tuned with conditional training reaches after only one round. The results support using preference information throughout pretraining rather than first learning undesirable behavior and attempting to remove it later.

  9. Knowl 9 — Preference conditioning changes diversity without severe degeneration

    empirical result

    The PHF objectives alter generation diversity in different ways. Conditional training, and to a lesser extent filtering, reduce sample entropy relative to MLE but retain a fraction of distinct unigrams closer to MLE. Unlikelihood, RWR, and AWR preserve diversity metrics closer to MLE but show slightly more degeneration, measured by repeated-token behavior. Across toxicity and PII experiments, none of the PHF objectives causes substantial entropy collapse or severe degeneration in absolute terms.

    This analysis qualifies the alignment-capability results: conditional training's lower KL divergence and strong alignment do not imply identical output diversity, but its diversity reduction remains moderate rather than catastrophic.

  10. Knowl 10 — Known limitations of PHF alignment

    limitation

    PHF depends on the reward model or detector accurately representing the desired human preference. The study uses narrow proxies for toxicity, PII, and PEP8, so its results do not establish alignment with general human values or safety in deployment. All PHF objectives remain vulnerable to black-box adversarial prompting, and some objectives have strong task-specific failures: unlikelihood is unreliable outside toxicity, while filtering can damage capabilities and conditional training can reduce diversity.

    PHF also requires running a reward model over the pretraining corpus and adding preference-related data processing, although the paper argues that this inference cost can be kept small with a compact, distilled, or low-precision reward model. The experiments use one 124-million-parameter architecture, three automatically scored tasks, and offline segment-level rewards; broader model scales, preference distributions, and deployment settings are not established by these experiments.

Coverage note — Detailed per-task GLUE scores, threshold-ablation plots, auxiliary tail-risk plots, individual adversarial prompt examples, and full hyperparameter tables were omitted because the knowls retain their main conclusions without reproducing every appendix measurement.

References

  1. 1.Abid, A., Farooqi, M., and Zou, J. Persistent anti-muslim bias in large language models. In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society, AIES ’21, pp. 298–306, New York, NY, USA, 2021. Association for Computing Machinery. ISBN 9781450384735. doi: 10.1145/3461702.3462624. URL https://doi.org/10.1145/3461702.3462624.
  2. 2.Askell, A., Bai, Y., Chen, A., Drain, D., Ganguli, D., Henighan, T., Jones, A., Joseph, N., Mann, B., Das-Sarma, N., Elhage, N., Hatfield-Dodds, Z., Hernandez, D., Kernion, J., Ndousse, K., Olsson, C., Amodei, D., Brown, T., Clark, J., McCandlish, S., Olah, C., and Kaplan, J. A general language assistant as a laboratory for alignment, 2021. URL https://arxiv.org/abs/2112.00861.
  3. 3.Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., Das-Sarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., Joseph, N., Kadavath, S., Kernion, J., Conerly, T., El-Showk, S., Elhage, N., Hatfield-Dodds, Z., Hernandez, D., Hume, T., Johnston, S., Kravec, S., Lovitt, L., Nanda, N., Olsson, C., Amodei, D., Brown, T., Clark, J., McCandlish, S., Olah, C., Mann, B., and Kaplan, J. Training a helpful and harmless assistant with reinforcement learning from human feedback, 2022.
  4. 4.Bar-Haim, R., Dagan, I., Dolan, B., Ferro, L., and Giampiccolo, D. The second pascal recognising textual entailment challenge. Proceedings of the Second PASCAL Challenges Workshop on Recognising Textual Entailment, 01 2006.
  5. 5.Bengio, Y., Ducharme, R., Vincent, P., and Janvin, C. A neural probabilistic language model. J. Mach. Learn. Res., 3(null):1137–1155, mar 2003. ISSN 1532-4435.
  6. 6.Bentivogli, L., Magnini, B., Dagan, I., Dang, H. T., and Giampiccolo, D. The fifth PASCAL recognizing textual entailment challenge. In Proceedings of the Second Text Analysis Conference, TAC 2009, Gaithersburg, Maryland, USA, November 16-17, 2009. NIST, 2009. URL https://tac.nist.gov/publications/2009/additional.papers/RTE5_overview.proceedings.pdf.
  7. 7.Borkan, D., Dixon, L., Sorensen, J., Thain, N., and Vasserman, L. Nuanced metrics for measuring unintended bias with real data for text classification. In Companion Proceedings of The 2019 World Wide Web Conference, WWW ’19, pp. 491–500, New York, NY, USA, 2019. Association for Computing Machinery. ISBN 9781450366755. doi: 10.1145/3308560.3317593. URL https://doi.org/10.1145/3308560.3317593.
  8. 8.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. Language models are few-shot learners. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., and Lin, H. (eds.), Advances in Neural Information Processing Systems, volume 33, pp. 1877–1901. Curran Associates, Inc., 2020. URL https://proceedings.neurips.cc/paper/2020/file/1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf.
  9. 9.Carlini, N., Liu, C., Erlingsson, U., Kos, J., and Song, D. The secret sharer: Evaluating and testing unintended memorization in neural networks. In Proceedings of the 28th USENIX Conference on Security Symposium, SEC’19, pp. 267–284, USA, 2019. USENIX Association. ISBN 9781939133069.
  10. 10.Carlini, N., Tramer, F., Wallace, E., Jagielski, M., Herbert-Voss, A., Lee, K., Roberts, A., Brown, T., Song, D., Erlingsson, U., Oprea, A., and Raffel, C. Extracting training data from large language models, 2020. URL https://arxiv.org/abs/2012.07805.
  11. 11.Carlini, N., Ippolito, D., Jagielski, M., Lee, K., Tramer, F., and Zhang, C. Quantifying memorization across neural language models, 2022. URL https://arxiv.org/abs/2202.07646.
  12. 12.Cer, D., Diab, M., Agirre, E., Lopez-Gazpio, I., and Specia, L. SemEval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation. In Proceedings of the 11th International Workshop on Semantic Evaluation (SemEval-2017), pp. 1–14, Vancouver, Canada, August 2017. Association for Computational Linguistics. doi: 10.18653/v1/S17-2001. URL https://aclanthology.org/S17-2001.
  13. 13.Chen, A., Scheurer, J., Korbak, T., Campos, J. A., Chan, J. S., Bowman, S. R., Cho, K., and Perez, E. Improving code generation by training with natural language feedback, 2023.
  14. 14.Chen, L., Lu, K., Rajeswaran, A., Lee, K., Grover, A., Laskin, M., Abbeel, P., Srinivas, A., and Mordatch, I. Decision transformer: Reinforcement learning via sequence modeling. In Ranzato, M., Beygelzimer, A., Dauphin, Y., Liang, P., and Vaughan, J. W. (eds.), Advances in Neural Information Processing Systems, volume 34, pp. 15084–15097. Curran Associates, Inc., 2021a. URL https://proceedings.neurips.cc/paper/2021/file/7f489f642a0ddb10272b5c31057f0663-Paper.pdf.
  15. 15.Chen, M., Tworek, J., Jun, H., Yuan, Q., de Oliveira Pinto, H. P., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., Ray, A., Puri, R., Krueger, G., Petrov, M., Khlaaf, H., Sastry, G., Mishkin, P., Chan, B., Gray, S., Ryder, N., Pavlov, M., Power, A., Kaiser, L., Bavarian, M., Winter, C., Tillet, P., Such, F. P., Cummings, D., Plappert, M., Chantzis, F., Barnes, E., Herbert-Voss, A., Guss, W. H., Nichol, A., Paino, A., Tezak, N., Tang, J., Babuschkin, I., Balaji, S., Jain, S., Saunders, W., Hesse, C., Carr, A. N., Leike, J., Achiam, J., Misra, V., Morikawa, E., Radford, A., Knight, M., Brundage, M., Murati, M., Mayer, K., Welinder, P., McGrew, B., Amodei, D., McCandlish, S., Sutskever, I., and Zaremba, W. Evaluating large language models trained on code. 2021b.
  16. 16.Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, Y., Wang, X., Dehghani, M., Brahma, S., Webson, A., Gu, S. S., Dai, Z., Suzgun, M., Chen, X., Chowdhery, A., Castro-Ros, A., Pellat, M., Robinson, K., Valter, D., Narang, S., Mishra, G., Yu, A., Zhao, V., Huang, Y., Dai, A., Yu, H., Petrov, S., Chi, E. H., Dean, J., Devlin, J., Roberts, A., Zhou, D., Le, Q. V., and Wei, J. Scaling instruction-finetuned language models, 2022. URL https://arxiv.org/abs/2210.11416.
  17. 17.Dagan, I., Glickman, O., and Magnini, B. The pascal recognising textual entailment challenge. In Proceedings of the First International Conference on Machine Learning Challenges: Evaluating Predictive Uncertainty Visual Object Classification, and Recognizing Textual Entailment, MLCW’05, pp. 177–190, Berlin, Heidelberg, 2005. Springer-Verlag. ISBN 3540334270. doi: 10.1007/11736790 9. URL https://doi.org/10.1007/11736790_9.
  18. 18.Dai, N., Liang, J., Qiu, X., and Huang, X. Style transformer: Unpaired text style transfer without disentangled latent representation. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pp. 5997–6007, Florence, Italy, July 2019. Association for Computational Linguistics. doi: 10.18653/v1/P19-1601. URL https://aclanthology.org/P19-1601.
  19. 19.Dettmers, T. and Zettlemoyer, L. The case for 4-bit precision: k-bit inference scaling laws, 2023.
  20. 20.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 4171–4186, Minneapolis, Minnesota, June 2019. Association for Computational Linguistics. doi: 10.18653/v1/N19-1423. URL https://aclanthology.org/N19-1423.
  21. 21.Dolan, W. B. and Brockett, C. Automatically constructing a corpus of sentential paraphrases. In Proceedings of the Third International Workshop on Paraphrasing (IWP2005), 2005. URL https://aclanthology.org/I05-5002.
  22. 22.Emmons, S., Eysenbach, B., Kostrikov, I., and Levine, S. Rvs: What is essential for offline RL via supervised learning? In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=S874XAIpkR-.
  23. 23.Fan, A., Lewis, M., and Dauphin, Y. Hierarchical neural story generation. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 889–898, Melbourne, Australia, July 2018. Association for Computational Linguistics. doi: 10.18653/v1/P18-1082. URL https://aclanthology.org/P18-1082.
  24. 24.Ficler, J. and Goldberg, Y. Controlling linguistic style aspects in neural language generation. In Proceedings of the Workshop on Stylistic Variation, pp. 94–104, Copenhagen, Denmark, September 2017. Association for Computational Linguistics. doi: 10.18653/v1/W17-4912. URL https://aclanthology.org/W17-4912.
  25. 25.Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., Presser, S., and Leahy, C. The Pile: An 800gb dataset of diverse text for language modeling. arXiv preprint arXiv:2101.00027, 2020.
  26. 26.Gehman, S., Gururangan, S., Sap, M., Choi, Y., and Smith, N. A. RealToxicityPrompts: Evaluating neural toxic degeneration in language models. In Findings of the Association for Computational Linguistics: EMNLP 2020, pp. 3356–3369, Online, November 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.findings-emnlp.301. URL https://aclanthology.org/2020.findings-emnlp.301.
  27. 27.Giampiccolo, D., Magnini, B., Dagan, I., and Dolan, B. The third PASCAL recognizing textual entailment challenge. In Proceedings of the ACL-PASCAL Workshop on Textual Entailment and Paraphrasing, pp. 1–9, Prague, June 2007. Association for Computational Linguistics. URL https://aclanthology.org/W07-1401.
  28. 28.Go, D., Korbak, T., Kruszewski, G., Rozen, J., Ryu, N., and Dymetman, M. Aligning language models with preferences through f-divergence minimization, 2023.
  29. 29.Hanu, L. and Unitary team. Detoxify. Github. https://github.com/unitaryai/detoxify, 2020.
  30. 30.Henderson, P., Sinha, K., Angelard-Gontier, N., Ke, N. R., Fried, G., Lowe, R., and Pineau, J. Ethical challenges in data-driven dialogue systems. In Proceedings of the 2018 AAAI/ACM Conference on AI, Ethics, and Society, AIES ’18, pp. 123–129, New York, NY, USA, 2018. Association for Computing Machinery. ISBN 9781450360128. doi: 10.1145/3278721.3278777. URL https://doi.org/10.1145/3278721.3278777.
  31. 31.Hendrycks, D., Mazeika, M., Kadavath, S., and Song, D. Using self-supervised learning can improve model robustness and uncertainty. Advances in Neural Information Processing Systems (NeurIPS), 2019.
  32. 32.Hendrycks, D., Liu, X., Wallace, E., Dziedzic, A., Krishnan, R., and Song, D. Pretrained transformers improve out-of-distribution robustness. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 2744–2751, Online, July 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.acl-main.244. URL https://aclanthology.org/2020.acl-main.244.
  33. 33.Hewitt, J. Initializing new word embeddings for pretrained language models, 2021.
  34. 34.Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., de las Casas, D., Hendricks, L. A., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, K., van den Driessche, G., Damoc, B., Guy, A., Osindero, S., Simonyan, K., Elsen, E., Vinyals, O., Rae, J. W., and Sifre, L. An empirical analysis of compute-optimal large language model training. In Oh, A. H., Agarwal, A., Belgrave, D., and Cho, K. (eds.), Advances in Neural Information Processing Systems, 2022. URL https://openreview.net/forum?id=iBBcRUlOAPR.
  35. 35.Holtzman, A., Buys, J., Du, L., Forbes, M., and Choi, Y. The curious case of neural text degeneration. In International Conference on Learning Representations, 2020. URL https://openreview.net/forum?id=rygGQyrFvH.
  36. 36.Honnibal, M., Montani, I., Van Landeghem, S., and Boyd, A. spaCy: Industrial-strength Natural Language Processing in Python. 2020. doi: 10.5281/zenodo.1212303.
  37. 37.Jang, Y., Lee, J., and Kim, K.-E. GPT-critic: Offline reinforcement learning for end-to-end task-oriented dialogue systems. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=qaxhBG1UUaS.
  38. 38.Janner, M., Li, Q., and Levine, S. Offline reinforcement learning as one big sequence modeling problem. In Advances in Neural Information Processing Systems, 2021.
  39. 39.Jaques, N., Shen, J. H., Ghandeharioun, A., Ferguson, C., Lapedriza, A., Jones, N., Gu, S., and Picard, R. Human-centric dialog training via offline reinforcement learning. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 3985–4003, Online, November 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.emnlp-main.327. URL https://aclanthology.org/2020.emnlp-main.327.
  40. 40.Keskar, N. S., McCann, B., Varshney, L. R., Xiong, C., and Socher, R. Ctrl: A conditional transformer language model for controllable generation, 2019. URL https://arxiv.org/abs/1909.05858.
  41. 41.Khalifa, M., Elsahar, H., and Dymetman, M. A distributional approach to controlled text generation. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net, 2021. URL https://openreview.net/forum?id=jWkw45-9AbL.
  42. 42.Korbak, T., Elsahar, H., Kruszewski, G., and Dymetman, M. On reinforcement learning and distribution matching for fine-tuning language models with no catastrophic forgetting. In Oh, A. H., Agarwal, A., Belgrave, D., and Cho, K. (eds.), Advances in Neural Information Processing Systems, 2022a. URL https://openreview.net/forum?id=XvI6h-s4un.
  43. 43.Korbak, T., Perez, E., and Buckley, C. RL with KL penalties is better viewed as Bayesian inference. In Findings of the Association for Computational Linguistics: EMNLP 2022, pp. 1083–1091, Abu Dhabi, United Arab Emirates, December 2022b. Association for Computational Linguistics. URL https://aclanthology.org/2022.findings-emnlp.77.
  44. 44.Kumar, A., Peng, X. B., and Levine, S. Reward-conditioned policies, 2019. URL https://arxiv.org/abs/1912.13465.
  45. 45.Kumar, A., Zhou, A., Tucker, G., and Levine, S. Conservative q-learning for offline reinforcement learning. In Proceedings of the 34th International Conference on Neural Information Processing Systems, NIPS’20, Red Hook, NY, USA, 2020. Curran Associates Inc. ISBN 9781713829546.
  46. 46.Levesque, H. J. The winograd schema challenge. In AAAI Spring Symposium: Logical Formalizations of Commonsense Reasoning. AAAI, 2011. URL http://dblp.uni-trier.de/db/conf/aaaiss/aaaiss2011-6.html#Levesque11.
  47. 47.Levine, S., Kumar, A., Tucker, G., and Fu, J. Offline reinforcement learning: Tutorial, review, and perspectives on open problems, 2020. URL https://arxiv.org/abs/2005.01643.
  48. 48.Li, J., Galley, M., Brockett, C., Gao, J., and Dolan, B. A diversity-promoting objective function for neural conversation models. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 110–119, San Diego, California, June 2016. Association for Computational Linguistics. doi: 10.18653/v1/N16-1014. URL https://aclanthology.org/N16-1014.
  49. 49.Lin, S., Hilton, J., and Evans, O. TruthfulQA: Measuring how models mimic human falsehoods. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 3214–3252, Dublin, Ireland, May 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.acl-long.229. URL https://aclanthology.org/2022.acl-long.229.
  50. 50.Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V. Roberta: A robustly optimized bert pretraining approach, 2019. URL https://arxiv.org/abs/1907.11692.
  51. 51.Lu, X., Welleck, S., Hessel, J., Jiang, L., Qin, L., West, P., Ammanabrolu, P., and Choi, Y. QUARK: Controllable text generation with reinforced unlearning. In Oh, A. H., Agarwal, A., Belgrave, D., and Cho, K. (eds.), Advances in Neural Information Processing Systems, 2022. URL https://openreview.net/forum?id=5HaIds3ux5O.
  52. 52.Menick, J., Trebacz, M., Mikulik, V., Aslanides, J., Song, F., Chadwick, M., Glaese, M., Young, S., Campbell-Gillingham, L., Irving, G., and McAleese, N. Teaching language models to support answers with verified quotes, 2022. URL https://arxiv.org/abs/2203.11147.
  53. 53.Mikolov, T. and Zweig, G. Context dependent recurrent neural network language model. In 2012 IEEE Spoken Language Technology Workshop (SLT), pp. 234–239, 2012. doi: 10.1109/SLT.2012.6424228.
  54. 54.Nair, A., Gupta, A., Dalal, M., and Levine, S. Awac: Accelerating online reinforcement learning with offline datasets, 2020. URL https://arxiv.org/abs/2006.09359.
  55. 55.Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Gray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R. Training language models to follow instructions with human feedback. In Oh, A. H., Agarwal, A., Belgrave, D., and Cho, K. (eds.), Advances in Neural Information Processing Systems, 2022. URL https://openreview.net/forum?id=TG8KACxEON.
  56. 56.Paperno, D., Kruszewski, G., Lazaridou, A., Pham, N. Q., Bernardi, R., Pezzelle, S., Baroni, M., Boleda, G., and Fernandez, R. The LAMBADA dataset: Word prediction ´ requiring a broad discourse context. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 1525–1534, Berlin, Germany, August 2016. Association for Computational Linguistics. doi: 10.18653/v1/P16-1144. URL https://aclanthology.org/P16-1144.
  57. 57.Peng, N., Ghazvininejad, M., May, J., and Knight, K. Towards controllable story generation. In Proceedings of the First Workshop on Storytelling, pp. 43–49, New Orleans, Louisiana, June 2018. Association for Computational Linguistics. doi: 10.18653/v1/W18-1505. URL https://aclanthology.org/W18-1505.
  58. 58.Peng, X. B., Kumar, A., Zhang, G., and Levine, S. Advantage-weighted regression: Simple and scalable off-policy reinforcement learning, 2019. URL https://arxiv.org/abs/1910.00177.
  59. 59.Perez, E., Huang, S., Song, F., Cai, T., Ring, R., Aslanides, J., Glaese, A., McAleese, N., and Irving, G. Red teaming language models with language models, 2022. URL https://arxiv.org/abs/2202.03286.
  60. 60.Peters, J. and Schaal, S. Reinforcement learning by reward-weighted regression for operational space control. In Proceedings of the 24th International Conference on Machine Learning, ICML ’07, pp. 745–750, New York, NY, USA, 2007. Association for Computing Machinery. ISBN 9781595937933. doi: 10.1145/1273496.1273590. URL https://doi.org/10.1145/1273496.1273590.
  61. 61.Radford, A. and Narasimhan, K. Improving language understanding by generative pre-training. 2018.
  62. 62.Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I. Language models are unsupervised multitask learners. 2019.
  63. 63.Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P. SQuAD: 100,000+ questions for machine comprehension of text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pp. 2383–2392, Austin, Texas, November 2016. Association for Computational Linguistics. doi: 10.18653/v1/D16-1264. URL https://aclanthology.org/D16-1264.
  64. 64.Ramasesh, V. V., Lewkowycz, A., and Dyer, E. Effect of scale on catastrophic forgetting in neural networks. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=GhVS8_yPeEa.
  65. 65.Sanh, V., Webson, A., Raffel, C., Bach, S., Sutawika, L., Alyafeai, Z., Chaffin, A., Stiegler, A., Raja, A., Dey, M., Bari, M. S., Xu, C., Thakker, U., Sharma, S. S., Szczechla, E., Kim, T., Chhablani, G., Nayak, N., Datta, D., Chang, J., Jiang, M. T.-J., Wang, H., Manica, M., Shen, S., Yong, Z. X., Pandey, H., Bawden, R., Wang, T., Neeraj, T., Rozen, J., Sharma, A., Santilli, A., Fevry, T., Fries, J. A., Teehan, R., Scao, T. L., Biderman, S., Gao, L., Wolf, T., and Rush, A. M. Multitask prompted training enables zero-shot task generalization. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=9Vrb9D0WI4.
  66. 66.Sap, M., Card, D., Gabriel, S., Choi, Y., and Smith, N. A. The risk of racial bias in hate speech detection. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pp. 1668–1678, Florence, Italy, July 2019. Association for Computational Linguistics. doi: 10.18653/v1/P19-1163. URL https://aclanthology.org/P19-1163.
  67. 67.Scheurer, J., Campos, J. A., Chan, J. S., Chen, A., Cho, K., and Perez, E. Training language models with language feedback, 2022. URL https://arxiv.org/abs/2204.14146.
  68. 68.Scheurer, J., Campos, J. A., Korbak, T., Chan, J. S., Chen, A., Cho, K., and Perez, E. Training language models with language feedback at scale, 2023.
  69. 69.Schmidhuber, J. Reinforcement learning upside down: Don’t predict rewards – just map them to actions, 2019. URL https://arxiv.org/abs/1912.02875.
  70. 70.Snell, C., Kostrikov, I., Su, Y., Yang, M., and Levine, S. Offline rl for natural language generation with implicit language q learning, 2022. URL https://arxiv.org/abs/2206.11871.
  71. 71.Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A., and Potts, C. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, pp. 1631–1642, Seattle, Washington, USA, October 2013. Association for Computational Linguistics. URL https://aclanthology.org/D13-1170.
  72. 72.Solaiman, I. and Dennison, C. Process for adapting language models to society (PALMS) with values-targeted datasets. In Beygelzimer, A., Dauphin, Y., Liang, P., and Vaughan, J. W. (eds.), Advances in Neural Information Processing Systems, 2021. URL https://openreview.net/forum?id=k-ghaB9VZBw.
  73. 73.Tang, R., Lu, Y., Liu, L., Mou, L., Vechtomova, O., and Lin, J. J. Distilling task-specific knowledge from bert into simple neural networks. ArXiv, abs/1903.12136, 2019.
  74. 74.Tay, Y., Wei, J., Chung, H. W., Tran, V. Q., So, D. R., Shakeri, S., Garcia, X., Zheng, H. S., Rao, J., Chowdhery, A., Zhou, D., Metzler, D., Petrov, S., Houlsby, N., Le, Q. V., and Dehghani, M. Transcending scaling laws with 0.1 URL https://arxiv.org/abs/2210.11399.
  75. 75.Tunstall, L., von Werra, L., and Wolf, T. Natural Language Processing with Transformers: Building Language Applications with Hugging Face. O’Reilly Media, Incorporated, 2022. ISBN 1098103246. URL https://books.google.ch/books?id=7hhyzgEACAAJ.
  76. 76.van Rossum, G., Warsaw, B., and Coghlan, N. Style guide for Python code. PEP 8, 2001. URL https://www.python.org/dev/peps/pep-0008/.
  77. 77.Villalobos, P., Sevilla, J., Heim, L., Besiroglu, T., Hobbhahn, M., and Ho, A. Will we run out of data? an analysis of the limits of scaling datasets in machine learning, 2022. URL https://arxiv.org/abs/2211.04325.
  78. 78.Vu, T., Barua, A., Lester, B., Cer, D., Iyyer, M., and Constant, N. Overcoming catastrophic forgetting in zero-shot cross-lingual generation. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pp. 9279–9300, Abu Dhabi, United Arab Emirates, December 2022. Association for Computational Linguistics. URL https://aclanthology.org/2022.emnlp-main.630.
  79. 79.Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. GLUE: A multi-task benchmark and analysis platform for natural language understanding. In Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pp. 353–355, Brussels, Belgium, November 2018. Association for Computational Linguistics. doi: 10.18653/v1/W18-5446. URL https://aclanthology.org/W18-5446.
  80. 80.Wang, B., Ping, W., Xiao, C., Xu, P., Patwary, M., Shoeybi, M., Li, B., Anandkumar, A., and Catanzaro, B. Exploring the limits of domain-adaptive training for detoxifying large-scale language models. In Oh, A. H., Agarwal, A., Belgrave, D., and Cho, K. (eds.), Advances in Neural Information Processing Systems, 2022. URL https://openreview.net/forum?id=v_0F4IZJZw.
  81. 81.Warstadt, A., Singh, A., and Bowman, S. R. Neural network acceptability judgments. arXiv preprint arXiv:1805.12471, 2018.
  82. 82.Welbl, J., Glaese, A., Uesato, J., Dathathri, S., Mellor, J., Hendricks, L. A., Anderson, K., Kohli, P., Coppin, B., and Huang, P.-S. Challenges in detoxifying language models. In Findings of the Association for Computational Linguistics: EMNLP 2021, pp. 2447–2469, Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.findings-emnlp.210. URL https://aclanthology.org/2021.findings-emnlp.210.
  83. 83.Welleck, S., Kulikov, I., Roller, S., Dinan, E., Cho, K., and Weston, J. Neural text generation with unlikelihood training. In International Conference on Learning Representations, 2020. URL https://openreview.net/forum?id=SJeYe0NtvH.
  84. 84.Williams, A., Nangia, N., and Bowman, S. A broad-coverage challenge corpus for sentence understanding through inference. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pp. 1112–1122, New Orleans, Louisiana, June 2018. Association for Computational Linguistics. doi: 10.18653/v1/N18-1101. URL https://aclanthology.org/N18-1101.
  85. 85.Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Davison, J., Shleifer, S., von Platen, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Le Scao, T., Gugger, S., Drame, M., Lhoest, Q., and Rush, A. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pp. 38–45, Online, October 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.emnlp-demos.6. URL https://aclanthology.org/2020.emnlp-demos.6.
  86. 86.Xu, A., Pathak, E., Wallace, E., Gururangan, S., Sap, M., and Klein, D. Detoxifying language models risks marginalizing minority voices. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 2390–2397, Online, June 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.naacl-main.190. URL https://aclanthology.org/2021.naacl-main.190.
  87. 87.Xu, J., Ju, D., Li, M., Boureau, Y.-L., Weston, J., and Dinan, E. Recipes for safety in open-domain chatbots, 2020. URL https://arxiv.org/abs/2010.07079.
  88. 88.Zhu, Y., Lu, S., Zheng, L., Guo, J., Zhang, W., Wang, J., and Yu, Y. Texygen: A benchmarking platform for text generation models. In The 41st International ACM SIGIR Conference on Research & Development in Information Retrieval, pp. 1097–1100, 2018.
  89. 89.Ziegler, D., Nix, S., Chan, L., Bauman, T., Schmidt-Nielsen, P., Lin, T., Scherlis, A., Nabeshima, N., Weinstein-Raun, B., de Haas, D., Shlegeris, B., and Thomas, N. Adversarial training for high-stakes reliability. In Oh, A. H., Agarwal, A., Belgrave, D., and Cho, K. (eds.), Advances in Neural Information Processing Systems, 2022. URL https://openreview.net/forum?id=NtJyGXo0nF.
  90. 90.Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G. Fine-tuning language models from human preferences. arXiv preprint arXiv:1909.08593, 2019.

Citation

MLA
Korbak, T., et al. “Pretraining Language Models with Human Preferences”. International Conference on Machine Learning, vol. 202, 2023, pp. 17506–33, https://proceedings.mlr.press/v202/korbak23a.html.
APA
Korbak, T., Shi, K., Chen, A., Bhalerao, R. V., Buckley, C., Phang, J., Bowman, S. R., & Perez, E. (2023). Pretraining Language Models with Human Preferences. International Conference on Machine Learning, 202, 17506–17533. https://proceedings.mlr.press/v202/korbak23a.html
Chicago
Korbak, T., K. Shi, A. Chen, et al. 2023. “Pretraining Language Models with Human Preferences”. International Conference on Machine Learning 202: 17506–33. https://proceedings.mlr.press/v202/korbak23a.html.
Harvard
Korbak, T. et al. (2023) “Pretraining Language Models with Human Preferences”, International Conference on Machine Learning. PMLR, pp. 17506–17533. Available at: https://proceedings.mlr.press/v202/korbak23a.html.
Vancouver
1. Korbak T, Shi K, Chen A, Bhalerao RV, Buckley C, Phang J, Bowman SR, Perez E (2023) Pretraining Language Models with Human Preferences. In: International Conference on Machine Learning. PMLR, pp 17506–17533

BibTeX

@InProceedings{pmlr-v202-korbak23a,
  title = 	 {Pretraining Language Models with Human Preferences},
  author =       {Korbak, Tomasz and Shi, Kejian and Chen, Angelica and Bhalerao, Rasika Vinayak and Buckley, Christopher and Phang, Jason and Bowman, Samuel R. and Perez, Ethan},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {17506--17533},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/korbak23a/korbak23a.pdf},
  url = 	 {https://proceedings.mlr.press/v202/korbak23a.html},
  abstract = 	 {Language models (LMs) are pretrained to imitate text from large and diverse datasets that contain content that would violate human preferences if generated by an LM: falsehoods, offensive comments, personally identifiable information, low-quality or buggy code, among others. Here, we explore alternative objectives for pretraining LMs in a way that also guides them to generate text aligned with human preferences. We benchmark five objectives for pretraining with human feedback across three tasks and study how they affect the alignment and capabilities of pretrained LMs. We find a Pareto-optimal and simple approach among those we explored: conditional training, or learning distribution over tokens conditional on their human preference scores. Conditional training reduces the rate of undesirable content by up to an order of magnitude, both when generating without a prompt and with an adversarially-chosen prompt. Moreover, conditional training maintains the downstream task performance of standard LM pretraining, both before and after task-specific finetuning. Pretraining with human feedback results in much better preference satisfaction than standard LM pretraining followed by finetuning with feedback, i.e., learning and then unlearning undesirable behavior. Our results suggest that we should move beyond imitation learning when pretraining LMs and incorporate human preferences from the start of training.}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/