Pretraining Language Models with Human Preferences
Tomasz KorbakKejian ShiAngelica ChenRasika Vinayak BhaleraoChristopher L. BuckleyJason PhangSamuel R. BowmanEthan Perez
Demonstrates that integrating human preferences directly into language model pretraining via conditional training reduces undesirable outputs by up to an order of magnitude while preserving downstream capabilities, outperforming the standard pipeline of pretraining followed by post-hoc alignment fine-tuning.
Modern language models are typically pretrained by imitating vast amounts of uncurated internet text. Consequently, they often internalize and reproduce harmful behaviors, such as generating offensive language, leaking personally identifiable information, and producing flawed code. The standard industry approach attempts to fix these issues after pretraining through safety filters or post-hoc finetuning techniques like reinforcement learning from human feedback. However, because large models strongly resist unlearning their initial training data, these downstream adjustments often fail or require costly interventions. Preemptively filtering datasets also introduces severe trade-offs, such as bottlenecking data scale, reducing model capabilities, and amplifying biases.
The article evaluates whether incorporating human preferences directly into the pretraining phase—a framework termed Pretraining with Human Feedback—can steer models away from generating undesirable text while preserving their overall performance and knowledge. Across a compute-optimal scale of 3.32 billion tokens and 124-million-parameter models, the authors benchmark standard imitation learning against five preference-guided pretraining objectives: conditional training, dataset filtering, unlikelihood loss, reward-weighted regression, and advantage-weighted regression. These objectives were systematically evaluated across three distinct tasks: mitigating toxic speech, preventing personal data leakage, and adhering to the PEP8 Python coding standard.
The investigation produced four central findings. First, conditional training—a technique that prepends segments with control tokens indicating preference scores—emerged as the most effective and Pareto-optimal approach across all tasks. Second, conditional training reduced the generation of undesirable content by up to an order of magnitude (for example, reducing toxicity scores from 0.0141 down to 0.0011) and demonstrated continuous improvement as training progressed without plateauing. Third, this method preserved general capabilities, matching standard imitation pretraining on zero-shot passage understanding and downstream language classification benchmarks. Finally, pretraining with human feedback substantially outperformed the conventional pipeline of standard pretraining followed by post-hoc finetuning, proving two to three times more effective at suppressing undesirable outputs and maintaining superior robustness against automated adversarial red-teaming.
These results demonstrate that language models should be aligned from the beginning of training rather than taught bad behaviors only to attempt unlearning them later. Conditional training enables models to retain broad knowledge from lower-quality or toxic text without imitating it during generation. In practice, this shifts the cost and risk structure of model development: the additional computational overhead of running segment-level reward scoring during data ingestion is minimal compared to the compounding costs and security risks of deploying poorly aligned base models.
Organizations developing or deploying foundation models should consider moving beyond pure imitation pretraining by adopting segment-level conditional training during initial pretraining runs. When planning training budgets, teams should incorporate reward scoring pipelines into early data preparation workflows rather than relying solely on post-hoc alignment stages. However, decision-makers must note that while pretraining with feedback dramatically improves safety and adversarial robustness, it does not guarantee complete immunity against determined adversarial prompts. Further validation at larger model parameter scales (such as tens or hundreds of billions of parameters) is recommended to confirm that these scaling trends fully generalize to enterprise-scale deployments.
- Paper: Fine-Tuning Language Models from Human Preferences, Daniel M. Ziegler et al. (2019). This foundational study establishes the preference-modeling and reinforcement-learning pipeline that the source contrasts with its pretraining-time alternatives.
- Paper: Training language models to follow instructions with human feedback, Long Ouyang et al. (2022). Its InstructGPT framework supplies the post-training human-feedback baseline against which pretraining with preferences is evaluated.
- Paper: Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback, Yuntao Bai et al. (2022). This work develops the helpfulness–harmlessness RLHF framework and evaluation logic that the source seeks to move earlier into pretraining.
- Paper: RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models, Samuel Gehman et al. (2020). Its RealToxicityPrompts benchmark provides the toxicity-generation problem and measurement context used by the source's safety experiments.
- Paper: Red Teaming Language Models with Language Models, Ethan Perez et al. (2022). Its automated red-teaming methodology motivates the source's emphasis on robustness against adversarial safety evaluation.
- Paper: CTRL: A Conditional Transformer Language Model for Controllable Generation, Nitish Shirish Keskar et al. (2019). CTRL demonstrates the control-token conditioning principle that directly underlies the source's preferred conditional-pretraining objective.
- Paper: Direct Preference Optimization: Your Language Model is Secretly a Reward Model, Rafael Rafailov et al. (2023). DPO extends preference alignment beyond the source's pretraining objectives by directly optimizing preference pairs without an explicit reward-model-and-RL pipeline.
- Paper: RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback, Harrison Lee et al. (2024). RLAIF continues the source's effort to scale preference supervision by replacing expensive human judgments with AI-generated feedback.
- Paper: Safe RLHF: Safe Reinforcement Learning from Human Feedback, Josef Dai et al. (2024). Safe RLHF extends preference alignment by separating helpfulness rewards from harmlessness costs rather than treating them as a single objective.
- Paper: Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!, Xiangyu Qi et al. (2023). This study stress-tests the source's premise by showing how subsequent fine-tuning can undo safety alignment established during earlier training.
- Paper: Fine-Tuning Language Models for Factuality, Katherine Tian et al. (2024). It applies preference optimization to factuality, extending the source's safety-focused pretraining framework to a distinct reliability objective.
- Paper: HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal, Mantas Mazeika et al. (2024). HarmBench extends the source's red-team evaluation concerns into a standardized framework for comparing attacks and refusal defenses.
