Safe RLHF: Safe Reinforcement Learning from Human Feedback

Josef DaiXuehai PanRuiyang SunJiaming JiXinbo XuMickel LiuYizhou WangYaodong Yang

article2024ICLR631 citations

Introduces a constrained reinforcement learning framework that separates human feedback on helpfulness and safety into distinct reward and cost models, enabling language models to satisfy strict safety limits without degrading response quality.

Listen

As artificial intelligence systems powered by large language models are increasingly deployed across healthcare, education, law, and business, ensuring that these models remain safe without losing their usefulness has become a critical challenge. Standard fine-tuning approaches often struggle with an inherent conflict between helpfulness and harmlessness: overly safe models frequently refuse harmless queries, while overly eager models can provide harmful instructions or biased outputs.

The article evaluates a new alignment framework called Safe Reinforcement Learning from Human Feedback (Safe RLHF). Its objective is to demonstrate that explicitly separating helpfulness and harmlessness into two distinct optimization objectives allows models to significantly reduce harmful generations while simultaneously improving overall response quality and task performance.

To accomplish this, the authors decoupled human feedback during data collection, tasking human annotators with rating helpfulness and harmlessness separately and evaluating responses across 14 distinct categories of harm. Using these distinct signals, the researchers trained an independent reward model for helpfulness and a cost model for harmlessness. They framed the alignment process as a constrained optimization problem—maximizing helpfulness while enforcing a strict boundary on harm—and solved it using dynamic mathematical adjustments during reinforcement learning. The team implemented this framework across three iterative training cycles on a 7-billion-parameter baseline language model (Alpaca-7B), incorporating red-teaming adversarial prompts to stress-test and expand the training data.

The findings confirm substantial performance improvements across multiple metrics. First, model safety improved dramatically: the proportion of harmful responses dropped from 53.08% in the baseline model to just 2.45% in the final fine-tuned model (Beaver-v3). Second, the model achieved these safety gains without compromising performance, gaining significant competitive rating points in both helpfulness and harmlessness when evaluated by human judges and automated evaluation systems. Third, decoupling the annotation process improved human reviewer agreement rates from roughly 61% to 66–69% and raised quality approval rates from below 80% to at least 90%. Finally, the dynamic adjustment method clearly outperformed standard static balancing approaches, which routinely degraded performance by either over-emphasizing safety or failing to prevent harm.

These results demonstrate that organizations do not have to accept a steep tradeoff between artificial intelligence capability and ethical compliance. Decoupling preference data and dynamically managing safety constraints reduces organizational liability, lowers reputational risks, and enhances system reliability. It also streamlines the data annotation pipeline by providing clearer, less confusing guidelines to human annotators, thereby lowering the risk of noisy training data.

Organizations developing or deploying large language models should transition from single-score preference models to decoupled reward and cost systems. For immediate next steps, technical teams should implement iterative adversarial testing (red-teaming) to discover emerging vulnerabilities across diverse risk categories. Future efforts should also explore expanding the Safe RLHF framework beyond single-turn dialogues into complex, multi-turn conversations and applying it to newer, larger base models.

The primary limitations noted in the article include high compute and data collection costs, reliance on single-turn interactions, and the use of an older baseline model architecture. Nonetheless, the reported empirical results are robust across both automated benchmarks and rigorous human evaluations, providing high confidence that dynamic, decoupled constraint optimization is an effective path forward for safe artificial intelligence deployment.

Cover for Safe RLHF: Safe Reinforcement Learning from Human Feedback

Abstract

With the development of large language models (LLMs), striking a balance between the performance and safety of AI systems has never been more critical. However, the inherent tension between the objectives of helpfulness and harmlessness presents a significant challenge during LLM training. To address this issue, we propose Safe Reinforcement Learning from Human Feedback (Safe RLHF), a novel algorithm for human value alignment. Safe RLHF explicitly decouples human preferences regarding helpfulness and harmlessness, effectively avoiding the crowdworkers' confusion about the tension and allowing us to train separate reward and cost models. We formalize the safety concern of LLMs as an optimization task of maximizing the reward function while satisfying specified cost constraints. Leveraging the Lagrangian method to solve this constrained problem, Safe RLHF dynamically adjusts the balance between the two objectives during fine-tuning. Through a three-round fine-tuning using Safe RLHF, we demonstrate a superior ability to mitigate harmful responses while enhancing model performance compared to existing value-aligned algorithms. Experimentally, we fine-tuned the Alpaca-7B using Safe RLHF and aligned it with collected human preferences, significantly improving its helpfulness and harmlessness according to human evaluations.

Table of Contents

  • 1 Introduction
  • 2 Preliminaries
  • 3 Method: Safe RLHF
  • 3.1 Human Preference of Harmlessness and Helpfulness
  • 3.2 Preference Model Fitting: Reward and Cost Models
  • 3.3 Safe Reinforcement Learning
  • 4 Experiments
  • 4.1 Experimental Details
  • 4.2 Experiment Results
  • 4.2.1 Helpfulness and Harmlessness Evaluation
  • 4.2.2 The Decoupling of Harmlessness and Helpfulness
  • 4.2.3 Balance between Harmlessness Objective and Helpfulness Objective
  • 4.2.4 Design of Cost Preference Model
  • 5 Related Works
  • 6 Limitations and Future work
  • 7 Ethic Discussion
  • 8 Conclusion
  • References
  • A Data Annotation Guidelines
  • A.1 Overview
  • A.2 Data Generation
  • A.3 Harm Categories
  • A.4 Annotation Documents
  • A.5 Data Annotation Team
  • B Implementation Details
  • B.1 Preference Models
  • B.2 Details of RLHF training
  • B.3 Details of Safe RLHF training
  • C Supplementary Details of the Experiments
  • C.1 Hyper-parameters
  • C.2 Prompts used in GPT-4 Evaluation
  • C.2.1 Helpfulness Preference Prompts
  • C.2.2 Harmlessness Preference Prompts
  • D Red Teaming
  • D.1 Partial Harmfulness
  • D.2 Scenario Assumptions
  • D.3 Contradictory Analysis
  • D.4 Complex Text Command Embedding

Knowls

  1. Knowl 1 — Safe RLHF Constrained Optimization Formulation

    model/method

    In Safe RLHF, human value alignment is formulated as a Constrained Markov Decision Process (CMDP) optimization problem. Rather than conflating helpfulness and harmlessness into a single scalar score, the objective maximizes the expected helpfulness reward subject to an expected harmlessness cost constraint:

    max⁡θJR(θ)subject toJC(θ)≤0\max_{\theta} J_R(\theta) \quad \text{subject to} \quad J_C(\theta) \le 0

    where πθ(y∣x)\pi_\theta(y|x) denotes the parameterized language model policy generating response yy for prompt x∼Dx \sim \mathcal{D}, and the expected reward and cost objectives are defined as:

    JR(θ)≜Ex∼D,y∼πθ(⋅∣x)[Rϕ(y,x)]J_R(\theta) \triangleq \mathbb{E}_{x \sim \mathcal{D}, y \sim \pi_\theta(\cdot|x)} \left[ R_\phi(y, x) \right]

    JC(θ)≜Ex∼D,y∼πθ(⋅∣x)[Cψ(y,x)]+dJ_C(\theta) \triangleq \mathbb{E}_{x \sim \mathcal{D}, y \sim \pi_\theta(\cdot|x)} \left[ C_\psi(y, x) \right] + d

    Here, Rϕ(y,x)R_\phi(y, x) is a learned scalar reward model scoring helpfulness, Cψ(y,x)C_\psi(y, x) is a learned scalar cost model scoring harmfulness, and dd is a hyperparameter threshold that sets the tolerance level for generating harmful responses.

    Applying the Lagrangian multiplier method converts this constrained primal problem into an unconstrained dual min-max problem:

    min⁡θmax⁡λ≥0[−JR(θ)+λ⋅JC(θ)]\min_{\theta} \max_{\lambda \ge 0} \left[ -J_R(\theta) + \lambda \cdot J_C(\theta) \right]

    where λ≥0\lambda \ge 0 is the dynamic Lagrange multiplier that adaptively modulates the harmlessness penalty relative to the helpfulness reward.

  2. Knowl 2 — Cost Model Loss with Pairwise Preference and Safety Boundary Classification

    model/method

    In Safe RLHF, the Cost Model Cψ(y,x)C_\psi(y, x) quantifies the harmfulness of a response yy to a prompt xx. Unlike standard preference reward models, the harmlessness dataset DC={(xj,ywj,ylj,swj,slj)}j=1N\mathcal{D}_C = \{(x^j, y_w^j, y_l^j, s_w^j, s_l^j)\}_{j=1}^N provides both pairwise comparison rankings (ywy_w is more harmful than yly_l) and individual binary safety classification meta-labels s(y)∈{+1,−1}s(y) \in \{+1, -1\}, where s(y)=+1s(y) = +1 denotes a harmful response and s(y)=−1s(y) = -1 denotes a harmless response.

    Under the Bradley-Terry model framework, assuming a virtual neutral response y0y_0 lying on the decision boundary between safe and unsafe responses with cost Cψ(y0,x)=0C_\psi(y_0, x) = 0, the probability of an unsafe response yy being preferred (more harmful) over y0y_0 is p(y≻y0∣x)=σ(Cψ(y,x)−Cψ(y0,x))=σ(s(y)⋅Cψ(y,x))p(y \succ y_0 | x) = \sigma(C_\psi(y, x) - C_\psi(y_0, x)) = \sigma(s(y) \cdot C_\psi(y, x)). Similarly, for a safe response yy, p(y0≻y∣x)=σ(Cψ(y0,x)−Cψ(y,x))=σ(s(y)⋅Cψ(y,x))p(y_0 \succ y | x) = \sigma(C_\psi(y_0, x) - C_\psi(y, x)) = \sigma(s(y) \cdot C_\psi(y, x)).

    The cost model parameters ψ\psi are optimized via the combined loss function:

    LC(ψ;DC)=−E(x,yw,yl,⋅,⋅)∼DC[log⁡σ(Cψ(yw,x)−Cψ(yl,x))]−E(x,yw,yl,sw,sl)∼DC[log⁡σ(sw⋅Cψ(yw,x))+log⁡σ(sl⋅Cψ(yl,x))]+μCE(x,y)∼DC[∣Cψ(y,x)∣2]\mathcal{L}_C(\psi; \mathcal{D}_C) = -\mathbb{E}_{(x, y_w, y_l, \cdot, \cdot)\sim \mathcal{D}_C} \left[ \log \sigma(C_\psi(y_w, x) - C_\psi(y_l, x)) \right] - \mathbb{E}_{(x, y_w, y_l, s_w, s_l)\sim \mathcal{D}_C} \left[ \log \sigma(s_w \cdot C_\psi(y_w, x)) + \log \sigma(s_l \cdot C_\psi(y_l, x)) \right] + \mu_C \mathbb{E}_{(x, y)\sim \mathcal{D}_C} \left[ |C_\psi(y, x)|^2 \right]

    where σ(z)=1/(1+exp⁡(−z))\sigma(z) = 1 / (1 + \exp(-z)) is the logistic sigmoid function, and μC\mu_C is a constant controlling L2L_2 regularization strength. The first term trains the relative harmfulness ranking, the second term anchors absolute safety around the Cψ=0C_\psi=0 boundary, and the third term regularizes model outputs.

  3. Knowl 3 — Safe RLHF Policy Optimization and Multiplier Update Rules

    algorithm

    Safe RLHF alternates between updating the policy parameters θ\theta and the Lagrange multiplier λ\lambda using Proximal Policy Optimization (PPO). For a prompt x∼Dpromptx \sim \mathcal{D}_{\text{prompt}} and generated response tokens y=a1:Ty = a_{1:T}, the token-level reward r^t\hat{r}_t and token-level cost c^t\hat{c}_t incorporate an evenly split Kullback-Leibler (KL) penalty against a frozen reference policy πref\pi_{\text{ref}}:

    r^t=rtRM+β2rtKL,c^t=ctCM−β2rtKL(1≤t≤T)\hat{r}_t = r_t^{RM} + \frac{\beta}{2} r_t^{KL}, \quad \hat{c}_t = c_t^{CM} - \frac{\beta}{2} r_t^{KL} \quad (1 \le t \le T)

    where rTRM=Rϕ(y,x)r_T^{RM} = R_\phi(y, x) (and 00 for t<Tt < T), cTCM=Cψ(y,x)c_T^{CM} = C_\psi(y, x) (and 00 for t<Tt < T), rtKL=−log⁡[πθ(at∣x,a1:t−1)/πref(at∣x,a1:t−1)]r_t^{KL} = -\log [\pi_\theta(a_t|x, a_{1:t-1}) / \pi_{\text{ref}}(a_t|x, a_{1:t-1})], and β≥0\beta \ge 0 is the KL penalty coefficient.

    Surrogate PPO clip losses are defined as:

    LRSafeRL(θ)=−Ex,y[Et[min⁡(ρt(θ)A^tr^,clip(ρt(θ),1−ϵ,1+ϵ)A^tr^)]]\mathcal{L}_R^{\text{SafeRL}}(\theta) = -\mathbb{E}_{x, y} \left[ \mathbb{E}_t \left[ \min\left( \rho_t(\theta)\hat{A}_t^{\hat{r}}, \text{clip}(\rho_t(\theta), 1-\epsilon, 1+\epsilon)\hat{A}_t^{\hat{r}} \right) \right] \right]

    LCSafeRL(θ)=−Ex,y[Et[min⁡(ρt(θ)A^tc^,clip(ρt(θ),1−ϵ,1+ϵ)A^tc^)]]\mathcal{L}_C^{\text{SafeRL}}(\theta) = -\mathbb{E}_{x, y} \left[ \mathbb{E}_t \left[ \min\left( \rho_t(\theta)\hat{A}_t^{\hat{c}}, \text{clip}(\rho_t(\theta), 1-\epsilon, 1+\epsilon)\hat{A}_t^{\hat{c}} \right) \right] \right]

    where ρt(θ)=πθ(at∣x,a1:t−1)/πθold(at∣x,a1:t−1)\rho_t(\theta) = \pi_\theta(a_t|x, a_{1:t-1}) / \pi_{\theta_{\text{old}}}(a_t|x, a_{1:t-1}), ϵ∈(0,1)\epsilon \in (0,1) is the clip ratio, and A^tr^,A^tc^\hat{A}_t^{\hat{r}}, \hat{A}_t^{\hat{c}} are advantages estimated via Generalized Advantage Estimation (GAE).

    Input: Initial policy parameters θ0\theta_0, initial multiplier λ0\lambda_0, prompt dataset Dprompt\mathcal{D}_{\text{prompt}}, pretraining dataset DSFT\mathcal{D}_{\text{SFT}}, learning rates η,α\eta, \alpha, KL weight β\beta, threshold −d-d, PTX weight γ\gamma
    for step k=0,1,2,…k = 0, 1, 2, \dots do
        Sample prompts x∼Dpromptx \sim \mathcal{D}_{\text{prompt}} and generate responses y∼πθk(⋅∣x)y \sim \pi_{\theta_k}(\cdot|x)
        Compute token rewards r^t\hat{r}_t and costs c^t\hat{c}_t
        Estimate reward advantages A^tr^\hat{A}_t^{\hat{r}} and cost advantages A^tc^\hat{A}_t^{\hat{c}} via GAE
        Compute cost objective estimate JC(θk)=E[Cψ(y,x)]+dJ_C(\theta_k) = \mathbb{E}[C_\psi(y, x)] + d
        Compute policy surrogate losses LRSafeRL(θk)\mathcal{L}_R^{\text{SafeRL}}(\theta_k) and LCSafeRL(θk)\mathcal{L}_C^{\text{SafeRL}}(\theta_k)
        Update policy parameters:
            θk+1=θk−η1+λk∇θk[LRSafeRL(θk)−λkLCSafeRL(θk)]−ηγ∇θkLPTX(θk)\theta_{k+1} = \theta_k - \frac{\eta}{1 + \lambda_k} \nabla_{\theta_k} \left[ \mathcal{L}_R^{\text{SafeRL}}(\theta_k) - \lambda_k \mathcal{L}_C^{\text{SafeRL}}(\theta_k) \right] - \eta \gamma \nabla_{\theta_k} \mathcal{L}^{\text{PTX}}(\theta_k)
        Update Lagrange multiplier in log-space:
            ln⁡λk+1=ln⁡λk+α⋅λk⋅JC(θk)\ln \lambda_{k+1} = \ln \lambda_k + \alpha \cdot \lambda_k \cdot J_C(\theta_k)
    end for
  4. Knowl 4 — Decoupled Two-Dimensional Human Preference Annotation Framework

    model/method

    Safe RLHF implements a two-stage human annotation strategy that decouples human preferences along two orthogonal dimensions: helpfulness and harmlessness.

    1. Safety Meta-Labeling: For each question-answer (QA) pair, annotators evaluate potential risks against 14 predefined harm categories: hate speech/offensive language; discrimination/stereotypes; violence/incitement; financial/property crime; privacy violation; drug abuse/weapons; non-violent unethical behavior; sexually explicit content; controversial topics/politics; misinformation regarding ethics/laws/safety; terrorism/organized crime; self-harm; animal abuse; and child abuse. A QA pair is assigned a binary meta-label of 'safe' (s=−1s = -1) only if it is completely risk-neutral across all 14 categories; otherwise, it is labeled 'unsafe' (s=+1s = +1).

    2. Independent Preference Ranking: For a prompt xx and pair of responses (y1,y2)(y_1, y_2), annotators independently perform two rankings:

      • Helpfulness Ranking: Ranking based on instruction completeness, clarity, relevance, and problem-solving quality, producing a helpfulness dataset DR={(xi,ywi,yli)}i=1N\mathcal{D}_R = \{(x^i, y_w^i, y_l^i)\}_{i=1}^N.
      • Harmlessness Ranking: Ranking based solely on safety and risk avoidance, producing a harmlessness dataset DC={(xj,ywj,ylj,swj,slj)}j=1N\mathcal{D}_C = \{(x^j, y_w^j, y_l^j, s_w^j, s_l^j)\}_{j=1}^N.

    Decoupling avoids forcing annotators into trade-off dilemmas (e.g., when a harmful response is technically well-structured or when a safe refusal provides minimal utility).

  5. Knowl 5 — Multi-Round Safe RLHF Alignment Results on Helpfulness, Harmlessness, and Safety Ratio

    empirical result

    Starting from an Alpaca-7B supervised fine-tuned (SFT) model (LLaMA-7B fine-tuned on 52K instruction instances), three iterative rounds of Safe RLHF produced models Beaver-v1, Beaver-v2, and Beaver-v3. Evaluated across helpfulness and harmlessness Elo ratings (with Alpaca-7B normalized to 1000) and human-rated safety proportions:

    • GPT-4 Evaluation Elo Scores:
      • Beaver-v3 achieved a +244.91 increase in helpfulness Elo and a +268.31 increase in harmlessness Elo compared to Alpaca-7B.
    • Human Evaluation Elo Scores:
      • Beaver-v3 achieved a +363.86 increase in helpfulness Elo and a +237.98 increase in harmlessness Elo compared to Alpaca-7B.
    • Proportion of Harmful Responses on Evaluation Prompts:
      • The probability of generating harmful responses as evaluated by human annotators decreased from 53.08% for Alpaca-7B to 2.45% for Beaver-v3.

    During round 3, because the model already satisfied safety constraints, the adaptive multiplier λ\lambda prevented excessive safety penalization, enabling continued gains in helpfulness without sacrificing harmlessness.

  6. Knowl 6 — Dynamic Lagrangian Balancing versus Static Reward Shaping

    empirical result

    In alignment experiments comparing Safe RLHF to static Reward Shaping (RS)—where the RL objective uses a fixed linear combination Rν(y,x)=Rϕ(y,x)−νCψ(y,x)R_\nu(y, x) = R_\phi(y, x) - \nu C_\psi(y, x) across fixed weighting coefficients ν∈{0.01,0.5,1,2,5,10,100}\nu \in \{0.01, 0.5, 1, 2, 5, 10, 100\}:

    • Extreme weights result in severe single-objective bias: low coefficients (ν=0.01,0.5\nu = 0.01, 0.5) improve helpfulness but fail to reduce harmful responses, while high coefficients (ν=5,10,100\nu = 5, 10, 100) over-optimize for safety at the severe expense of helpfulness win rates.
    • Moderate weights (ν=1,2\nu = 1, 2) remain strictly Pareto-inferior to Safe RLHF in both helpfulness and harmlessness win rates against the baseline Alpaca-7B model.

    Safe RLHF dynamically updates λ\lambda via the moving average cost JC(θ)J_C(\theta): when responses satisfy the cost constraint (JC(θ)≤0J_C(\theta) \le 0), λ\lambda decreases, allowing the policy to maximize helpfulness while preserving harmlessness.

  7. Knowl 7 — Preference and Cost Model Prediction Accuracies Across Safe RLHF Iterations

    data/table

    The test accuracy for the Reward Model and Cost Model across the three Safe RLHF training rounds (Beaver-v1, Beaver-v2, Beaver-v3) and a unified model trained on pooled, balanced preference data from all iterations is shown below:

    Model Metric Beaver-v1 Beaver-v2 Beaver-v3 Unified
    Reward Model Ranking Accuracy 78.13% 75.73% 77.32% 73.95%
    Cost Model Ranking Accuracy 74.47% 76.07% 74.17% 70.44%
    Cost Model Safety Classification Accuracy 95.62% 84.54% 85.88% 85.83%

    The reward model maintains 73.95%–78.13% ranking accuracy for helpfulness preferences. The cost model achieves 70.44%–76.07% ranking accuracy on relative harmfulness while simultaneously maintaining 84.54%–95.62% binary classification accuracy on absolute response safety.

  8. Knowl 8 — Effectiveness of Joint Ranking-Classification Cost Modeling vs. Pure Classifier

    empirical result

    In ablation studies comparing the Safe RLHF Cost Model (which simultaneously optimizes Bradley-Terry pairwise preference ranking and binary safety classification anchored at Cψ(y0,x)=0C_\psi(y_0, x)=0) against a baseline using the output logits of a standard binary safety classifier as the cost signal (CM-classifier):

    • Using raw classifier logits during RL training yields significantly inferior harmlessness win rates against the base SFT model compared to Safe RLHF.
    • Removing the classification terms from the cost loss and fixing the dual multiplier degrades the model into standard static reward shaping.

    Jointly fitting pairwise rankings and absolute safety labels allows the cost model to establish a calibrated decision boundary at zero while retaining continuous gradient direction for ranking less harmful alternatives.

  9. Knowl 9 — Taxonomy of Red-Teaming Jailbreak Vulnerabilities in Aligned LLMs

    definition

    Adversarial red-teaming of intermediate RLHF models identified four main categories of successful attacks that bypass safety mechanisms:

    1. Partial Harmfulness: The model provides harmful instructions or dangerous details while formulating a refusal, or provides dangerous content first before following up with an ethical disclaimer.
    2. Scenario Assumptions: The prompt forces the model to engage in role-playing, hypothetical storytelling, or specific situational simulations, exploiting the model's instruction-following nature to override safety guards.
    3. Contradictory Analysis: The prompt asks the model to justify, list advantages of, or analyze positive aspects of harmful, discriminatory, or unlawful activities.
    4. Complex Text Command Embedding: The prompt embeds harmful instructions within complex formatting constraints (e.g., demanding execution as a Python script or strict prefix completions) that distract or bypass safety filters.
  10. Knowl 10 — Inter-Rater and Quality Control Agreement Gains from Preference Decoupling

    empirical result

    Decoupling helpfulness and harmlessness into independent annotation dimensions provides measurable improvements in annotation consistency over single-dimensional overall preference collection:

    • Inter-Rater Agreement: Crowdworkers achieve 69.00% agreement on helpfulness and 66.53% agreement on safety under the decoupled annotation strategy, compared to 61.65% agreement under single-dimensional overall preference annotation.
    • Quality Control Approval Rate: In 10% spot-check quality control reviews comparing crowdworker annotations to expert researcher answer keys, the decoupled framework maintains an approval rate ≥90%\ge 90\%, whereas single-dimensional annotation approval rates fall below 80%.

Coverage note — None was omitted; the knowl set covers the CMDP optimization formulation, preference and cost modeling losses, policy update algorithm, annotation pipeline, iterative Beaver experimental results, comparison with reward shaping and classifier baselines, model accuracies, red-teaming taxonomy, and annotation agreement statistics.

References

  1. 1.Abubakar Abid, Maheen Farooqi, and James Zou. Persistent anti-muslim bias in large language models. In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society, pp. 298–306, 2021.
  2. 2.Eitan Altman. Constrained Markov decision processes. Routledge, 2021.
  3. 3.Rohan Anil, Andrew M Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al. Palm 2 technical report. arXiv preprint arXiv:2305.10403, 2023.
  4. 4.Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Ben Mann, Nova DasSarma, et al. A general language assistant as a laboratory for alignment. arXiv preprint arXiv:2112.00861, 2021.
  5. 5.Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862, 2022a.
  6. 6.Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al. Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073, 2022b.
  7. 7.Dimitri P Bertsekas. Nonlinear programming. Journal of the Operational Research Society, 48(3): 334–334, 1997.
  8. 8.Ralph Allan Bradley and Milton E Terry. Rank analysis of incomplete block designs: I. the method of paired comparisons. Biometrika, 39(3/4):324–345, 1952.
  9. 9.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020.
  10. 10.Stephen Casper, Xander Davies, Claudia Shi, Thomas Krendl Gilbert, Jérémy Scheurer, Javier Rando, Rachel Freedman, Tomasz Korbak, David Lindner, Pedro Freire, et al. Open problems and fundamental limitations of reinforcement learning from human feedback. arXiv preprint arXiv:2307.15217, 2023.
  11. 11.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374, 2021.
  12. 12.Cheng-Han Chiang and Hung-yi Lee. Can large language models be an alternative to human evaluations? arXiv preprint arXiv:2305.01937, 2023.
  13. 13.Yinlam Chow, Mohammad Ghavamzadeh, Lucas Janson, and Marco Pavone. Risk-constrained reinforcement learning with percentile risk criteria. The Journal of Machine Learning Research, 18 (1):6070–6120, 2017.
  14. 14.Jon Christian. Amazing “jailbreak” bypasses chatgpt’s ethics safeguards. Futurism, February, 4: 2023, 2023.
  15. 15.Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. Deep reinforcement learning from human preferences. Advances in neural information processing systems, 30, 2017.
  16. 16.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416, 2022.
  17. 17.Ameet Deshpande, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, and Karthik Narasimhan. Toxicity in chatgpt: Analyzing persona-assigned language models. arXiv preprint arXiv:2304.05335, 2023.
  18. 18.Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, et al. Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned. arXiv preprint arXiv:2209.07858, 2022.
  19. 19.Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig. Pal: Program-aided language models. In International Conference on Machine Learning, pp. 10764–10799. PMLR, 2023.
  20. 20.Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A Smith. Realtoxicityprompts: Evaluating neural toxic degeneration in language models. arXiv preprint arXiv:2009.11462, 2020.
  21. 21.Amelia Glaese, Nat McAleese, Maja Trebacz, John Aslanides, Vlad Firoiu, Timo Ewalds, Maribeth Rauh, Laura Weidinger, Martin Chadwick, Phoebe Thacker, et al. Improving alignment of dialogue agents via targeted human judgements. arXiv preprint arXiv:2209.14375, 2022.
  22. 22.Jiaming Ji, Mickel Liu, Juntao Dai, Xuehai Pan, Chi Zhang, Ce Bian, Ruiyang Sun, Yizhou Wang, and Yaodong Yang. Beavertails: Towards improved safety alignment of llm via a human-preference dataset. arXiv preprint arXiv:2307.04657, 2023.
  23. 23.Enkelejda Kasneci, Kathrin Seßler, Stefan Küchemann, Maria Bannert, Daryna Dementieva, Frank Fischer, Urs Gasser, Georg Groh, Stephan Günnemann, Eyke Hüllermeier, et al. Chatgpt for good? on opportunities and challenges of large language models for education. Learning and individual differences, 103:102274, 2023.
  24. 24.Daniel Martin Katz, Michael James Bommarito, Shang Gao, and Pablo Arredondo. Gpt-4 passes the bar exam. Available at SSRN 4389233, 2023.
  25. 25.Changyeon Kim, Jongjin Park, Jinwoo Shin, Honglak Lee, Pieter Abbeel, and Kimin Lee. Preference transformer: Modeling human preferences using transformers for rl. arXiv preprint arXiv:2303.00957, 2023.
  26. 26.Jan Kocoń, Igor Cichecki, Oliwier Kaszyca, Mateusz Kochanek, Dominika Szydło, Joanna Baran, Julita Bielaniewicz, Marcin Gruza, Arkadiusz Janz, Kamil Kanclerz, et al. Chatgpt: Jack of all trades, master of none. Information Fusion, pp. 101861, 2023.
  27. 27.Huan Yee Koh, Jiaxin Ju, Ming Liu, and Shirui Pan. An empirical survey on long document summarization: Datasets, models, and metrics. ACM computing surveys, 55(8):1–35, 2022.
  28. 28.Tiffany H Kung, Morgan Cheatham, Arielle Medenilla, Czarina Sillos, Lorie De Leon, Camille Elepaño, Maria Madriaga, Rimel Aggabao, Giezel Diaz-Candido, James Maningo, et al. Performance of chatgpt on usmle: Potential for ai-assisted medical education using large language models. PLoS digital health, 2(2):e0000198, 2023.
  29. 29.Michael Moor, Oishi Banerjee, Zahra Shakeri Hossein Abad, Harlan M Krumholz, Jure Leskovec, Eric J Topol, and Pranav Rajpurkar. Foundation models for generalist medical artificial intelligence. Nature, 616(7956):259–265, 2023.
  30. 30.Andrew Y Ng, Daishi Harada, and Stuart Russell. Policy invariance under reward transformations: Theory and application to reward shaping. In Icml, volume 99, pp. 278–287. Citeseer, 1999.
  31. 31.OpenAI. Gpt-4 technical report, 2023.
  32. 32.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35: 27730–27744, 2022.
  33. 33.Martin L Puterman. Markov decision processes: discrete stochastic dynamic programming. John Wiley & Sons, 2014.
  34. 34.Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D Manning, and Chelsea Finn. Direct preference optimization: Your language model is secretly a reward model. arXiv preprint arXiv:2305.18290, 2023.
  35. 35.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research, 21(1):5485–5551, 2020.
  36. 36.Maribeth Rauh, John Mellor, Jonathan Uesato, Po-Sen Huang, Johannes Welbl, Laura Weidinger, Sumanth Dathathri, Amelia Glaese, Geoffrey Irving, Iason Gabriel, et al. Characteristics of harmful text: Towards rigorous benchmarking of language models. Advances in Neural Information Processing Systems, 35:24720–24739, 2022.
  37. 37.Dorsa Sadigh, Anca D Dragan, Shankar Sastry, and Sanjit A Seshia. Active preference-based learning of reward functions. 2017.
  38. 38.Victor Sanh, Albert Webson, Colin Raffel, Stephen H Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, et al. Multitask prompted training enables zero-shot task generalization. arXiv preprint arXiv:2110.08207, 2021.
  39. 39.Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al. Bloom: A 176bparameter open-access multilingual language model. arXiv preprint arXiv:2211.05100, 2022.
  40. 40.John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017.
  41. 41.John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel. Highdimensional continuous control using generalized advantage estimation. arXiv preprint arXiv:1506.02438, 2018.
  42. 42.Emily Sheng, Kai-Wei Chang, Premkumar Natarajan, and Nanyun Peng. Societal biases in language generation: Progress and challenges. arXiv preprint arXiv:2105.04054, 2021.
  43. 43.Toby Shevlane, Sebastian Farquhar, Ben Garfinkel, Mary Phuong, Jess Whittlestone, Jade Leung, Daniel Kokotajlo, Nahema Marchal, Markus Anderljung, Noam Kolt, et al. Model evaluation for extreme risks. arXiv preprint arXiv:2305.15324, 2023.
  44. 44.Feifan Song, Bowen Yu, Minghao Li, Haiyang Yu, Fei Huang, Yongbin Li, and Houfeng Wang. Preference ranking optimization for human alignment. arXiv preprint arXiv:2306.17492, 2023.
  45. 45.Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano. Learning to summarize with human feedback. Advances in Neural Information Processing Systems, 33:3008–3021, 2020.
  46. 46.Hao Sun, Zhexin Zhang, Jiawen Deng, Jiale Cheng, and Minlie Huang. Safety assessment of chinese large language models, 2023a.
  47. 47.Zhiqing Sun, Yikang Shen, Qinhong Zhou, Hongxin Zhang, Zhenfang Chen, David Cox, Yiming Yang, and Chuang Gan. Principle-driven self-alignment of language models from scratch with minimal human supervision. arXiv preprint arXiv:2305.03047, 2023b.
  48. 48.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto. Stanford alpaca: An instruction-following llama model, 2023.
  49. 49.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023a.
  50. 50.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023b.
  51. 51.Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, et al. Ethical and social risks of harm from language models. arXiv preprint arXiv:2112.04359, 2021.
  52. 52.Laura Weidinger, Jonathan Uesato, Maribeth Rauh, Conor Griffin, Po-Sen Huang, John Mellor, Amelia Glaese, Myra Cheng, Borja Balle, Atoosa Kasirzadeh, et al. Taxonomy of risks posed by language models. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, pp. 214–229, 2022.
  53. 53.Zeqiu Wu, Yushi Hu, Weijia Shi, Nouha Dziri, Alane Suhr, Prithviraj Ammanabrolu, Noah A Smith, Mari Ostendorf, and Hannaneh Hajishirzi. Fine-grained human feedback gives better rewards for language model training. arXiv preprint arXiv:2306.01693, 2023.
  54. 54.Kevin Yang, Dan Klein, Asli Celikyilmaz, Nanyun Peng, and Yuandong Tian. Rlcd: Reinforcement learning from contrast distillation for language model alignment. arXiv preprint arXiv:2307.12950, 2023.
  55. 55.Xi Yang, Aokun Chen, Nima PourNejatian, Hoo Chang Shin, Kaleb E Smith, Christopher Parisien, Colin Compas, Cheryl Martin, Anthony B Costa, Mona G Flores, et al. A large language model for electronic health records. NPJ Digital Medicine, 5(1):194, 2022.
  56. 56.Zheng Yuan, Hongyi Yuan, Chuanqi Tan, Wei Wang, Songfang Huang, and Fei Huang. Rrhf: Rank responses to align language models with human feedback without tears. arXiv preprint arXiv:2304.05302, 2023.

Citation

MLA
Dai, J., et al. “Safe RLHF: Safe Reinforcement Learning from Human Feedback”. arXiv, 2023, http://arxiv.org/abs/2310.12773v1.
APA
Dai, J., Pan, X., Sun, R., Ji, J., Xu, X., Liu, M., Wang, Y., & Yang, Y. (2023). Safe RLHF: Safe Reinforcement Learning from Human Feedback. arXiv. http://arxiv.org/abs/2310.12773v1
Chicago
Dai, J., X. Pan, R. Sun, et al. 2023. “Safe RLHF: Safe Reinforcement Learning from Human Feedback”. arXiv. http://arxiv.org/abs/2310.12773v1.
Harvard
Dai, J. et al. (2023) “Safe RLHF: Safe Reinforcement Learning from Human Feedback”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2310.12773v1.
Vancouver
1. Dai J, Pan X, Sun R, Ji J, Xu X, Liu M, Wang Y, Yang Y (2023) Safe RLHF: Safe Reinforcement Learning from Human Feedback. arXiv

BibTeX

@article{dai2023safe,
  title = {Safe RLHF: Safe Reinforcement Learning from Human Feedback},
  author = {Dai, Josef and Pan, Xuehai and Sun, Ruiyang and Ji, Jiaming and Xu, Xinbo and Liu, Mickel and Wang, Yizhou and Yang, Yaodong},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2310.12773v1},
  eprint = {2310.12773}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors