Mitigating the Alignment Tax of RLHF

Yong LinHangyu LinWei XiongShizhe DiaoJianmeng LiuJipeng ZhangRui PanHaoxiang WangWenbin HuHanning Zhang

article2024EMNLP226 citations

Proposes Heterogeneous Model Averaging, a layer-adaptive weight interpolation technique between pre- and post-RLHF models that optimizes alignment rewards while preserving general NLP capabilities across diverse model scales.

Listen

Large language models acquire broad general capabilities during pre-training, such as reading comprehension, common sense reasoning, and translation. However, post-training alignment techniques like Reinforcement Learning with Human Feedback (RLHF), designed to make models helpful, honest, and harmless, frequently cause catastrophic forgetting of these core abilities—a phenomenon known as the "alignment tax." As organizations increasingly rely on aligned models for production tasks, mitigating this performance degradation without sacrificing safety and human preference alignment has become a critical operational challenge.

The article evaluates methods to resolve this alignment-forgetting trade-off and demonstrates that model weight averaging provides an exceptionally effective and computationally practical solution. It introduces and evaluates Heterogeneous Model Averaging (HMA), an approach that dynamically assigns distinct interpolation weights to different model layers to maximize alignment rewards while preserving pre-trained skills.

To conduct this evaluation, the researchers tested alignment algorithms—including Rejection Sampling Fine-Tuning (RSF), Direct Preference Optimization (DPO), and Proximal Policy Optimization (PPO)—on the OpenLLaMA-3B foundation model and extended their validation to larger architectures like Mistral-7B and Gemma-7B (specifically Zephyr-7B variants). They measured alignment tax across multiple standard natural language benchmarks (such as ARC, SQuAD, DROP, and WMT translation) and assessed alignment quality using both specialized reward models and GPT-4 evaluations. The analysis systematically compared simple model averaging and HMA against established alternatives, including parameter regularization (L1/L2 penalties), knowledge distillation, low-rank adaptation (LoRA), reward penalties, and experience replay using subsets of pre-training data.

The study yielded several key findings. First, simple model weight averaging between pre-RLHF and post-RLHF checkpoints consistently established a superior trade-off boundary compared to complex regularization, distillation, LoRA, and even experience replay methods. Second, replaying pre-training data failed to match model averaging on two of three core benchmark suites, despite adding four times the data volume of the RLHF dataset (400 million tokens) and incurring heavy computational overhead. Third, the researchers established both theoretically and empirically that averaging lower-level transformer layers yields the greatest joint improvements in alignment and general task performance, because these layers share broad, foundational feature spaces across tasks. Fourth, the proposed Heterogeneous Model Averaging framework pushed the performance boundary further across all algorithms: setting a moderate average ratio (around 0.2) preserved baseline capabilities while outperforming baseline models, achieving higher human-preference win rates on AlpacaEval (e.g., 9.32% vs. 8.10% against GPT-4 on Zephyr-7B-β) and improved scores across reading comprehension, common sense, and translation tasks.

These findings indicate that teams deploying aligned language models do not need to choose between severe capability regression and expensive, complex mitigation pipelines. Model averaging and HMA operate strictly as post-processing steps on existing checkpoints, eliminating the extreme computational costs, training instabilities, and proprietary data access hurdles associated with data replay or constrained optimization. This directly lowers infrastructure expenses, reduces project timelines, and enhances model safety and performance simultaneously.

Organizations training or fine-tuning language models should adopt model averaging techniques as a standard post-alignment pipeline stage. Practitioners should prioritize setting an overall averaging ratio around 0.2, retaining heavier weight from the pre-RLHF model on lower layers while allowing higher layers to retain more alignment-specific adjustments. When feasible, HMA should be applied via reward proxy distillation on a small sample of generated outputs rather than relying on heavy multi-task tuning.

While the article demonstrates high confidence and consistent empirical validation across multiple model sizes and alignment frameworks, it notes that HMA substantially mitigates but does not entirely eliminate the alignment tax. Decision-makers should recognize that optimal averaging ratios may still require minor empirical verification across distinct domain distributions, and future work is required to establish the theoretical lower limit of capability degradation during alignment.

Cover for Mitigating the Alignment Tax of RLHF

Abstract

LLMs acquire a wide range of abilities during pre-training, but aligning LLMs under Reinforcement Learning with Human Feedback (RLHF) can lead to forgetting pretrained abilities, which is also known as the alignment tax. To investigate alignment tax, we conducted experiments with existing RLHF algorithms using OpenLLaMA-3B, which revealed a pronounced alignment tax in NLP tasks. Whereas, despite various techniques to mitigate forgetting, they are often at odds with the RLHF performance, leading to a trade-off between alignment performance and forgetting mitigation, leading to an alignment-forgetting trade-off.

In this paper we show that model averaging, which simply interpolates between pre and post RLHF model weights, surprisingly achieves the most strongest alignment-forgetting Pareto front among a wide range of competing methods. To understand its effectiveness, we offer theoretical insights into model averaging, revealing that it enhances performance Pareto front by increasing feature diversity on the layers where tasks share overlapped feature spaces. Empirical evidence corroborates our analysis by showing the benefits of averaging low-level transformer layers. Building on the analysis and the observation that averaging different layers of the transformer leads to significantly different alignment-forgetting trade-offs, we propose Heterogeneous Model Averaging (HMA) to Heterogeneously find various combination ratios of model layers. HMA seeks to maximize the alignment performance while incurring minimal alignment tax. Moreover, we validate HMA’s performance across a range of RLHF algorithms over OpenLLaMA-3B and further extend our findings to Mistral-7B which is evaluated by open-sourced preference model and GPT4. Code available here¹.

Table of Contents

  • 1 Introduction
  • 2 Discussion with existing works.
  • 3 Experimental Settings
  • 4 Evaluating Existing Methods
  • 4.1 Basic Methods
  • 5 Unravelling the Mysteries of Model Averaging for Alleviating Alignment Tax
  • 6 Heterogeneous Model Averaging
  • 7 Conclusion
  • Limitations
  • References
  • A Related Work
  • B RLHF Basics
  • B.1 Algorithm of Heterogeneous Model Averaging
  • C More Results
  • C.1 Experience Replay
  • C.2 Reward Penalty
  • C.3 Consistency of different combination ratios among various tasks
  • C.4 Results of α = 0.2
  • D Implementation Details
  • D.1 Rejection Sampling Fine-tuning Implementation
  • D.2 Implementation of PPO
  • D.3 Implementation of DPO
  • D.4 Implementations of Existing Methods to Alleviate Alignment Tax
  • D.5 Implementations of Heterogeneous Model Averaging
  • E More Results
  • E.1 The Alignment Tax during Training (Results of Early Stopping)
  • E.2 More Results of Averaging Different Parts
  • E.3 Comparison of RLHF Algorithms
  • E.4 Results of AdaMerging (Yang et al., 2023)
  • E.5 Detailed Results of Heterogeneous Model Averaging
  • F Theoretical Settings, Proofs and Discussions
  • F.1 Re-statement of Formal Settings
  • F.2 Proof of Proposition 5.1
  • F.3 Discussion on the Effect of Task Similarity on Model Averaging
  • F.4 Close Form of F_p(x)
  • G Hyper-Parameters

Knowls

  1. Knowl 1 — Heterogeneous Model Averaging Framework for Mitigating Alignment Tax

    model/method

    Heterogeneous Model Averaging (HMA) is a post-RLHF parameter merging strategy designed to mitigate the forgetting of pre-trained capabilities (the "alignment tax") while maximizing human preference alignment. Given an initial instruction-tuned checkpoint θ0\theta_0 and a post-RLHF aligned policy checkpoint θ\theta, the LL-layer transformer architecture is partitioned into KK sequential block segments (for example, K=3K=3 corresponding to lower/input, middle, and upper/output layers). Each kk-th block segment θ[k]\theta^{[k]} is assigned a distinct averaging coefficient αk∈[0,1]\alpha_k \in [0, 1], defining the merged parameters θ(K)\theta(K) component-wise as:

    θ[k](K):=αkθ[k]+(1−αk)θ0[k],∀k∈{1,…,K}\theta^{[k]}(K) := \alpha_k \theta^{[k]} + (1 - \alpha_k) \theta_0^{[k]}, \quad \forall k \in \{1, \dots, K\}

    To ensure that the overall retention of pre-RLHF general capability remains controlled, HMA constrains the arithmetic mean of the layer coefficients to a pre-selected target scalar merge ratio α∈[0,1]\alpha \in [0, 1]:

    Ω:={(α1,…,αK)∈[0,1]K  |  1K∑k=1Kαk=α}\Omega := \left\{ (\alpha_1, \dots, \alpha_K) \in [0, 1]^K \;\middle|\; \frac{1}{K} \sum_{k=1}^K \alpha_k = \alpha \right\}

    HMA solves for the optimal combination of block coefficients by maximizing the expected reward r∗(x,a)r^*(x, a) under the policy πθ(K)\pi_{\theta(K)} induced by the merged parameters:

    max⁡(α1,…,αK)∈ΩExEa∼πθ(K)(⋅∣x)[r∗(x,a)]\max_{(\alpha_1, \dots, \alpha_K) \in \Omega} \mathbb{E}_{x} \mathbb{E}_{a \sim \pi_{\theta(K)}(\cdot \mid x)} \left[ r^*(x, a) \right]

  2. Knowl 2 — HMA Optimization via Reparameterized Proxy Distillation

    algorithm

    To optimize the block combination ratios (α1,…,αK)(\alpha_1, \dots, \alpha_K) without executing expensive reinforcement learning rollouts, HMA uses a proxy distillation dataset and a reparameterization trick that enforces the mean constraint 1K∑k=1Kαk=α\frac{1}{K}\sum_{k=1}^K \alpha_k = \alpha.

    First, using the post-RLHF aligned policy πθ\pi_\theta, responses are sampled for prompts x∈Xx \in \mathcal{X} to construct a high-reward proxy distillation dataset Dθ={(x,a)∣a∼πθ(⋅∣x)}\mathcal{D}_\theta = \{(x, a) \mid a \sim \pi_\theta(\cdot \mid x)\} containing the top reward-scoring prompt-response pairs. The objective function is transformed into maximizing the log-likelihood of Dθ\mathcal{D}_\theta:

    max⁡α1,…,αK∈Ω1∣Dθ∣∑(x,a)∈Dθlog⁡πθ(K)(a∣x)\max_{\alpha_1, \dots, \alpha_K \in \Omega} \frac{1}{|\mathcal{D}_\theta|} \sum_{(x, a) \in \mathcal{D}_\theta} \log \pi_{\theta(K)}(a \mid x)

    To enforce (α1,…,αK)∈Ω(\alpha_1, \dots, \alpha_K) \in \Omega, unconstrained continuous variables s1,…,sK∈Rs_1, \dots, s_K \in \mathbb{R} are optimized via:

    α^i=σ(si)+ϵ,αi=α^i∑j=1Kα^j⋅(Kα)\hat{\alpha}_i = \sigma(s_i) + \epsilon, \quad \alpha_i = \frac{\hat{\alpha}_i}{\sum_{j=1}^K \hat{\alpha}_j} \cdot (K \alpha)

    where σ(z)=(1+exp⁡(−z))−1\sigma(z) = (1 + \exp(-z))^{-1} is the sigmoid function and ϵ≥0\epsilon \ge 0 is a boundary control parameter.

    Input: Initial instruction policy πθ0\pi_{\theta_0}, post-RLHF aligned policy πθ\pi_\theta, prompt set Dx\mathcal{D}_x, partition count KK, target merge ratio α\alpha, boundary parameter ϵ\epsilon.
    Output: Optimized merged policy πθ(K)\pi_{\theta(K)}.
    Sample responses a∼πθ(⋅∣x)a \sim \pi_\theta(\cdot \mid x) for x∈Dxx \in \mathcal{D}_x and select high-reward pairs to build proxy dataset Dθ\mathcal{D}_\theta
    Initialize unconstrained parameters s1,…,sK∈Rs_1, \dots, s_K \in \mathbb{R}
    for step = 1 to max_steps do
        Compute normalized layer coefficients αi=σ(si)+ϵ∑j=1K(σ(sj)+ϵ)⋅Kα\alpha_i = \frac{\sigma(s_i) + \epsilon}{\sum_{j=1}^K (\sigma(s_j) + \epsilon)} \cdot K\alpha for all i∈{1,…,K}i \in \{1, \dots, K\}
        Form merged model weights θ[i](K)=αiθ[i]+(1−αi)θ0[i]\theta^{[i]}(K) = \alpha_i \theta^{[i]} + (1 - \alpha_i) \theta_0^{[i]} for each block ii
        Compute proxy distillation loss L=−1∣Dθ∣∑(x,a)∈Dθlog⁡πθ(K)(a∣x)\mathcal{L} = -\frac{1}{|\mathcal{D}_\theta|} \sum_{(x, a) \in \mathcal{D}_\theta} \log \pi_{\theta(K)}(a \mid x)
        Update s1,…,sKs_1, \dots, s_K via gradient descent on L\mathcal{L}
    end for
    Construct final merged model θ(K)\theta(K) using optimized ratios (α1,…,αK)(\alpha_1, \dots, \alpha_K)
    return πθ(K)\pi_{\theta(K)}
  3. Knowl 3 — Theoretical Gain of Model Averaging in Shared Feature Spaces

    theoretical result

    Let two tasks TaT_a and TbT_b be defined on an input feature space Sx={xi}i=1DS_x = \{x_i\}_{i=1}^D where each feature fails with independent probability pp. Let model fa(x)=waΦa(x)f_a(x) = w_a \Phi_a(x) select feature set Sx,a⊂SxS_{x,a} \subset S_x and model fb(x)=wbΦb(x)f_b(x) = w_b \Phi_b(x) select feature set Sx,b⊂SxS_{x,b} \subset S_x, where ∣Sx,a∣=∣Sx,b∣=n|S_{x,a}| = |S_{x,b}| = n, and the number of overlapping features is no=∣Sx,a∩Sx,b∣n_o = |S_{x,a} \cap S_{x,b}|. The averaged model is defined as favg(x)=14(wa+wb)⊤x(Φa+Φb)f_{\text{avg}}(x) = \frac{1}{4}(w_a + w_b)^\top x (\Phi_a + \Phi_b).

    The effective averaging robustness gain across the two tasks is defined by:

    ξ=12(Aa(favg)−Aa(fa)+Ab(favg)−Ab(fb))\xi = \frac{1}{2} \left( A_a(f_{\text{avg}}) - A_a(f_a) + A_b(f_{\text{avg}}) - A_b(f_b) \right)

    where At(f)A_t(f) denotes the classification accuracy of model ff on task t∈{a,b}t \in \{a, b\}.

    Under Case (1), where tasks share identical label spaces and project onto a shared feature space, the robustness gain is denoted ξ(1)\xi^{(1)}. Under Case (2), where tasks have disjoint label and feature spaces (∣Ya∩Yb∣=0|\mathcal{Y}_a \cap \mathcal{Y}_b| = 0), the robustness gain is ξ(2)=0\xi^{(2)} = 0.

    Assuming unit-norm orthogonal features (∥μi(k)∥2=1\Vert\mu_i(k)\Vert_2 = 1 and μi(k)⊥μi′(k′)\mu_i(k) \perp \mu_{i'}(k') for (i,k)≠(i′,k′)(i, k) \neq (i', k')) and bounded Gaussian feature noise, Proposition 5.1 establishes:

    ξ(1)−ξ(2)=Fp(2(1−p)nn+no)−Fp((1−p)n)≥0\xi^{(1)} - \xi^{(2)} = F_p\left( \frac{\sqrt{2}(1 - p)n}{\sqrt{n + n_o}} \right) - F_p\left((1 - p)\sqrt{n}\right) \ge 0

    where Fp(x)=P(η1>0,…,ηK−1>0)F_p(x) = \mathbb{P}(\eta_1 > 0, \dots, \eta_{K-1} > 0) for η∼N(x,M)\eta \sim \mathcal{N}(x, M) is a strictly monotonically increasing cumulative distribution function with covariance matrix Mi,i=p(K+2−pK)KM_{i,i} = \frac{p(K+2-pK)}{K} and Mi,j=p(K+1−pK)KM_{i,j} = \frac{p(K+1-pK)}{K} for i≠ji \neq j. The equality ξ(1)−ξ(2)=0\xi^{(1)} - \xi^{(2)} = 0 holds if and only if no=nn_o = n.

    This proves that weight averaging provides strict performance improvements when tasks share lower-level representation spaces with partially non-overlapping feature selections (no<nn_o < n), as weight averaging increases the diversity of utilized features and decreases the probability of joint feature failure.

  4. Knowl 4 — Pareto Superiority of Linear Model Averaging over Existing Anti-Forgetting Techniques

    empirical result

    Simple linear model averaging (MA) between the pre-RLHF instruction-tuned checkpoint θ0\theta_0 and the post-RLHF policy checkpoint θ\theta, defined by θavg=(1−α)θ0+αθ\theta_{\text{avg}} = (1 - \alpha)\theta_0 + \alpha \theta with α∈[0,1]\alpha \in [0, 1], achieves a strictly superior alignment-forgetting Pareto front compared to existing continual learning, regularization, and parameter-efficient methods.

    Evaluated on OpenLLaMA-3B aligned with Rejection Sampling Finetuning (RSF) on the Anthropic Helpful and Harmless (HH) RLHF dataset, MA was tested against:

    1. Early stopping checkpoints during RLHF (iterations 2, 4, 6, 8).
    2. Weight-space parameter penalties during alignment: L1L_1 penalty λ∥θ−θ0∥1\lambda \Vert\theta - \theta_0\Vert_1 with λ∈{0.04,0.1,0.4,0.6,1.0}\lambda \in \{0.04, 0.1, 0.4, 0.6, 1.0\}, and L2L_2 penalty λ∥θ−θ0∥22\lambda \Vert\theta - \theta_0\Vert_2^2 with λ∈{0.01,0.04,0.06,0.08,0.1}\lambda \in \{0.01, 0.04, 0.06, 0.08, 0.1\}.
    3. Knowledge Distillation (KD) penalty λ∥πθ(x)−πθ0(x)∥22\lambda \Vert\pi_\theta(x) - \pi_{\theta_0}(x)\Vert_2^2 with λ∈{10−5,10−3,10−1}\lambda \in \{10^{-5}, 10^{-3}, 10^{-1}\}.
    4. Low-Rank Adaptation (LoRA) on MLP blocks (ranks 16 to 512) and attention blocks (rank 16).
    5. Stochastic Moving Averaging (SMA).

    Across Reading Comprehension (SQuAD and DROP F1 scores), Commonsense QA (ARC-Easy, ARC-Challenge, Race, PIQA accuracy), and French-to-English translation (WMT 2014 BLEU), linear model averaging consistently forms an outer Pareto envelope over all regularized and parameter-efficient fine-tuning baselines.

  5. Knowl 5 — Layer-Asymmetric Sensitivity in Model Averaging

    empirical result

    Averaging different architectural blocks of the transformer during post-RLHF model averaging yields starkly different alignment-forgetting trade-offs:

    1. Input/Low-Level Layers (Layers 1–8 of OpenLLaMA-3B): Interpolating only the input block θ[1]\theta^{[1]} towards θ0[1]\theta_0^{[1]} while keeping middle and output layers fixed at θ\theta produces a simultaneous increase in both alignment reward (HH RLHF reward) and general NLP metrics (Reading Comprehension F1, Commonsense QA accuracy, Translation BLEU).
    2. Middle Layers (Layers 9–17) and Output Layers (Layers 18–26): Averaging middle or output layers towards θ0\theta_0 results in a steep decline in alignment reward without a commensurate gain in pre-trained NLP capability retention.

    This asymmetric pattern is empirically consistent across Rejection Sampling Finetuning (RSF), Direct Preference Optimization (DPO), and Proximal Policy Optimization (PPO), confirming that lower transformer layers capture foundational, shared feature representations beneficial to both pre-training and alignment tasks, whereas output layers contain task-specific alignment parameters.

  6. Knowl 6 — GPT-4 and NLP Benchmark Evaluation of HMA on 7B Scale Models

    data/table

    Applying Heterogeneous Model Averaging (HMA) with target merge ratio α=0.2\alpha = 0.2 to 7B language models (Zephyr-7B-β\beta aligned via DPO from Mistral-7B-SFT-β\beta, and Zephyr-7B-Gemma aligned from Gemma-7B) achieves higher GPT-4 win rates on AlpacaEval 2.0 while simultaneously reversing the alignment tax across reading comprehension, commonsense reasoning, and translation benchmarks.

    Model Win-Rate Reading (F1) CommonSense (Acc %) Trans (BLEU)
    Zephyr-7B-β\beta 8.10% 37.47 66.34 36.55
    HMA (Ours) 9.32% 38.93 66.55 37.23
    Zephyr-7B-Gemma 11.3% 41.15 66.30 38.09
    HMA (Ours) 11.5% 42.45 66.40 38.71

    In the evaluation table:

    • Win-Rate is measured on AlpacaEval 2.0 against GPT-4.
    • Reading is evaluated by macro F1 score across SQuAD and DROP.
    • CommonSense is evaluated by accuracy across ARC-Easy, ARC-Challenge, Race, and PIQA.
    • Trans is French-to-English translation evaluated by BLEU on WMT 2014.

    On both Mistral-7B and Gemma-7B architectures, HMA improves instruction-following win rates against GPT-4 while raising Reading Comprehension F1 by +1.30 to +1.46 points, Translation BLEU by +0.62 to +0.68 points, and Commonsense accuracy by +0.10 to +0.21%.

  7. Knowl 7 — Empirical Superiority of Model Averaging over Experience Replay, KL Penalties, and AdaMerging

    empirical result

    Comparison against traditional continual learning and multi-task merging paradigms demonstrates the advantages of post-hoc model averaging for preserving pre-trained abilities:

    1. Experience Replay (ER): Mixing pre-training data during RLHF with replay ratios from 0.25×0.25\times to 4×4\times the alignment data volume (up to 800M tokens sampled from the 1.2T token OpenLLaMA pre-training dataset) only outperforms model averaging on Reading Comprehension, while failing on Commonsense QA and Translation. Because the pre-training dataset is massive, a 4×4\times replay buffer covers only ≈0.03%\approx 0.03\% of the pre-training distribution, leaving diverse abilities underrepresented.
    2. KL Reward Penalties in PPO: Incorporating KL penalties −ηKL(πθ∥πθ0)-\eta \text{KL}(\pi_\theta \parallel \pi_{\theta_0}) with η∈{0.05,0.1,0.2}\eta \in \{0.05, 0.1, 0.2\} into PPO objective functions partially mitigates regression but produces an alignment-forgetting trade-off curve strictly inferior to uniform model averaging.
    3. AdaMerging: Optimizing layer-wise merging coefficients using labeled training data from a single downstream task (e.g., Reading Comprehension) fails to preserve performance on other disjoint tasks (e.g., Commonsense QA). In contrast, HMA optimizes block coefficients solely on alignment data without requiring task-specific downstream datasets.
  8. Knowl 8 — Effect of Partition Granularity K on Overfitting in HMA

    empirical result

    In Heterogeneous Model Averaging (HMA), the number of partitioned transformer blocks KK controls the trade-off between model expressiveness and in-domain overfitting during ratio optimization:

    1. Coarse Partition (K=3K=3): Dividing the transformer into input, middle, and output thirds yields the most optimal alignment-forgetting Pareto frontier, consistently improving over vanilla uniform model averaging without degrading general task retention.
    2. Finer Partitions (K=6,9K=6, 9): While increasing KK to 6 or 9 yields slightly higher training alignment reward on the proxy distillation set Dθ\mathcal{D}_\theta, downstream task performance (such as Reading Comprehension F1) systematically drops.

    Optimizing a larger parameter vector (s1,…,sK)(s_1, \dots, s_K) causes the layer ratios to overfit the in-domain preference data, thereby amplifying catastrophic forgetting of pre-trained NLP skills.

  9. Knowl 9 — Optimal Universal Averaging Ratio Choice

    empirical result

    Empirical evaluations across diverse model architectures (OpenLLaMA-3B, Zephyr-7B-β\beta, Zephyr-7B-Gemma), alignment algorithms (RSF, DPO, PPO), and multiple downstream NLP benchmarks identify α=0.2\alpha = 0.2 as an optimal, robust default averaging ratio.

    Setting α=0.2\alpha = 0.2 in uniform Model Averaging and HMA consistently recovers substantial pre-trained NLP abilities (Reading Comprehension F1, Commonsense accuracy, and Translation BLEU) without degrading human preference rewards or GPT-4 win rates relative to the fully aligned checkpoint (α=1.0\alpha = 1.0).

  10. Knowl 10 — Limitations of Heterogeneous Model Averaging

    limitation

    Although Heterogeneous Model Averaging (HMA) significantly alleviates the alignment tax across various benchmarks, it exhibits specific limitations:

    1. Incomplete Elimination of Alignment Tax: HMA expands the alignment-forgetting Pareto frontier but does not entirely eliminate performance drops on all downstream NLP tasks relative to the unaligned instruction-tuned base model θ0\theta_0 when targeting high preference rewards.
    2. Absence of Theoretical Lower Bounds: The theoretical lower bound of the alignment tax trade-off under preference optimization remains uncharacterized, leaving open the question of whether an optimal zero-tax alignment exists.
    3. Proxy Distillation Dependency: The ratio optimization relies on empirical proxy distillation datasets Dθ\mathcal{D}_\theta, which introduces potential overfitting risks when scaling the partition granularity KK beyond coarse block segmentations.

Coverage note — None was omitted. All key theoretical formulations (Proposition 5.1, feature diversity framework), algorithms (HMA, proxy distillation reparameterization), empirical comparisons (MA vs LoRA, KD, L1/L2, ER, KL penalty, AdaMerging), layer-wise ablation studies, and multi-model evaluations (OpenLLaMA-3B, Zephyr-7B, Zephyr-7B-Gemma on GPT-4/PairRM) are fully covered.

References

  1. 1.Takuya Akiba, Makoto Shing, Yujin Tang, Qi Sun, and David Ha. 2024. Evolutionary optimization of model merging recipes. arXiv preprint arXiv:2403.13187.
  2. 2.Rahaf Aljundi, Francesca Babiloni, Mohamed Elhoseiny, Marcus Rohrbach, and Tinne Tuytelaars. 2018. Memory aware synapses: Learning what (not) to forget. In Proceedings of the European conference on computer vision (ECCV), pages 139–154.
  3. 3.Zeyuan Allen-Zhu and Yuanzhi Li. 2020. Towards understanding ensemble, knowledge distillation and self-distillation in deep learning. arXiv preprint arXiv:2012.09816.
  4. 4.Anders Andreassen, Yasaman Bahri, Behnam Neyshabur, and Rebecca Roelofs. 2021. The evolution of out-of-distribution robustness throughout fine-tuning. arXiv preprint arXiv:2106.15831.
  5. 5.Anthropic. 2023. Introducing claude.
  6. 6.Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Ben Mann, Nova DasSarma, et al. 2021. A general language assistant as a laboratory for alignment. arXiv preprint arXiv:2112.00861.
  7. 7.Mohammad Gheshlaghi Azar, Mark Rowland, Bilal Piot, Daniel Guo, Daniele Calandriello, Michal Valko, and Rémi Munos. 2023. A general theoretical paradigm to understand learning from human preferences. arXiv preprint arXiv:2310.12036.
  8. 8.Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al. 2022. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862.
  9. 9.Yonatan Bisk, Rowan Zellers, Jianfeng Gao, Yejin Choi, et al. 2020. Piqa: Reasoning about physical commonsense in natural language. In Proceedings of the AAAI conference on artificial intelligence, volume 34, pages 7432–7439.
  10. 10.Ondřej Bojar, Christian Buck, Christian Federmann, Barry Haddow, Philipp Koehn, Johannes Leveling, Christof Monz, Pavel Pecina, Matt Post, Herve Saint-Amand, et al. 2014. Findings of the 2014 workshop on statistical machine translation. In Proceedings of the ninth workshop on statistical machine translation, pages 12–58.
  11. 11.Ralph Allan Bradley and Milton E Terry. 1952. Rank analysis of incomplete block designs: I. the method of paired comparisons. Biometrika, 39(3/4):324–345.
  12. 12.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  13. 13.Pietro Buzzega, Matteo Boschini, Angelo Porrello, Davide Abati, and Simone Calderara. 2020. Dark experience for general continual learning: a strong, simple baseline. Advances in neural information processing systems, 33:15920–15930.
  14. 14.Lucas Caccia, Rahaf Aljundi, Nader Asadi, Tinne Tuytelaars, Joelle Pineau, and Eugene Belilovsky. 2021. New insights on reducing abrupt representation change in online continual learning. arXiv preprint arXiv:2104.05025.
  15. 15.Lucas Caccia, Eugene Belilovsky, Massimo Caccia, and Joelle Pineau. 2020. Online learned continual compression with adaptive quantization modules. In International Conference on Machine Learning, pages 1240–1250. PMLR.
  16. 16.Stephen Casper, Xander Davies, Claudia Shi, Thomas Krendl Gilbert, Jérémy Scheurer, Javier Rando, Rachel Freedman, Tomasz Korbak, David Lindner, Pedro Freire, et al. 2023. Open problems and fundamental limitations of reinforcement learning from human feedback. arXiv preprint arXiv:2307.15217.
  17. 17.Hyuntak Cha, Jaeho Lee, and Jinwoo Shin. 2021a. Co2l: Contrastive continual learning. In Proceedings of the IEEE/CVF International conference on computer vision, pages 9516–9525.
  18. 18.Junbum Cha, Sanghyuk Chun, Kyungjae Lee, Han-Cheol Cho, Seunghyun Park, Yunsung Lee, and Sungrae Park. 2021b. Swad: Domain generalization by seeking flat minima. Advances in Neural Information Processing Systems, 34:22405–22418.
  19. 19.Arslan Chaudhry, Marc’Aurelio Ranzato, Marcus Rohrbach, and Mohamed Elhoseiny. 2018. Efficient lifelong learning with a-gem. arXiv preprint arXiv:1812.00420.
  20. 20.Wuyang Chen, Yanqi Zhou, Nan Du, Yanping Huang, James Laudon, Zhifeng Chen, and Claire Cui. 2023. Lifelong language pretraining with distribution-specialized experts. In International Conference on Machine Learning, pages 5383–5395. PMLR.
  21. 21.Leshem Choshen, Lior Fox, Zohar Aizenbud, and Omri Abend. 2019. On the weaknesses of reinforcement learning for neural machine translation. arXiv preprint arXiv:1907.01752.
  22. 22.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2023. Palm: Scaling language modeling with pathways. Journal of Machine Learning Research, 24(240):1–113.
  23. 23.Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017. Deep reinforcement learning from human preferences. Advances in neural information processing systems, 30.
  24. 24.Xu Chu, Yujie Jin, Wenwu Zhu, Yasha Wang, Xin Wang, Shanghang Zhang, and Hong Mei. 2022. Dna: Domain generalization with diversified neural averaging. In International Conference on Machine Learning, pages 4010–4034. PMLR.
  25. 25.Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv preprint arXiv:1803.05457.
  26. 26.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805.
  27. 27.Shizhe Diao, Rui Pan, Hanze Dong, Ka Shun Shum, Jipeng Zhang, Wei Xiong, and Tong Zhang. 2023. Lmflow: An extensible toolkit for finetuning and inference of large foundation models. arXiv preprint arXiv:2306.12420.
  28. 28.Hanze Dong, Wei Xiong, Deepanshu Goyal, Rui Pan, Shizhe Diao, Jipeng Zhang, Kashun Shum, and Tong Zhang. 2023. Raft: Reward ranked finetuning for generative foundation model alignment. arXiv preprint arXiv:2304.06767.
  29. 29.Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2019. Drop: A reading comprehension benchmark requiring discrete reasoning over paragraphs. arXiv preprint arXiv:1903.00161.
  30. 30.Logan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Firdaus Janoos, Larry Rudolph, and Aleksander Madry. 2020. Implementation matters in deep policy gradients: A case study on ppo and trpo. arXiv preprint arXiv:2005.12729.
  31. 31.Leo Gao, John Schulman, and Jacob Hilton. 2023. Scaling laws for reward model overoptimization. In International Conference on Machine Learning, pages 10835–10866. PMLR.
  32. 32.Xinyang Geng and Hao Liu. 2023. Openllama: An open reproduction of llama.
  33. 33.Zheng Gong, Kun Zhou, Wayne Xin Zhao, Jing Sha, Shijin Wang, and Ji-Rong Wen. 2022. Continual pre-training of language models for math problem understanding with syntax-aware memory network. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5923–5933.
  34. 34.Google. 2023. Bard.
  35. 35.Sachin Goyal, Ananya Kumar, Sankalp Garg, Zico Kolter, and Aditi Raghunathan. 2022. Finetune like you pretrain: Improved finetuning of zero-shot vision models. arXiv preprint arXiv:2212.00638.
  36. 36.Caglar Gulcehre, Tom Le Paine, Srivatsan Srinivasan, Ksenia Konyushkova, Lotte Weerts, Abhishek Sharma, Aditya Siddhant, Alex Ahern, Miaosen Wang, Chenjie Gu, et al. 2023. Reinforced self-training (rest) for language modeling. arXiv preprint arXiv:2308.08998.
  37. 37.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778.
  38. 38.J. Edward Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, and Weizhu Chen. 2021. Lora: Low-rank adaptation of large language models. ArXiv, abs/2106.09685.
  39. 39.Yufan Huang, Yanzhe Zhang, Jiaao Chen, Xuezhi Wang, and Diyi Yang. 2021. Continual learning for text classification with information disentanglement based regularization. arXiv preprint arXiv:2104.05489.
  40. 40.Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023a. Mistral 7b. arXiv preprint arXiv:2310.06825.
  41. 41.Dongfu Jiang, Xiang Ren, and Bill Yuchen Lin. 2023b. Llm-blender: Ensembling large language models with pairwise ranking and generative fusion. arXiv preprint arXiv:2306.02561.
  42. 42.Xisen Jin, Dejiao Zhang, Henghui Zhu, Wei Xiao, Shang-Wen Li, Xiaokai Wei, Andrew Arnold, and Xiang Ren. 2021. Lifelong pretraining: Continually adapting language models to emerging corpora. arXiv preprint arXiv:2110.08534.
  43. 43.James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al. 2017. Overcoming catastrophic forgetting in neural networks. Proceedings of the national academy of sciences, 114(13):3521–3526.
  44. 44.Ananya Kumar, Aditi Raghunathan, Robbie Jones, Tengyu Ma, and Percy Liang. 2022. Fine-tuning can distort pretrained features and underperform out-of-distribution. arXiv preprint arXiv:2202.10054.
  45. 45.Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard Hovy. 2017. Race: Large-scale reading comprehension dataset from examinations. arXiv preprint arXiv:1704.04683.
  46. 46.Seunghyun Lee, Younggyo Seo, Kimin Lee, Pieter Abbeel, and Jinwoo Shin. 2022a. Offline-to-online reinforcement learning via balanced replay and pessimistic q-ensemble. In Conference on Robot Learning, pages 1702–1712. PMLR.
  47. 47.Yoonho Lee, Annie S. Chen, Fahim Tajwar, Ananya Kumar, Huaxiu Yao, Percy Liang, and Chelsea Finn. 2022b. Surgical fine-tuning improves adaptation to distribution shifts. ArXiv, abs/2210.11466.
  48. 48.Shengzhi Li, Rongyu Lin, and Shichao Pei. 2024. Multimodal preference alignment remedies regression of visual instruction tuning on language model. arXiv preprint arXiv:2402.10884.
  49. 49.Xuechen Li, Tianyi Zhang, Yann Dubois, Rohan Taori, Ishaan Gulrajani, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Alpacaeval: An automatic evaluator of instruction-following models. https://github.com/tatsu-lab/alpaca_eval.
  50. 50.Yong Lin, Hanze Dong, Hao Wang, and Tong Zhang. 2022a. Bayesian invariant risk minimization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16021–16030.
  51. 51.Yong Lin, Lu Tan, Yifan Hao, Honam Wong, Hanze Dong, Weizhong Zhang, Yujiu Yang, and Tong Zhang. 2023. Spurious feature diversification improves out-of-distribution generalization. arXiv preprint arXiv:2309.17230.
  52. 52.Yong Lin, Shengyu Zhu, Lu Tan, and Peng Cui. 2022b. Zin: When and how to learn invariance without environment partition? Advances in Neural Information Processing Systems, 35:24529–24542.
  53. 53.Tianqi Liu, Yao Zhao, Rishabh Joshi, Misha Khalman, Mohammad Saleh, Peter J Liu, and Jialu Liu. 2023. Statistical rejection sampling improves preference optimization. arXiv preprint arXiv:2309.06657.
  54. 54.Zihan Liu, Genta Indra Winata, and Pascale Fung. 2021. Continual mixed-language pre-training for extremely low-resource neural machine translation. arXiv preprint arXiv:2105.03953.
  55. 55.Andrea Madotto, Zhaojiang Lin, Zhenpeng Zhou, Seungwhan Moon, Paul Crook, Bing Liu, Zhou Yu, Eunjoon Cho, and Zhiguang Wang. 2020. Continual learning in task-oriented dialogue systems. arXiv preprint arXiv:2012.15504.
  56. 56.James L McClelland, Bruce L McNaughton, and Randall C O’Reilly. 1995. Why there are complementary learning systems in the hippocampus and neocortex: insights from the successes and failures of connectionist models of learning and memory. Psychological review, 102(3):419.
  57. 57.Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al. 2021. Webgpt: Browser-assisted question-answering with human feedback. arXiv preprint arXiv:2112.09332.
  58. 58.Michael Noukhovitch, Samuel Lavoie, Florian Strub, and Aaron Courville. 2023. Language model alignment with elastic reset. arXiv preprint arXiv:2312.07551.
  59. 59.Michael Noukhovitch, Samuel Lavoie, Florian Strub, and Aaron C Courville. 2024. Language model alignment with elastic reset. Advances in Neural Information Processing Systems, 36.
  60. 60.OpenAI. 2023. Gpt-4 technical report. ArXiv, abs/2303.08774.
  61. 61.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744.
  62. 62.Abhishek Panigrahi, Nikunj Saunshi, Haoyu Zhao, and Sanjeev Arora. 2023. Task-specific skill localization in fine-tuned language models. arXiv preprint arXiv:2302.06600.
  63. 63.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics, pages 311–318.
  64. 64.Yujia Qin, Jiajie Zhang, Yankai Lin, Zhiyuan Liu, Peng Li, Maosong Sun, and Jie Zhou. 2022. Elle: Efficient lifelong pre-training for emerging data. arXiv preprint arXiv:2203.06311.
  65. 65.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, pages 8748–8763. PMLR.
  66. 66.Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D Manning, and Chelsea Finn. 2023. Direct preference optimization: Your language model is secretly a reward model. arXiv preprint arXiv:2305.18290.
  67. 67.Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018. Know what you don’t know: Unanswerable questions for squad. arXiv preprint arXiv:1806.03822.
  68. 68.Rajkumar Ramamurthy, Prithviraj Ammanabrolu, Kianté Brantley, Jack Hessel, Rafet Sifa, Christian Bauckhage, Hannaneh Hajishirzi, and Yejin Choi. 2022. Is reinforcement learning (not) for natural language processing?: Benchmarks, baselines, and building blocks for natural language policy optimization. arXiv preprint arXiv:2210.01241.
  69. 69.Alexandre Rame, Guillaume Couairon, Corentin Dancette, Jean-Baptiste Gaya, Mustafa Shukor, Laure Soulier, and Matthieu Cord. 2024. Rewarded soups: towards pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards. Advances in Neural Information Processing Systems, 36.
  70. 70.Alexandre Ramé, Nino Vieillard, Léonard Hussenot, Robert Dadashi, Geoffrey Cideron, Olivier Bachem, and Johan Ferret. 2024. Warm: On the benefits of weight averaged reward models. arXiv preprint arXiv:2401.12187.
  71. 71.Anastasia Razdaibiedina, Yuning Mao, Rui Hou, Madian Khabsa, Mike Lewis, and Amjad Almahairi. 2023. Progressive prompts: Continual learning for language models. arXiv preprint arXiv:2301.12314.
  72. 72.Sylvestre-Alvise Rebuffi, Alexander Kolesnikov, Georg Sperl, and Christoph H Lampert. 2017. icarl: Incremental classifier and representation learning. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, pages 2001–2010.
  73. 73.Sylvestre-Alvise Rebuffi, Alexander Kolesnikov, Georg Sperl, Christoph H Lampert, and icarl. Incremental classifier and representation learning. In Conference on Computer Vision and Pattern Recognition (CVPR), pages 5533–5542.
  74. 74.Matthew Riemer, Ignacio Cases, Robert Ajemian, Miao Liu, Irina Rish, Yuhai Tu, and Gerald Tesauro. 2018. Learning to learn without forgetting by maximizing transfer and minimizing interference. arXiv preprint arXiv:1810.11910.
  75. 75.Hippolyt Ritter, Aleksandar Botev, and David Barber. 2018. Online structured laplace approximations for overcoming catastrophic forgetting. Advances in Neural Information Processing Systems, 31.
  76. 76.Victor Sanh, Albert Webson, Colin Raffel, Stephen H Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, et al. 2021. Multitask prompted training enables zero-shot task generalization. arXiv preprint arXiv:2110.08207.
  77. 77.Sunny Sanyal, Atula Tejaswi Neerkaje, Jean Kaddour, Abhishek Kumar, et al. 2023. Early weight averaging meets high learning rates for llm pre-training. In Workshop on Advancing Neural Network Training: Computational Efficiency, Scalability, and Resource Optimization (WANT@ NeurIPS 2023).
  78. 78.Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilic, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al. 2022. Bloom: A 176b-parameter open-access multilingual language model. arXiv preprint arXiv:2211.05100.
  79. 79.John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347.
  80. 80.Jonathan Schwarz, Wojciech Czarnecki, Jelena Luketina, Agnieszka Grabska-Barwinska, Yee Whye Teh, Razvan Pascanu, and Raia Hadsell. 2018. Progress & compress: A scalable framework for continual learning. In International conference on machine learning, pages 4528–4537. PMLR.
  81. 81.Hanul Shin, Jung Kwon Lee, Jaehong Kim, and Jiwon Kim. 2017. Continual learning with deep generative replay. Advances in neural information processing systems, 30.
  82. 82.Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556.
  83. 83.Ziang Song, Tianle Cai, Jason D Lee, and Weijie J Su. 2023. Reward collapse in aligning large language models. arXiv preprint arXiv:2305.17608.
  84. 84.Fan-Keng Sun, Cheng-Hao Ho, and Hung-Yi Lee. 2019. Lamol: Language modeling for lifelong language learning. arXiv preprint arXiv:1909.03329.
  85. 85.Xiaoyu Tan, LIN Yong, Shengyu Zhu, Chao Qu, Xihe Qiu, Xu Yinghui, Peng Cui, and Yuan Qi. 2023. Provably invariant learning without domain information.
  86. 86.Ross Taylor, Marcin Kardas, Guillem Cucurull, Thomas Scialom, Anthony Hartshorn, Elvis Saravia, Andrew Poulton, Viktor Kerkez, and Robert Stojnic. 2022. Galactica: A large language model for science. arXiv preprint arXiv:2211.09085.
  87. 87.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971.
  88. 88.Haoqin Tu, Bingchen Zhao, Chen Wei, and Cihang Xie. 2023. Sight beyond text: Multi-modal training enhances llms in truthfulness and ethics. arXiv preprint arXiv:2309.07120.
  89. 89.Lewis Tunstall, Edward Beeching, Nathan Lambert, Nazneen Rajani, Kashif Rasul, Younes Belkada, Shengyi Huang, Leandro von Werra, Clémentine Fourrier, Nathan Habib, et al. 2023. Zephyr: Direct distillation of lm alignment. arXiv preprint arXiv:2310.16944.
  90. 90.Jeffrey S Vitter. 1985. Random sampling with a reservoir. ACM Transactions on Mathematical Software (TOMS), 11(1):37–57.
  91. 91.Chaoqi Wang, Yibo Jiang, Chenghao Yang, Han Liu, and Yuxin Chen. 2023a. Beyond reverse kl: Generalizing direct preference optimization with diverse divergence constraints. arXiv preprint arXiv:2309.16240.
  92. 92.Liyuan Wang, Xingxing Zhang, Hang Su, and Jun Zhu. 2023b. A comprehensive survey of continual learning: Theory, method and application. arXiv preprint arXiv:2302.00487.
  93. 93.Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2022. Self-instruct: Aligning language model with self generated instructions. arXiv preprint arXiv:2212.10560.
  94. 94.Mitchell Wortsman, Gabriel Ilharco, Samir Ya Gadre, Rebecca Roelofs, Raphael Gontijo-Lopes, Ari S Morcos, Hongseok Namkoong, Ali Farhadi, Yair Carmon, Simon Kornblith, et al. 2022a. Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time. In International Conference on Machine Learning, pages 23965–23998. PMLR.
  95. 95.Mitchell Wortsman, Gabriel Ilharco, Jong Wook Kim, Mike Li, Simon Kornblith, Rebecca Roelofs, Raphael Gontijo Lopes, Hannaneh Hajishirzi, Ali Farhadi, Hongseok Namkoong, et al. 2022b. Robust fine-tuning of zero-shot models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7959–7971.
  96. 96.Mitchell Wortsman, Gabriel Ilharco, Mike Li, Jong Wook Kim, Hannaneh Hajishirzi, Ali Farhadi, Hongseok Namkoong, and Ludwig Schmidt. 2021. Robust fine-tuning of zero-shot models. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 7949–7961.
  97. 97.Jeff Wu, Long Ouyang, Daniel M Ziegler, Nisan Stiennon, Ryan Lowe, Jan Leike, and Paul Christiano. 2021a. Recursively summarizing books with human feedback. arXiv preprint arXiv:2109.10862.
  98. 98.Tongtong Wu, Massimo Caccia, Zhuang Li, Yuan-Fang Li, Guilin Qi, and Gholamreza Haffari. 2021b. Pretrained language model in continual learning: A comparative study. In International Conference on Learning Representations.
  99. 99.Tongtong Wu, Massimo Caccia, Zhuang Li, Yuan-Fang Li, Guilin Qi, and Gholamreza Haffari. 2022. Pretrained language model in continual learning: A comparative study. In International Conference on Learning Representations.
  100. 100.Wei Xiong, Hanze Dong, Chen Ye, Han Zhong, Nan Jiang, and Tong Zhang. 2023. Gibbs sampling from human feedback: A provable kl- constrained framework for rlhf.
  101. 101.LI Xuhong, Yves Grandvalet, and Franck Davoine. 2018. Explicit inductive bias for transfer learning with convolutional networks. In International Conference on Machine Learning, pages 2825–2834. PMLR.
  102. 102.Enneng Yang, Zhenyi Wang, Li Shen, Shiwei Liu, Guibing Guo, Xingwei Wang, and Dacheng Tao. 2023. Adamerging: Adaptive model merging for multi-task learning. arXiv preprint arXiv:2310.02575.
  103. 103.Jason Yosinski, Jeff Clune, Anh Nguyen, Thomas Fuchs, and Hod Lipson. 2015. Understanding neural networks through deep visualization. arXiv preprint arXiv:1506.06579.
  104. 104.Pengfei Yu and Heng Ji. 2023. Self information update for large language models through mitigating exposure bias. In arxiv.
  105. 105.Pengfei Yu, Heng Ji, and Premkumar Natarajan. 2021. Lifelong event detection with knowledge transfer. In Proc. The 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP2021).
  106. 106.Zheng Yuan, Hongyi Yuan, Chuanqi Tan, Wei Wang, Songfang Huang, and Fei Huang. 2023. Rrhf: Rank responses to align language models with human feedback without tears. arXiv preprint arXiv:2304.05302.
  107. 107.Matthew D Zeiler and Rob Fergus. 2014. Visualizing and understanding convolutional networks. In Computer Vision– ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part I 13, pages 818–833. Springer.
  108. 108.Michael Zhang and Christopher Ré. 2022. Contrastive adapters for foundation model group robustness. arXiv preprint arXiv:2207.07180.
  109. 109.Tong Zhang. 2023. Mathematical Analysis of Machine Learning Algorithms. Cambridge University Press.
  110. 110.Yanzhe Zhang, Xuezhi Wang, and Diyi Yang. 2022. Continual sequence generation with adaptive compositional modules. arXiv preprint arXiv:2203.10652.
  111. 111.Yao Zhao, Rishabh Joshi, Tianqi Liu, Misha Khalman, Mohammad Saleh, and Peter J Liu. 2023. Slic-hf: Sequence likelihood calibration with human feedback. arXiv preprint arXiv:2305.10425.
  112. 112.Rui Zheng, Shihan Dou, Songyang Gao, Yuan Hua, Wei Shen, Binghai Wang, Yan Liu, Senjie Jin, Qin Liu, Yuhao Zhou, et al. 2023. Secrets of rlhf in large language models part i: Ppo. arXiv preprint arXiv:2307.04964.
  113. 113.Xiao Zhou, Yong Lin, Renjie Pi, Weizhong Zhang, Renzhe Xu, Peng Cui, and Tong Zhang. 2022a. Model agnostic sample reweighting for out-of-distribution learning. In International Conference on Machine Learning, pages 27203–27221. PMLR.
  114. 114.Xiao Zhou, Yong Lin, Weizhong Zhang, and Tong Zhang. 2022b. Sparse invariant risk minimization. In International Conference on Machine Learning, pages 27222–27244. PMLR.
  115. 115.Banghua Zhu, Jiantao Jiao, and Michael I Jordan. 2023. Principled reinforcement learning with human feedback from pairwise or k-wise comparisons. arXiv preprint arXiv:2301.11270.
  116. 116.Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving. 2019. Fine-tuning language models from human preferences. arXiv preprint arXiv:1909.08593.

Citation

MLA
Lin, Y., et al. “Mitigating the Alignment Tax of RLHF”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024, pp. 580–606, https://doi.org/10.18653/v1/2024.emnlp-main.35.
APA
Lin, Y., Lin, H., Xiong, W., Diao, S., Liu, J., Zhang, J., Pan, R., Wang, H., Hu, W., Zhang, H., Dong, H., Pi, R., Zhao, H., Jiang, N., Ji, H., Yao, Y., & Zhang, T. (2024). Mitigating the Alignment Tax of RLHF. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 580–606. https://doi.org/10.18653/v1/2024.emnlp-main.35
Chicago
Lin, Y., H. Lin, W. Xiong, et al. 2024. “Mitigating the Alignment Tax of RLHF”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 580–606. https://doi.org/10.18653/v1/2024.emnlp-main.35.
Harvard
Lin, Y. et al. (2024) “Mitigating the Alignment Tax of RLHF”, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 580–606. Available at: https://doi.org/10.18653/v1/2024.emnlp-main.35.
Vancouver
1. Lin Y, Lin H, Xiong W, et al (2024) Mitigating the Alignment Tax of RLHF. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 580–606

BibTeX

@inproceedings{lin-etal-2024-mitigating,
    title = "Mitigating the Alignment Tax of {RLHF}",
    author = "Lin, Yong  and
      Lin, Hangyu  and
      Xiong, Wei  and
      Diao, Shizhe  and
      Liu, Jianmeng  and
      Zhang, Jipeng  and
      Pan, Rui  and
      Wang, Haoxiang  and
      Hu, Wenbin  and
      Zhang, Hanning  and
      Dong, Hanze  and
      Pi, Renjie  and
      Zhao, Han  and
      Jiang, Nan  and
      Ji, Heng  and
      Yao, Yuan  and
      Zhang, Tong",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.35/",
    doi = "10.18653/v1/2024.emnlp-main.35",
    pages = "580--606"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/