Mitigating the Alignment Tax of RLHF
Yong LinHangyu LinWei XiongShizhe DiaoJianmeng LiuJipeng ZhangRui PanHaoxiang WangWenbin HuHanning Zhang
Proposes Heterogeneous Model Averaging, a layer-adaptive weight interpolation technique between pre- and post-RLHF models that optimizes alignment rewards while preserving general NLP capabilities across diverse model scales.
Large language models acquire broad general capabilities during pre-training, such as reading comprehension, common sense reasoning, and translation. However, post-training alignment techniques like Reinforcement Learning with Human Feedback (RLHF), designed to make models helpful, honest, and harmless, frequently cause catastrophic forgetting of these core abilities—a phenomenon known as the "alignment tax." As organizations increasingly rely on aligned models for production tasks, mitigating this performance degradation without sacrificing safety and human preference alignment has become a critical operational challenge.
The article evaluates methods to resolve this alignment-forgetting trade-off and demonstrates that model weight averaging provides an exceptionally effective and computationally practical solution. It introduces and evaluates Heterogeneous Model Averaging (HMA), an approach that dynamically assigns distinct interpolation weights to different model layers to maximize alignment rewards while preserving pre-trained skills.
To conduct this evaluation, the researchers tested alignment algorithms—including Rejection Sampling Fine-Tuning (RSF), Direct Preference Optimization (DPO), and Proximal Policy Optimization (PPO)—on the OpenLLaMA-3B foundation model and extended their validation to larger architectures like Mistral-7B and Gemma-7B (specifically Zephyr-7B variants). They measured alignment tax across multiple standard natural language benchmarks (such as ARC, SQuAD, DROP, and WMT translation) and assessed alignment quality using both specialized reward models and GPT-4 evaluations. The analysis systematically compared simple model averaging and HMA against established alternatives, including parameter regularization (L1/L2 penalties), knowledge distillation, low-rank adaptation (LoRA), reward penalties, and experience replay using subsets of pre-training data.
The study yielded several key findings. First, simple model weight averaging between pre-RLHF and post-RLHF checkpoints consistently established a superior trade-off boundary compared to complex regularization, distillation, LoRA, and even experience replay methods. Second, replaying pre-training data failed to match model averaging on two of three core benchmark suites, despite adding four times the data volume of the RLHF dataset (400 million tokens) and incurring heavy computational overhead. Third, the researchers established both theoretically and empirically that averaging lower-level transformer layers yields the greatest joint improvements in alignment and general task performance, because these layers share broad, foundational feature spaces across tasks. Fourth, the proposed Heterogeneous Model Averaging framework pushed the performance boundary further across all algorithms: setting a moderate average ratio (around 0.2) preserved baseline capabilities while outperforming baseline models, achieving higher human-preference win rates on AlpacaEval (e.g., 9.32% vs. 8.10% against GPT-4 on Zephyr-7B-β) and improved scores across reading comprehension, common sense, and translation tasks.
These findings indicate that teams deploying aligned language models do not need to choose between severe capability regression and expensive, complex mitigation pipelines. Model averaging and HMA operate strictly as post-processing steps on existing checkpoints, eliminating the extreme computational costs, training instabilities, and proprietary data access hurdles associated with data replay or constrained optimization. This directly lowers infrastructure expenses, reduces project timelines, and enhances model safety and performance simultaneously.
Organizations training or fine-tuning language models should adopt model averaging techniques as a standard post-alignment pipeline stage. Practitioners should prioritize setting an overall averaging ratio around 0.2, retaining heavier weight from the pre-RLHF model on lower layers while allowing higher layers to retain more alignment-specific adjustments. When feasible, HMA should be applied via reward proxy distillation on a small sample of generated outputs rather than relying on heavy multi-task tuning.
While the article demonstrates high confidence and consistent empirical validation across multiple model sizes and alignment frameworks, it notes that HMA substantially mitigates but does not entirely eliminate the alignment tax. Decision-makers should recognize that optimal averaging ratios may still require minor empirical verification across distinct domain distributions, and future work is required to establish the theoretical lower limit of capability degradation during alignment.
- Paper: Training language models to follow instructions with human feedback, Long Ouyang et al. (2022). It introduces the foundational RLHF alignment pipeline that the source specifically analyzes to resolve the resulting catastrophic forgetting of core pre-trained capabilities.
- Paper: Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback, Yuntao Bai et al. (2022). It formalizes the challenge of training helpful and harmless assistants using RLHF while introducing the concept of the alignment tax on NLP benchmarks.
- Paper: Scaling Laws for Reward Model Overoptimization, Leo Gao et al. (2023). It establishes empirical scaling laws and KL-divergence trade-offs in reward model optimization, explaining why intense RLHF degrades baseline model behaviors.
- Paper: LoRA: Low-Rank Adaptation of Large Language Models, Edward J. Hu et al. (2022). It introduces Low-Rank Adaptation (LoRA), which serves as one of the baseline parameter-efficient mitigation techniques directly benchmarked against model averaging in the source.
- Paper: Experience Replay for Continual Learning, David Rolnick et al. (2018). It establishes experience replay strategies for continual learning, providing the baseline replay framework that the source tests and outperforms with post-hoc model averaging.
- Paper: An Empirical Investigation of the Role of Pre-training in Lifelong Learning, Sanket Vaibhav Mehta et al. (2023). It demonstrates how pre-trained weight initializations mitigate catastrophic forgetting during sequential fine-tuning, underpinning the source's reliance on pre-trained checkpoint representations.
- Paper: Lifelong Pretraining: Continually Adapting Language Models to Emerging Corpora, Xisen Jin et al. (2022). It evaluates continual learning and distillation baselines against catastrophic forgetting in sequential updates, contextualizing the alternative mitigation methods tested in the source.
- Paper: Theoretical guarantees on the best-of-n alignment policy, Ahmad Beirami et al. (2025). It provides theoretical guarantees and bounds on distribution drift and policy quality under best-of-n sampling, expanding the mathematical understanding of post-training alignment trade-offs.
- Paper: From RLHF to Direct Alignment: A Theoretical Unification of Preference Learning for Large Language Models, Tarun Raheja et al. (2026). It unifies RLHF and direct preference methods under a single theoretical framework, offering deeper insight into the various alignment objectives evaluated in the source.
- Paper: DeAL: Decoding-time Alignment for Large Language Models, James Y. Huang et al. (2025). It explores decoding-time alignment as an alternative paradigm to avoid retraining and mitigate capability loss without modifying model parameters.
- Paper: Aligning Large Language Models with Representation Editing: A Control Perspective, Lingkai Kong et al. (2024). It introduces representation editing during inference to steer safety and helpfulness dynamically without fine-tuning or altering base model weights.
- Paper: Lisa: Lazy Safety Alignment for Large Language Models against Harmful Fine-tuning Attack, Tiansheng Huang et al. (2024). It develops proximal constraints to prevent alignment drift and preserve safety guardrails during downstream user fine-tuning.
- Paper: What Makes and Breaks Safety Fine-tuning? A Mechanistic Study, Samyak Jain et al. (2024). It analyzes the internal geometric and weight-space mechanisms of safety fine-tuning to explain why alignment alters specific representation subspaces.
- Paper: Unintended Impacts of LLM Alignment on Global Representation, Michael J. Ryan et al. (2024). It evaluates the unintended downstream effects of standard alignment procedures on multilingual capabilities and global dialect representation.
- Paper: HybridFlow: A Flexible and Efficient RLHF Framework, Guangming Sheng et al. (2024). It tackles the system-level distributed infrastructure challenges of executing multi-model RLHF pipelines efficiently at scale.
