Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study
Shusheng XuWei FuJiaxuan GaoWenjie YeWeilin LiuZhiyu MeiGuangju WangChao YuYi Wu
Demonstrates that properly configured Proximal Policy Optimization consistently surpasses Direct Preference Optimization across diverse alignment benchmarks, identifying fundamental limitations in reward-free methods and isolating the critical training factors needed to achieve superior performance in complex tasks like code generation.
Aligning large language models with human expectations is critical for ensuring reliable, safe, and effective artificial intelligence systems. Organizations currently face a major strategic dilemma between two competing alignment techniques: reward-based methods such as Proximal Policy Optimization, which are used by leading industry applications like ChatGPT and Claude, and newer reward-free methods like Direct Preference Optimization, which have gained rapid popularity in open-source research due to their simplicity. This tension creates uncertainty for engineering teams deciding how to invest computing resources and design model post-training pipelines.
The article investigates whether Direct Preference Optimization is fundamentally superior to Proximal Policy Optimization and determines the specific training practices needed to achieve optimal performance. To answer these questions, the authors analyze the theoretical properties of Direct Preference Optimization, conduct controlled synthetic experiments, and perform extensive empirical evaluations across dialogue generation (using the HH-RLHF and SafeRLHF datasets) and complex programming benchmarks (using APPS and CodeContest). The study evaluates language models ranging from 7 billion to 34 billion parameters across diverse task difficulties and reward mechanisms.
The findings show that properly tuned Proximal Policy Optimization consistently outperforms Direct Preference Optimization across all evaluated domains. First, theoretical and empirical analyses reveal that Direct Preference Optimization is vulnerable to distribution shifts between model outputs and static preference data, often assigning high probabilities to out-of-distribution responses and producing erratic behaviors. Second, the authors identify three key operational practices that dramatically enhance Proximal Policy Optimization: normalizing policy advantages, utilizing large batch sizes, and updating the baseline reference model with an exponential moving average. Third, while iterative variants of Direct Preference Optimization mitigate some distribution mismatch in dialogue tasks, they fail on complex reasoning tasks; in competitive programming evaluations, a 34-billion-parameter model aligned with Proximal Policy Optimization achieved state-of-the-art results (increasing the CodeContest pass rate from 16.4% to 22.4%), whereas Direct Preference Optimization degraded to near-zero pass rates.
These results demonstrate that organizations pursuing advanced model alignment should not abandon reinforcement learning frameworks in favor of simpler reward-free short-cuts. While Direct Preference Optimization offers lower initial engineering overhead, it presents significant risks in complex reasoning domains where distribution shifts cause severe performance degradation. Proximal Policy Optimization remains the more robust and capable alignment framework when configured with adequate batch sizes and proper reference model updating.
For practical implementation, technical teams should maintain reinforcement learning pipelines utilizing Proximal Policy Optimization for high-stakes and complex reasoning tasks. When employing Direct Preference Optimization for simpler conversational applications, teams should adopt iterative retraining on newly generated model outputs rather than relying on static datasets. Decision-makers should note that these conclusions assume access to reliable reward models or direct execution feedback; future work remains necessary to establish best practices for training robust reward models across non-deterministic domains.
- Paper: Direct Preference Optimization: Your Language Model is Secretly a Reward Model, Rafael Rafailov et al. (2023). This paper introduces Direct Preference Optimization (DPO), defining the core reward-free alignment method that the source paper directly investigates and critiques.
- Paper: Proximal Policy Optimization Algorithms, John Schulman et al. (2017). This work establishes the Proximal Policy Optimization (PPO) algorithm, providing the foundational reinforcement learning mechanics that the source paper evaluates and optimizes for LLM alignment.
- Paper: Fine-Tuning Language Models from Human Preferences, Daniel M. Ziegler et al. (2019). This seminal paper introduces fine-tuning language models with human preferences via reward modeling and PPO, establishing the standard reward-based RLHF pipeline examined in the source.
- Paper: Learning to summarize from human feedback, Nisan Stiennon et al. (2020). This work formulates the standard modern RLHF workflow using reward models and KL-regularized PPO fine-tuning that serves as the baseline pipeline in the source paper's study.
- Paper: Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback, Yuntao Bai et al. (2022). This paper demonstrates scaling PPO-based RLHF with iterated online data collection on dialogue benchmarks, providing key background on the practical performance of reward-based alignment.
- Paper: Principled Reinforcement Learning with Human Feedback from Pairwise or K-wise Comparisons, Banghua Zhu et al. (2023). This paper provides theoretical foundations for Bradley-Terry preference estimation and policy learning in RLHF that underpin the algorithmic analyses of DPO and PPO.
- Paper: Scaling Laws for Reward Model Overoptimization, Leo Gao et al. (2023). This work formalizes scaling laws and policy degradation under reward model overoptimization with PPO, contextualizing the fundamental trade-offs between reward-based and reward-free alignment.
- Paper: From RLHF to Direct Alignment: A Theoretical Unification of Preference Learning for Large Language Models, Tarun Raheja et al. (2026). This work theoretically unifies DPO, PPO, and subsequent direct alignment algorithms, extending the source paper's comparative insights into a broader taxonomy of preference learning.
- Paper: Human Alignment of Large Language Models through Online Preference Optimisation, Daniele Calandriello et al. (2024). This paper addresses the offline limitations of DPO identified in the source by developing online preference optimization techniques through dynamic sampling and self-play.
- Paper: Mitigating the Alignment Tax of RLHF, Yong Lin et al. (2024). This study builds on the alignment trade-offs of DPO and PPO by introducing weight-averaging strategies to alleviate the alignment tax and catastrophic forgetting in tuned models.
- Paper: DAPO: An Open-Source LLM Reinforcement Learning System at Scale, Qiying Yu et al. (2025). This paper continues the source's investigation into scaling LLM reinforcement learning by providing an open-source system and algorithmic stabilization techniques for complex reasoning tasks.
- Paper: VIMPO: Value-Implicit Policy Optimization for LLMs, Zhewei Kang et al. (2026). This work proposes a value-implicit policy optimization method that aims to overcome the instability of critic training in PPO while avoiding the token-level credit assignment issues of reward-free methods.
- Paper: HybridFlow: A Flexible and Efficient RLHF Framework, Guangming Sheng et al. (2024). This paper presents an efficient distributed execution system that resolves practical computational and memory bottlenecks when deploying large-scale PPO workflows.
- Paper: ORPO: Monolithic Preference Optimization without Reference Model, Jiwoo Hong et al. (2024). This work develops monolithic odds ratio preference optimization to streamline alignment into a single training step without the multi-model complexity of PPO or reference models.
- Paper: Dense Reward for Free in Reinforcement Learning from Human Feedback, Alex James Chan et al. (2024). This paper enhances PPO training by extracting dense, token-level reward signals directly from standard preference reward models, mitigating sparse feedback inefficiencies.
