DPO Meets PPO: Reinforced Token Optimization for RLHF
Han ZhongZikang ShanGuhao FengWei XiongXinle ChengLi ZhaoDi HeJiang BianLiwei Wang
Proposes Reinforced Token Optimization, a framework that extracts fine-grained token-level rewards from Direct Preference Optimization to guide Proximal Policy Optimization, significantly outperforming standard RLHF methods on AlpacaEval 2 and Arena-Hard.
Aligning large language models with human preferences is essential for making artificial intelligence systems helpful, safe, and reliable. The standard industry approach, Reinforcement Learning from Human Feedback, typically relies on Proximal Policy Optimization to train models using overall sentence-level reward scores. However, open-source implementations of this standard approach remain suboptimal and sample-inefficient because sparse sentence-level rewards are assigned only at the very end of generated text, creating a mismatch with multi-step reinforcement learning algorithms that benefit from feedback at every generation step.
The article develops a new framework that formulates model alignment as a Markov decision process with fine-grained, token-level feedback, and introduces an algorithm named Reinforced Token Optimization. The objective of the article is to prove theoretically and demonstrate empirically that extracting dense, word-by-word reward signals from preference data and optimizing them with reinforcement learning significantly improves model performance and training efficiency.
To evaluate this approach, the authors established mathematical guarantees on sample complexity and designed a practical two-step pipeline. In this pipeline, Direct Preference Optimization is used to derive an implicit, step-by-step token reward from offline preference data, which is then paired with standard Proximal Policy Optimization updates and an optional compact sentence-level reward. The authors tested this implementation using a Llama-3-8B model fine-tuned on the UltraFeedback dataset, evaluating performance against standard benchmarks such as AlpacaEval 2 and Arena-Hard, alongside further experiments on text summarization.
The findings show that Reinforced Token Optimization substantially outperforms existing alignment techniques. Specifically, it exceeds standard Proximal Policy Optimization by 7.5 points on length-controlled AlpacaEval 2 and by 4.1 points on style-controlled Arena-Hard, while also surpassing popular direct preference learning methods like SimPO and token-level Direct Preference Optimization. In addition, the method matches standard Proximal Policy Optimization performance using only one-eighth of the training data and continues to scale in quality as dataset size increases, whereas conventional methods plateau early. Ablation analyses confirmed that assigning dense rewards across individual tokens acts as an effective reward-shaping mechanism that accelerates and stabilizes model training.
These results demonstrate that organizations can achieve superior model alignment and conversational quality at comparable computational costs without requiring expensive human token-level annotations. The framework reduces data collection requirements and mitigates reward over-optimization issues by combining fine-grained token shaping with lightweight sequence-level scoring. When making deployment decisions, organizations should adopt this hybrid token-level reinforcement approach for fine-tuning foundation models, keeping the token-level scaling factor small to prevent policy distortion while testing alternative reinforcement algorithms like REINFORCE-type updates for simpler pipelines.
The authors express strong theoretical and empirical confidence in their results across standard open-source evaluation benchmarks and diverse tasks like summarization. However, practical implementations remain bounded by assumptions of deterministic state transitions in standard text generation, meaning caution and further empirical analysis are advised if applying the framework to stochastic environments or external tool-use scenarios.
- Paper: Proximal Policy Optimization Algorithms, John Schulman et al. (2017). PPO’s clipped policy-update objective is the optimization backbone that RTO adapts to token-level rewards.
- Paper: Direct Preference Optimization: Your Language Model is Secretly a Reward Model, Rafael Rafailov et al. (2023). RTO explicitly derives its token-wise quality signal from DPO, so DPO’s preference-learning formulation clarifies the first stage of the method.
- Paper: Fine-Tuning Language Models from Human Preferences, Daniel M. Ziegler et al. (2019). This foundational RLHF pipeline explains how human comparisons train a reward model that guides PPO, the sparse-reward setup RTO seeks to refine.
- Paper: Learning to summarize from human feedback, Nisan Stiennon et al. (2020). Its human-preference reward modeling and PPO summarization pipeline provides a concrete antecedent for RTO’s attempt to replace sequence-level feedback with token-level signals.
- Paper: Dense Reward for Free in Reinforcement Learning from Human Feedback, Alex James Chan et al. (2024). Its method densifies RLHF feedback into token-level rewards, making it a direct prerequisite for understanding RTO’s fine-grained credit assignment.
- Paper: Trust Region Policy Optimization, John Schulman et al. (2015). TRPO’s trust-region policy optimization supplies the stability-oriented policy-gradient lineage from which PPO, used by RTO, developed.
- Paper: Test-Time Preference Optimization: On-the-Fly Alignment via Iterative Textual Feedback, Yafu Li et al. (2025). This inference-time alternative carries preference-guided alignment beyond RTO’s learned token rewards and parameter-updating PPO into iterative response refinement.
