DPO Meets PPO: Reinforced Token Optimization for RLHF

Han ZhongZikang ShanGuhao FengWei XiongXinle ChengLi ZhaoDi HeJiang BianLiwei Wang

article2025ICML144 citations

Proposes Reinforced Token Optimization, a framework that extracts fine-grained token-level rewards from Direct Preference Optimization to guide Proximal Policy Optimization, significantly outperforming standard RLHF methods on AlpacaEval 2 and Arena-Hard.

Listen

Aligning large language models with human preferences is essential for making artificial intelligence systems helpful, safe, and reliable. The standard industry approach, Reinforcement Learning from Human Feedback, typically relies on Proximal Policy Optimization to train models using overall sentence-level reward scores. However, open-source implementations of this standard approach remain suboptimal and sample-inefficient because sparse sentence-level rewards are assigned only at the very end of generated text, creating a mismatch with multi-step reinforcement learning algorithms that benefit from feedback at every generation step.

The article develops a new framework that formulates model alignment as a Markov decision process with fine-grained, token-level feedback, and introduces an algorithm named Reinforced Token Optimization. The objective of the article is to prove theoretically and demonstrate empirically that extracting dense, word-by-word reward signals from preference data and optimizing them with reinforcement learning significantly improves model performance and training efficiency.

To evaluate this approach, the authors established mathematical guarantees on sample complexity and designed a practical two-step pipeline. In this pipeline, Direct Preference Optimization is used to derive an implicit, step-by-step token reward from offline preference data, which is then paired with standard Proximal Policy Optimization updates and an optional compact sentence-level reward. The authors tested this implementation using a Llama-3-8B model fine-tuned on the UltraFeedback dataset, evaluating performance against standard benchmarks such as AlpacaEval 2 and Arena-Hard, alongside further experiments on text summarization.

The findings show that Reinforced Token Optimization substantially outperforms existing alignment techniques. Specifically, it exceeds standard Proximal Policy Optimization by 7.5 points on length-controlled AlpacaEval 2 and by 4.1 points on style-controlled Arena-Hard, while also surpassing popular direct preference learning methods like SimPO and token-level Direct Preference Optimization. In addition, the method matches standard Proximal Policy Optimization performance using only one-eighth of the training data and continues to scale in quality as dataset size increases, whereas conventional methods plateau early. Ablation analyses confirmed that assigning dense rewards across individual tokens acts as an effective reward-shaping mechanism that accelerates and stabilizes model training.

These results demonstrate that organizations can achieve superior model alignment and conversational quality at comparable computational costs without requiring expensive human token-level annotations. The framework reduces data collection requirements and mitigates reward over-optimization issues by combining fine-grained token shaping with lightweight sequence-level scoring. When making deployment decisions, organizations should adopt this hybrid token-level reinforcement approach for fine-tuning foundation models, keeping the token-level scaling factor small to prevent policy distortion while testing alternative reinforcement algorithms like REINFORCE-type updates for simpler pipelines.

The authors express strong theoretical and empirical confidence in their results across standard open-source evaluation benchmarks and diverse tasks like summarization. However, practical implementations remain bounded by assumptions of deterministic state transitions in standard text generation, meaning caution and further empirical analysis are advised if applying the framework to stochastic environments or external tool-use scenarios.

Cover for DPO Meets PPO: Reinforced Token Optimization for RLHF

Abstract

In the classical Reinforcement Learning from Human Feedback (RLHF) framework, Proximal Policy Optimization (PPO) is employed to learn from sparse, sentence-level rewards—a challenging scenario in traditional deep reinforcement learning. Despite the great successes of PPO in the alignment of large language models, its open-source implementation is still largely sub-optimal. To address these issues, we introduce a framework that models RLHF problems as a Markov decision process (MDP), enabling the capture of fine-grained token-wise information. Under this framework, we introduce an algorithm Reinforced Token Optimization (RTO), which learns the token-wise reward function from preference data and performs policy optimization based on this learned token-wise reward signal.Theoretically, RTO is proven to have the capability of finding the near-optimal policy sample-efficiently. For its practical implementation, RTO innovatively integrates Direct Preference Optimization (DPO) and PPO. DPO, originally derived from sparse sentence rewards, surprisingly provides us with a token-wise characterization of response quality, which is seamlessly incorporated into our subsequent PPO training stage. Extensive experiments demonstrate that RTO performs better than PPO and other direct preference learning algorithms. In particular, RTO outperforms PPO by 7.5 points on the AlpacaEval 2 benchmark and by 4.1 points on Arena-Hard. Our code and models are available at https://github.com/zkshan2002/RTO.

Table of Contents

  • 1. Introduction
  • 1.1. Related Works
  • 1.2. Notation
  • 2. Preliminaries
  • 3. RLHF Formulation: From Bandit to MDP
  • 3.1. MDP Formulation for RLHF
  • 3.2. Learning Objective
  • 3.3. Advantages of Token-Wise MDP over Sentence-Wise Bandit
  • 4. Reinforced Token Optimization
  • 4.1. Theoretical Version with Sample Complexity Guarantee
  • 4.2. Practical Implementation
  • 5. Experiments
  • 5.1. Benchmark Results
  • 5.2. In-depth Analysis of RTO Performance
  • 6. Conclusion
  • Impact Statement
  • Acknowledgement
  • References
  • A. Detailed Proofs
  • A.1. Proof of Proposition 3.2
  • A.2. Proof of Theorem 4.2
  • A.3. Proof of Lemma A.2
  • B. Variants of Reinforced Token Optimization
  • C. Additional Discussions
  • C.1. Direct Preference Optimization
  • C.2. Autoregressive Policy
  • D. Implementation Details
  • E. Additional Experiments on REINFORCE-type Algorithm
  • F. Additional Experiments on Summarization Task
  • F.1. Experimental Setup
  • F.2. Examples of Datasets
  • F.3. Training Configurations of TL;DR
  • F.4. Evaluation Details

Knowls

  1. Knowl 1 — RLHF as an autoregressive token-level MDP

    model/method

    For a prompt xx drawn from distribution ρ\rho, text generation can be modeled as an MDP whose state at step hh is sh=(x,y1:h−1)s_h=(x,y_{1:h-1}) and whose action is the next vocabulary token ah=yha_h=y_h. The transition appends the selected token and is deterministic in ordinary autoregressive generation; generation ends at an end-of-sentence token or the maximum length HH. A token reward r(sh,ah)r(s_h,a_h) assigns feedback to each generation step. For two trajectories τ1,τ2\tau^1,\tau^2 of at most HH tokens, the Bradley–Terry preference model is

    P(τ1≻τ2)=σ ⁣(∑h=1Hr(sh1,ah1)−∑h=1Hr(sh2,ah2)),P(\tau^1\succ\tau^2)=\sigma\!\left(\sum_{h=1}^{H}r(s_h^1,a_h^1)-\sum_{h=1}^{H}r(s_h^2,a_h^2)\right),

    where σ(z)=1/(1+e−z)\sigma(z)=1/(1+e^{-z}) and trajectories that terminate early can be padded with absorbing, zero-reward steps. Unlike a sentence-level bandit reward, this formulation represents both the autoregressive generation process and the reward at each token.

  2. Knowl 2 — KL-regularized objective for token-level RLHF

    equation

    The policy is optimized for cumulative token reward while remaining close to a reference policy, usually an SFT language model. For prompt distribution ρ\rho, policy π\pi, reference policy πref\pi_{\mathrm{ref}}, reward rr, penalty coefficient β>0\beta>0, and maximum response length HH, define

    Jβ(π)=Ex∼ρ, y∼π(⋅∣x)[∑h=1H(r(sh,ah)−βlog⁡π(ah∣sh)πref(ah∣sh))],J_\beta(\pi)=\mathbb{E}_{x\sim\rho,\,y\sim\pi(\cdot\mid x)}\left[\sum_{h=1}^{H}\left(r(s_h,a_h)-\beta\log\frac{\pi(a_h\mid s_h)}{\pi_{\mathrm{ref}}(a_h\mid s_h)}\right)\right],

    where sh=(x,y1:h−1)s_h=(x,y_{1:h-1}) and ah=yha_h=y_h. The optimal policy maximizes JβJ_\beta; the per-token log-ratio penalty implements the KL regularization and preserves a stochastic, reference-relative objective rather than unconstrained reward maximization.

  3. Knowl 3 — Token feedback reduces the response-search sample complexity

    theoretical result

    Consider a deterministic generation problem with a fixed prompt, vocabulary/action-set size AA, and response horizon HH. Let π∗\pi^* be an autoregressive policy for which at least one response yy satisfies π∗(y∣x)≥A−ξ\pi^*(y\mid x)\ge A^{-\xi}, with 0≤ξ≤H0\le\xi\le H. The paper establishes that identifying the best length-HH response from sentence-level rewards requires AHA^H response/reward observations. When token-wise rewards are available, an algorithm can identify the optimal response with at most Amin⁡{ξ+1,H}A^{\min\{\xi+1,H\}} observations. Thus, when ξ\xi is much smaller than HH, token-level feedback can substantially reduce the exploration burden.

  4. Knowl 4 — Offline RTO learns a pessimistic token reward

    model/method

    The theoretical offline version of Reinforced Token Optimization (RTO) assumes preference pairs of trajectories sharing a prompt and a linear token reward rθ(s,a)=ϕ(s,a)⊤θr_\theta(s,a)=\phi(s,a)^\top\theta, where the known feature vector ϕ(s,a)∈Rd\phi(s,a)\in\mathbb{R}^d obeys ∥ϕ(s,a)∥2≤L\|\phi(s,a)\|_2\le L and the unknown parameter obeys ∥θ∗∥2≤B\|\theta^*\|_2\le B. For a trajectory τ\tau, write ψ(τ)=∑h=1Hϕ(sh,ah)\psi(\tau)=\sum_{h=1}^{H}\phi(s_h,a_h). Fit θMLE\theta_{\mathrm{MLE}} by maximizing ∑(τw,τl)∈Dlog⁡σ(θ⊤[ψ(τw)−ψ(τl)])\sum_{(\tau^w,\tau^l)\in D}\log\sigma(\theta^\top[\psi(\tau^w)-\psi(\tau^l)]) over ∥θ∥2≤B\|\theta\|_2\le B, where DD contains preferred/rejected trajectory pairs. Form

    ΣD=∑(τw,τl)∈D[ψ(τw)−ψ(τl)][ψ(τw)−ψ(τl)]⊤+λId,r^(s,a)=ϕ(s,a)⊤θMLE−ϱ ∥ϕ(s,a)∥ΣD−1,\Sigma_D=\sum_{(\tau^w,\tau^l)\in D}[\psi(\tau^w)-\psi(\tau^l)][\psi(\tau^w)-\psi(\tau^l)]^\top+\lambda I_d, \qquad \widehat r(s,a)=\phi(s,a)^\top\theta_{\mathrm{MLE}}-\varrho\,\|\phi(s,a)\|_{\Sigma_D^{-1}},

    with regularization λ>0\lambda>0, confidence coefficient ϱ\varrho, identity matrix IdI_d, and ∥v∥ΣD−1=v⊤ΣD−1v\|v\|_{\Sigma_D^{-1}}=\sqrt{v^\top\Sigma_D^{-1}v}. RTO outputs the policy maximizing the KL-regularized value under r^\widehat r. The subtraction makes the reward estimate pessimistic in poorly supported feature directions.

  5. Knowl 5 — Finite-sample guarantee for theoretical RTO

    theoretical result

    Under the bounded linear-reward assumptions and offline preference-pair setup described here, theoretical RTO selects the policy π^\widehat\pi optimal for its pessimistic reward estimate. With probability at least 1−δ1-\delta, for β>0\beta>0, λ>0\lambda>0, and confidence radius ϱ=O~(d)\varrho=\widetilde O(\sqrt d), its suboptimality satisfies

    Vβ∗(ρ)−Vβπ^(ρ)≤2ϱ E(s,a)∼d∗ ⁣[∥ϕ(s,a)∥ΣD−1]−β Es∼d∗ ⁣[KL ⁣(πβ∗(⋅∣s) ∥ π^(⋅∣s))].V^*_{\beta}(\rho)-V^{\widehat\pi}_{\beta}(\rho) \le 2\varrho\,\mathbb{E}_{(s,a)\sim d^*}\!\left[\|\phi(s,a)\|_{\Sigma_D^{-1}}\right] -\beta\,\mathbb{E}_{s\sim d^*}\!\left[\mathrm{KL}\!\left(\pi^*_{\beta}(\cdot\mid s)\,\|\,\widehat\pi(\cdot\mid s)\right)\right].

    Here Vβ∗V^*_{\beta} is the optimal KL-regularized value for the true reward, πβ∗\pi^*_{\beta} its optimal policy, and d∗d^* the cumulative state-action (or state) visitation measure of that policy from initial prompts drawn from ρ\rho. The data-dependent first term measures coverage of the optimal policy's visits by the offline comparisons; under the paper's stated mild partial-coverage conditions, it typically decreases at rate ∣D∣−1/2|D|^{-1/2} as the number of preference pairs grows.

  6. Knowl 6 — DPO supplies an implicit token-wise reward

    equation

    For an autoregressive policy, the sequence log-probability ratio used by Direct Preference Optimization decomposes into token-level log-probability ratios. Under the paper's deterministic-transition MDP formulation, the token reward associated with an optimal KL-regularized policy is

    r∗(sh,ah)=βlog⁡πβ∗(ah∣sh)πref(ah∣sh)=βlog⁡πβ∗(yh∣x,y1:h−1)πref(yh∣x,y1:h−1).r^*(s_h,a_h)=\beta\log\frac{\pi^*_{\beta}(a_h\mid s_h)}{\pi_{\mathrm{ref}}(a_h\mid s_h)} =\beta\log\frac{\pi^*_{\beta}(y_h\mid x,y_{1:h-1})}{\pi_{\mathrm{ref}}(y_h\mid x,y_{1:h-1})}.

    Here xx is the prompt, y1:h−1y_{1:h-1} is the generated prefix, yhy_h is the current token, and β\beta is the KL coefficient. Summing these token rewards gives the response-level implicit reward used in DPO; for preference comparisons with a shared prompt, the prompt-dependent value term cancels. Practical RTO substitutes the DPO-trained policy πdpo\pi_{\mathrm{dpo}} for the unknown πβ∗\pi^*_{\beta}. The token signal is therefore derived from a model trained on sequence preferences, rather than from separately collected human token labels.

  7. Knowl 7 — Practical RTO uses DPO rewards in PPO updates

    algorithm

    Practical RTO trains a DPO policy on offline preference pairs, keeps the SFT policy πref\pi_{\mathrm{ref}} as the reference, and uses the DPO policy to provide a dense reward while PPO updates the actor. For a sampled response y1:Hy_{1:H} from prompt xx, current policy π\pi, DPO policy πdpo\pi_{\mathrm{dpo}}, and sentence-level reward model rMLEr_{\mathrm{MLE}}, the reward at token hh is

    rRTO(x,y1:h)=β1log⁡πdpo(yh∣x,y1:h−1)πref(yh∣x,y1:h−1)−β2log⁡π(yh∣x,y1:h−1)πref(yh∣x,y1:h−1)+1{h=H}β3rMLE(x,y1:H).r_{\mathrm{RTO}}(x,y_{1:h})=\beta_1\log\frac{\pi_{\mathrm{dpo}}(y_h\mid x,y_{1:h-1})}{\pi_{\mathrm{ref}}(y_h\mid x,y_{1:h-1})}-\beta_2\log\frac{\pi(y_h\mid x,y_{1:h-1})}{\pi_{\mathrm{ref}}(y_h\mid x,y_{1:h-1})}+\mathbf{1}\{h=H\}\beta_3r_{\mathrm{MLE}}(x,y_{1:H}).

    The first term rewards DPO-preferred token choices; the second regularizes the current policy against the reference; the optional terminal sentence reward helps prevent excessively short or long outputs. The training procedure is: (1) fit πdpo\pi_{\mathrm{dpo}} on the preference data and initialize π0=πref\pi_0=\pi_{\mathrm{ref}}; (2) at each iteration, sample a prompt batch from the offline data, generate responses from the current policy, and compute the reward for every generated token; (3) perform PPO updates on those prompt-response trajectories; (4) return the final policy after the chosen training budget. In the Llama-3-8B experiments, the main settings were β1=0.05\beta_1=0.05, β2=0.01\beta_2=0.01, and β3=1\beta_3=1, with actor and critic learning rates 5×10−75\times10^{-7} and 9×10−69\times10^{-6}, batch size 128, eight PPO update steps, maximum prompt and response lengths of 1024 tokens, PPO clip coefficient 0.2, and GAE λ=0.95\lambda=0.95. The sentence-level reward model was 1B parameters; the actor was trained for one epoch.

  8. Knowl 8 — RTO outperforms PPO and preference-learning baselines

    empirical result

    The benchmark comparison on page 8 evaluates Llama-3-8B models initialized from the same SFT model and trained using binarized UltraFeedback preferences. AlpacaEval 2 reports length-controlled (LC) and standard win rate (WR); Arena-Hard reports style-controlled (SC) and standard win rate. The controlled metrics are intended to reduce verbosity bias. The reported scores are:

    Metric SFT DPO R-DPO SimPO TDPO PPO RTO
    AE (LC) 13.22 17.40 18.34 25.46 20.13 19.47 27.00
    AE (WR) 8.58 12.23 12.03 20.20 11.97 12.89 22.45
    AH (SC) 9.2 13.2 14.2 14.5 13.2 16.2 20.3
    AH (WR) 8.9 13.8 14.1 15.2 12.3 15.6 21.4

    RTO leads the listed methods on all four metrics. Compared with PPO, it gains 7.53 points on AlpacaEval 2 LC and 4.1 points on Arena-Hard SC; it also exceeds the best direct-preference baseline on those controlled metrics.

  9. Knowl 9 — Ablations isolate dense reward assignment and DPO reward shaping

    empirical result

    In the reward-granularity comparison, RTO assigns token rewards at each token, Semi-RTO moves each response's token rewards to its delimiter, and DPPO delays the reward to the end-of-sentence token. RS-PPO uses the RTO reward for shaping but removes the DPO implicit reward from the final token, making the total episode reward equal to the sentence-level MLE reward. Results on the same Llama-3-8B/UltraFeedback benchmarks are:

    Metric RTO Semi-RTO DPPO RS-PPO
    AE (LC) 27.00 23.77 21.09 27.52
    AE (WR) 22.45 19.17 13.06 21.69
    AH (SC) 20.3 19.0 13.1 19.2
    AH (WR) 21.4 19.7 12.1 19.9

    Token-level RTO exceeds both coarser reward assignments on every reported metric, and DPPO is lowest on each. RS-PPO improves on DPPO despite having the same total sentence reward, supporting the authors' conclusion that DPO's token signal contributes substantially through reward shaping, not only through changing the total reward. The reward-training curves plotted on page 9 likewise show higher learned-reward trajectories for denser reward assignment and for DPO shaping than for delayed-reward DPPO.

  10. Knowl 10 — RTO reaches PPO-level performance with less training data

    empirical result

    The data-fraction comparison plots AlpacaEval 2 performance for PPO and RTO as the amount of UltraFeedback training data increases. The curve on page 9 shows that RTO reaches approximately PPO's performance using about one-eighth of the data. With additional data, RTO continues to improve and ultimately surpasses PPO, while PPO's performance saturates earlier. This empirical scaling behavior is consistent with the paper's motivation that token-wise feedback provides more usable learning signal than a sentence-level reward.

Coverage note — The supplementary transfer experiments using REINFORCE++ and the TL;DR summarization task are omitted because they provide additional validation beyond the central MDP formulation, RTO method, theory, and primary benchmark evidence.

References

  1. 1.Agarwal, A., Kakade, S. M., Lee, J. D., and Mahajan, G. On the theory of policy gradient methods: Optimality, approximation, and distribution shift. The Journal of Machine Learning Research, 22(1):4431–4506, 2021.
  2. 2.Ahmadian, A., Cremer, C., Galle, M., Fadaee, M., Kreutzer, J., Üstün, A., and Hooker, S. Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms. arXiv preprint arXiv:2402.14740, 2024.
  3. 3.Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Pieter Abbeel, O., and Zaremba, W. Hindsight experience replay. Advances in neural information processing systems, 30, 2017.
  4. 4.Anthropic. Introducing claude. 2023. URL https://www.anthropic.com/index/introducing-claude.
  5. 5.Azar, M. G., Rowland, M., Piot, B., Guo, D., Calandriello, D., Valko, M., and Munos, R. A general theoretical paradigm to understand learning from human preferences. arXiv preprint arXiv:2310.12036, 2023.
  6. 6.Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., Das-Sarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., et al. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862, 2022.
  7. 7.Bengs, V., Busa-Fekete, R., El Mesaoudi-Paul, A., and Hüllermeier, E. Preference-based online learning with dueling bandits: A survey. The Journal of Machine Learning Research, 22(1):278–385, 2021.
  8. 8.Biderman, S., Schoelkopf, H., Anthony, Q. G., Bradley, H., O’Brien, K., Hallahan, E., Khan, M. A., Purohit, S., Prashanth, U. S., Raff, E., et al. Pythia: A suite for analyzing large language models across training and scaling. In International Conference on Machine Learning, pp. 2397–2430. PMLR, 2023.
  9. 9.Bradley, R. A. and Terry, M. E. Rank analysis of incomplete block designs: I. the method of paired comparisons. Biometrika, 39(3/4):324–345, 1952.
  10. 10.Cai, Q., Yang, Z., Jin, C., and Wang, Z. Provably efficient exploration in policy optimization. In International Conference on Machine Learning, pp. 1283–1294. PMLR, 2020.
  11. 11.Casper, S., Davies, X., Shi, C., Gilbert, T. K., Scheurer, J., Rando, J., Freedman, R., Korbak, T., Lindner, D., Freire, P., et al. Open problems and fundamental limitations of reinforcement learning from human feedback. arXiv preprint arXiv:2307.15217, 2023.
  12. 12.Cen, S., Cheng, C., Chen, Y., Wei, Y., and Chi, Y. Fast global convergence of natural policy gradient methods with entropy regularization. Operations Research, 70(4):2563–2578, 2022.
  13. 13.Chan, A. J., Sun, H., Holt, S., and van der Schaar, M. Dense reward for free in reinforcement learning from human feedback. arXiv preprint arXiv:2402.00782, 2024.
  14. 14.Chen, L., Lu, K., Rajeswaran, A., Lee, K., Grover, A., Laskin, M., Abbeel, P., Srinivas, A., and Mordatch, I. Decision transformer: Reinforcement learning via sequence modeling. Advances in neural information processing systems, 34:15084–15097, 2021.
  15. 15.Chen, X., Zhong, H., Yang, Z., Wang, Z., and Wang, L. Human-in-the-loop: Provably efficient preference-based reinforcement learning with general function approximation. In International Conference on Machine Learning, pp. 3773–3793. PMLR, 2022.
  16. 16.Choshen, L., Fox, L., Aizenbud, Z., and Abend, O. On the weaknesses of reinforcement learning for neural machine translation. arXiv preprint arXiv:1907.01752, 2019.
  17. 17.Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D. Deep reinforcement learning from human preferences. Advances in neural information processing systems, 30, 2017.
  18. 18.Coste, T., Anwar, U., Kirk, R., and Krueger, D. Reward model ensembles help mitigate overoptimization. arXiv preprint arXiv:2310.02743, 2023.
  19. 19.Cui, G., Yuan, L., Ding, N., Yao, G., Zhu, W., Ni, Y., Xie, G., Liu, Z., and Sun, M. Ultrafeedback: Boosting language models with high-quality feedback, 2023.
  20. 20.Cui, G., Yuan, L., Wang, Z., Wang, H., Li, W., He, B., Fan, Y., Yu, T., Xu, Q., Chen, W., Yuan, J., Chen, H., Zhang, K., Lv, X., Wang, S., Yao, Y., Peng, H., Cheng, Y., Liu, Z., Sun, M., Zhou, B., and Ding, N. Process reinforcement through implicit rewards, 2025.
  21. 21.Dong, H., Xiong, W., Goyal, D., Zhang, Y., Chow, W., Pan, R., Diao, S., Zhang, J., SHUM, K., and Zhang, T. RAFT: Reward ranked finetuning for generative foundation model alignment. Transactions on Machine Learning Research, 2023. ISSN 2835-8856.
  22. 22.Dong, H., Xiong, W., Pang, B., Wang, H., Zhao, H., Zhou, Y., Jiang, N., Sahoo, D., Xiong, C., and Zhang, T. Rlhf workflow: From reward modeling to online rlhf, 2024.
  23. 23.Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783, 2024.
  24. 24.Faury, L., Abeille, M., Calauzenes, C., and Fercoq, O. Improved optimistic algorithms for logistic bandits. In International Conference on Machine Learning, pp. 3052–3060. PMLR, 2020.
  25. 25.Gulcehre, C., Paine, T. L., Srinivasan, S., Konyushkova, K., Weerts, L., Sharma, A., Siddhant, A., Ahern, A., Wang, M., Gu, C., et al. Reinforced self-training (rest) for language modeling. arXiv preprint arXiv:2308.08998, 2023.
  26. 26.Hoang Tran, Chris Glaze, B. H. Snorkel-mistral-pairrm-dpo. 2024. URL https://huggingface.co/snorkelai/Snorkel-Mistral-PairRM-DPO.
  27. 27.Hu, J. Reinforce++: A simple and efficient approach for aligning large language models. arXiv preprint arXiv:2501.03262, 2025.
  28. 28.Hu, J., Tao, L., Yang, J., and Zhou, C. Aligning language models with offline reinforcement learning from human feedback. arXiv preprint arXiv:2308.12050, 2023.
  29. 29.Hu, J., Wu, X., Zhu, Z., Xianyu, Wang, W., Zhang, D., and Cao, Y. Openrlhf: An easy-to-use, scalable and high-performance rlhf framework. arXiv preprint arXiv:2405.11143, 2024.
  30. 30.Huang, J., Yardim, B., and He, N. On the statistical efficiency of mean-field reinforcement learning with general function approximation. In International Conference on Artificial Intelligence and Statistics, pp. 289–297. PMLR, 2024.
  31. 31.Jang, J., Kim, S., Lin, B. Y., Wang, Y., Hessel, J., Zettlemoyer, L., Hajishirzi, H., Choi, Y., and Ammanabrolu, P. Personalized soups: Personalized large language model alignment via post-hoc parameter merging. arXiv preprint arXiv:2310.11564, 2023.
  32. 32.Jin, Y., Yang, Z., and Wang, Z. Is pessimism provably efficient for offline rl? In International Conference on Machine Learning, pp. 5084–5096. PMLR, 2021.
  33. 33.Kakade, S. and Langford, J. Approximately optimal approximate reinforcement learning. In Proceedings of the Nineteenth International Conference on Machine Learning, pp. 267–274, 2002.
  34. 34.Kingma, D. P. and Ba, J. Adam: A method for stochastic optimization, 2017. URL https://arxiv.org/abs/1412.6980.
  35. 35.Lambert, N., Pyatkin, V., Morrison, J., Miranda, L., Lin, B. Y., Chandu, K., Dziri, N., Kumar, S., Zick, T., Choi, Y., et al. Rewardbench: Evaluating reward models for language modeling. arXiv preprint arXiv:2403.13787, 2024.
  36. 36.Li, T., Chiang, W.-L., Frick, E., Dunlap, L., Wu, T., Zhu, B., Gonzalez, J. E., and Stoica, I. From crowdsourced data to high-quality benchmarks: Arena-hard and benchbuilder pipeline. arXiv preprint arXiv:2406.11939, 2024.
  37. 37.Li, X., Zhang, T., Dubois, Y., Taori, R., Gulrajani, I., Guestrin, C., Liang, P., and Hashimoto, T. B. Alpacaeval: An automatic evaluator of instruction-following models. https://github.com/tatsu-lab/alpaca_eval, 5 2023a.
  38. 38.Li, Z., Xu, T., Zhang, Y., Yu, Y., Sun, R., and Luo, Z.-Q. Remax: A simple, effective, and efficient reinforcement learning method for aligning large language models. arXiv e-prints, pp. arXiv–2310, 2023b.
  39. 39.Li, Z., Yang, Z., and Wang, M. Reinforcement learning with human feedback: Learning dynamic choices via pessimism. arXiv preprint arXiv:2305.18438, 2023c.
  40. 40.Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., and Cobbe, K. Let’s verify step by step. arXiv preprint arXiv:2305.20050, 2023.
  41. 41.Liu, Q., Chung, A., Szepesvari, C., and Jin, C. When is partially observable reinforcement learning not scary? In Conference on Learning Theory, pp. 5175–5220. PMLR, 2022.
  42. 42.Liu, Z., Lu, M., Xiong, W., Zhong, H., Hu, H., Zhang, S., Zheng, S., Yang, Z., and Wang, Z. Maximize to explore: One objective function fusing estimation, planning, and exploration. In Thirty-seventh Conference on Neural Information Processing Systems, 2023.
  43. 43.Meng, Y., Xia, M., and Chen, D. Simpo: Simple preference optimization with a reference-free reward. arXiv preprint arXiv:2405.14734, 2024.
  44. 44.Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., et al. Webgpt: Browser-assisted question-answering with human feedback. arXiv preprint arXiv:2112.09332, 2021.
  45. 45.OpenAI. Gpt-4 technical report. ArXiv, abs/2303.08774, 2023.
  46. 46.Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744, 2022.
  47. 47.Pacchiano, A., Saha, A., and Lee, J. Dueling rl: reinforcement learning with trajectory preferences. arXiv preprint arXiv:2111.04850, 2021.
  48. 48.Park, R., Rafailov, R., Ermon, S., and Finn, C. Disentangling length from quality in direct preference optimization. arXiv preprint arXiv:2403.19159, 2024.
  49. 49.Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C. Direct preference optimization: Your language model is secretly a reward model. arXiv preprint arXiv:2305.18290, 2023.
  50. 50.Rafailov, R., Hejna, J., Park, R., and Finn, C. From r to q∗: Your language model is secretly a q-function. arXiv preprint arXiv:2404.12358, 2024.
  51. 51.Rashidinejad, P., Zhu, B., Ma, C., Jiao, J., and Russell, S. Bridging offline reinforcement learning and imitation learning: A tale of pessimism. Advances in Neural Information Processing Systems, 34:11702–11716, 2021.
  52. 52.Rosset, C., Cheng, C.-A., Mitra, A., Santacroce, M., Awadallah, A., and Xie, T. Direct nash optimization: Teaching language models to self-improve with general preferences. arXiv preprint arXiv:2404.03715, 2024.
  53. 53.Saha, A. Optimal algorithms for stochastic contextual preference bandits. Advances in Neural Information Processing Systems, 34:30050–30062, 2021.
  54. 54.Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017.
  55. 55.Tang, Y., Guo, Z. D., Zheng, Z., Calandriello, D., Munos, R., Rowland, M., Richemond, P. H., Valko, M., Pires, B. A., and Piot, B. Generalized preference optimization: A unified approach to offline alignment. arXiv preprint arXiv:2402.05749, 2024.
  56. 56.Team, G., Anil, R., Borgeaud, S., Wu, Y., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., et al. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805, 2023.
  57. 57.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023.
  58. 58.Uehara, M. and Sun, W. Pessimistic model-based offline reinforcement learning under partial coverage. arXiv preprint arXiv:2107.06226, 2021.
  59. 59.Uesato, J., Kushman, N., Kumar, R., Song, F., Siegel, N., Wang, L., Creswell, A., Irving, G., and Higgins, I. Solving math word problems with process-and outcome-based feedback. arXiv preprint arXiv:2211.14275, 2022.
  60. 60.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  61. 61.Volske, M., Potthast, M., Syed, S., and Stein, B. Tl; dr: Mining reddit to learn automatic summarization. In Proceedings of the Workshop on New Frontiers in Summarization, pp. 59–63, 2017.
  62. 62.Wang, H., Lin, Y., Xiong, W., Yang, R., Diao, S., Qiu, S., Zhao, H., and Zhang, T. Arithmetic control of llms for diverse user preferences: Directional preference alignment with multi-objective rewards. arXiv preprint arXiv:2402.18571, 2024.
  63. 63.Wang, Y., Liu, Q., and Jin, C. Is rlhf more difficult than standard rl? arXiv preprint arXiv:2306.14111, 2023.
  64. 64.Williams, R. J. Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine learning, 8:229–256, 1992.
  65. 65.Williams, R. J. and Peng, J. Function optimization using connectionist reinforcement learning algorithms. Connection Science, 3(3):241–268, 1991.
  66. 66.Wu, T., Yang, Y., Zhong, H., Wang, L., Du, S., and Jiao, J. Nearly optimal policy optimization with stable at any time guarantee. In International Conference on Machine Learning, pp. 24243–24265. PMLR, 2022.
  67. 67.Wu, Z., Hu, Y., Shi, W., Dziri, N., Suhr, A., Ammanabrolu, P., Smith, N. A., Ostendorf, M., and Hajishirzi, H. Fine-grained human feedback gives better rewards for language model training. Advances in Neural Information Processing Systems, 36, 2024.
  68. 68.Xiong, W., Zhong, H., Shi, C., Shen, C., Wang, L., and Zhang, T. Nearly minimax optimal offline reinforcement learning with linear function approximation: Single-agent mdp and markov game. arXiv preprint arXiv:2205.15512, 2022.
  69. 69.Xiong, W., Dong, H., Ye, C., Wang, Z., Zhong, H., Ji, H., Jiang, N., and Zhang, T. Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint. In ICLR 2024 Workshop on Mathematical and Empirical Understanding of Foundation Models, 2023.
  70. 70.Yang, R., Pan, X., Luo, F., Qiu, S., Zhong, H., Yu, D., and Chen, J. Rewards-in-context: Multi-objective alignment of foundation models with dynamic preference adjustment. arXiv preprint arXiv:2402.10207, 2024a.
  71. 71.Yang, S., Zhang, S., Xia, C., Feng, Y., Xiong, C., and Zhou, M. Preference-grounded token-level guidance for language model fine-tuning. Advances in Neural Information Processing Systems, 36, 2024b.
  72. 72.Ye, C., Xiong, W., Zhang, Y., Jiang, N., and Zhang, T. A theoretical analysis of nash learning from human feedback under general kl-regularized preference. arXiv preprint arXiv:2402.07314, 2024.
  73. 73.Yin, Y., Yang, S., Xie, Y., Yang, Z., Sun, Y., Awadalla, H., Chen, W., and Zhou, M. Segmenting text and learning their rewards for improved rlhf in language model. arXiv preprint arXiv:2501.02790, 2025.
  74. 74.Yue, Y., Broder, J., Kleinberg, R., and Joachims, T. The k-armed dueling bandits problem. Journal of Computer and System Sciences, 78(5):1538–1556, 2012.
  75. 75.Zeng, Y., Liu, G., Ma, W., Yang, N., Zhang, H., and Wang, J. Token-level direct preference optimization. arXiv preprint arXiv:2404.11999, 2024.
  76. 76.Zhan, W., Uehara, M., Kallus, N., Lee, J. D., and Sun, W. Provable offline reinforcement learning with human feedback. arXiv preprint arXiv:2305.14816, 2023a.
  77. 77.Zhan, W., Uehara, M., Sun, W., and Lee, J. D. How to query human feedback efficiently in rl? arXiv preprint arXiv:2305.18505, 2023b.
  78. 78.Zhang, T. Mathematical analysis of machine learning algorithms. Cambridge University Press, 2023.
  79. 79.Zhao, Y., Joshi, R., Liu, T., Khalman, M., Saleh, M., and Liu, P. J. Slic-hf: Sequence likelihood calibration with human feedback. arXiv preprint arXiv:2305.10425, 2023.
  80. 80.Zhong, H. and Zhang, T. A theoretical analysis of optimistic proximal policy optimization in linear markov decision processes. Advances in Neural Information Processing Systems, 36, 2024.
  81. 81.Zhong, H., Xiong, W., Zheng, S., Wang, L., Wang, Z., Yang, Z., and Zhang, T. Gec: A unified framework for interactive decision making in mdp, pomdp, and beyond. arXiv preprint arXiv:2211.01962, 2022.
  82. 82.Zhu, B., Jiao, J., and Jordan, M. I. Principled reinforcement learning with human feedback from pairwise or k-wise comparisons. arXiv preprint arXiv:2301.11270, 2023.
  83. 83.Ziebart, B. D. Modeling purposeful adaptive behavior with the principle of maximum causal entropy. Carnegie Mellon University, 2010.
  84. 84.Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G. Fine-tuning language models from human preferences. arXiv preprint arXiv:1909.08593, 2019.

Citation

MLA
Zhong, H., et al. “DPO Meets PPO: Reinforced Token Optimization for RLHF”. arXiv, 2024, http://arxiv.org/abs/2404.18922v4.
APA
Zhong, H., Shan, Z., Feng, G., Xiong, W., Cheng, X., Zhao, L., He, D., Bian, J., & Wang, L. (2024). DPO Meets PPO: Reinforced Token Optimization for RLHF. arXiv. http://arxiv.org/abs/2404.18922v4
Chicago
Zhong, H., Z. Shan, G. Feng, et al. 2024. “DPO Meets PPO: Reinforced Token Optimization for RLHF”. arXiv. http://arxiv.org/abs/2404.18922v4.
Harvard
Zhong, H. et al. (2024) “DPO Meets PPO: Reinforced Token Optimization for RLHF”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2404.18922v4.
Vancouver
1. Zhong H, Shan Z, Feng G, Xiong W, Cheng X, Zhao L, He D, Bian J, Wang L (2024) DPO Meets PPO: Reinforced Token Optimization for RLHF. arXiv

BibTeX

@article{zhong2024dpo,
  title = {DPO Meets PPO: Reinforced Token Optimization for RLHF},
  author = {Zhong, Han and Shan, Zikang and Feng, Guhao and Xiong, Wei and Cheng, Xinle and Zhao, Li and He, Di and Bian, Jiang and Wang, Liwei},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2404.18922v4},
  eprint = {2404.18922}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/