Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study

Shusheng XuWei FuJiaxuan GaoWenjie YeWeilin LiuZhiyu MeiGuangju WangChao YuYi Wu

article2024ICML271 citations

Demonstrates that properly configured Proximal Policy Optimization consistently surpasses Direct Preference Optimization across diverse alignment benchmarks, identifying fundamental limitations in reward-free methods and isolating the critical training factors needed to achieve superior performance in complex tasks like code generation.

Listen

Aligning large language models with human expectations is critical for ensuring reliable, safe, and effective artificial intelligence systems. Organizations currently face a major strategic dilemma between two competing alignment techniques: reward-based methods such as Proximal Policy Optimization, which are used by leading industry applications like ChatGPT and Claude, and newer reward-free methods like Direct Preference Optimization, which have gained rapid popularity in open-source research due to their simplicity. This tension creates uncertainty for engineering teams deciding how to invest computing resources and design model post-training pipelines.

The article investigates whether Direct Preference Optimization is fundamentally superior to Proximal Policy Optimization and determines the specific training practices needed to achieve optimal performance. To answer these questions, the authors analyze the theoretical properties of Direct Preference Optimization, conduct controlled synthetic experiments, and perform extensive empirical evaluations across dialogue generation (using the HH-RLHF and SafeRLHF datasets) and complex programming benchmarks (using APPS and CodeContest). The study evaluates language models ranging from 7 billion to 34 billion parameters across diverse task difficulties and reward mechanisms.

The findings show that properly tuned Proximal Policy Optimization consistently outperforms Direct Preference Optimization across all evaluated domains. First, theoretical and empirical analyses reveal that Direct Preference Optimization is vulnerable to distribution shifts between model outputs and static preference data, often assigning high probabilities to out-of-distribution responses and producing erratic behaviors. Second, the authors identify three key operational practices that dramatically enhance Proximal Policy Optimization: normalizing policy advantages, utilizing large batch sizes, and updating the baseline reference model with an exponential moving average. Third, while iterative variants of Direct Preference Optimization mitigate some distribution mismatch in dialogue tasks, they fail on complex reasoning tasks; in competitive programming evaluations, a 34-billion-parameter model aligned with Proximal Policy Optimization achieved state-of-the-art results (increasing the CodeContest pass rate from 16.4% to 22.4%), whereas Direct Preference Optimization degraded to near-zero pass rates.

These results demonstrate that organizations pursuing advanced model alignment should not abandon reinforcement learning frameworks in favor of simpler reward-free short-cuts. While Direct Preference Optimization offers lower initial engineering overhead, it presents significant risks in complex reasoning domains where distribution shifts cause severe performance degradation. Proximal Policy Optimization remains the more robust and capable alignment framework when configured with adequate batch sizes and proper reference model updating.

For practical implementation, technical teams should maintain reinforcement learning pipelines utilizing Proximal Policy Optimization for high-stakes and complex reasoning tasks. When employing Direct Preference Optimization for simpler conversational applications, teams should adopt iterative retraining on newly generated model outputs rather than relying on static datasets. Decision-makers should note that these conclusions assume access to reliable reward models or direct execution feedback; future work remains necessary to establish best practices for training robust reward models across non-deterministic domains.

Cover for Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study

Abstract

Reinforcement Learning from Human Feedback (RLHF) is currently the most widely used method to align large language models (LLMs) with human preferences. Existing RLHF methods can be roughly categorized as either reward-based or reward-free. Novel applications such as ChatGPT and Claude leverage reward-based methods that first learn a reward model and apply actor-critic algorithms, such as Proximal Policy Optimization (PPO). However, in academic benchmarks, state-of-the-art results are often achieved via reward-free methods, such as Direct Preference Optimization (DPO). Is DPO truly superior to PPO? Why does PPO perform poorly on these benchmarks? In this paper, we first conduct both theoretical and empirical studies on the algorithmic properties of DPO and show that DPO may have fundamental limitations. Moreover, we also comprehensively examine PPO and reveal the key factors for the best performances of PPO in fine-tuning LLMs. Finally, we benchmark DPO and PPO across a collection of RLHF testbeds, ranging from dialogue to code generation. Experiment results demonstrate that PPO is able to surpass other alignment methods in all cases and achieve state-of-the-art results in challenging code competitions. Our code is publicly available at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Preliminary
  • 4 Understanding the Limitation of DPO
  • 4.1 Theoretical Analysis
  • 4.2 Empirical Validation in A Synthetic Scenario
  • 4.3 Experiments on Real Preference Datasets
  • 5 Key Factors to PPO for RLHF
  • 6 Benchmark Results
  • 7 Conclusion
  • References
  • A Implementation Details
  • A.1 DPO Details
  • A.2 PPO Details
  • B GPT-4 Evaluation
  • C Additional Experiments
  • C.1 Varying the Reference Model
  • C.2 Varying β\beta
  • C.3 Varying Preference Dataset
  • C.4 Human Evaluation

Knowls

  1. Knowl 1 — Theoretical Containment of PPO in DPO Policy Classes

    theoretical result

    For a ground-truth reward function rr and a preference dataset D={(x,yw,yl)}\mathcal{D} = \{(x, y_w, y_l)\}, let ΠPPO\Pi_{\mathrm{PPO}} be the class of policies induced by training a Bradley-Terry reward model rϕr_\phi on D\mathcal{D} and optimizing the standard KL-regularized reinforcement learning objective:

    Jrϕ(πθ)=Ex∼pdata,y∼πθ[rϕ(x,y)−βlog⁡πθ(y∣x)πref(y∣x)]J_{r_\phi}(\pi_\theta) = \mathbb{E}_{x \sim p_{\mathrm{data}}, y \sim \pi_\theta} \left[ r_\phi(x, y) - \beta \log \frac{\pi_\theta(y \mid x)}{\pi_{\mathrm{ref}}(y \mid x)} \right]

    Let ΠDPO\Pi_{\mathrm{DPO}} be the class of policies induced by minimizing the Direct Preference Optimization (DPO) objective:

    LDPO(πθ)=−E(x,yw,yl)∼D[log⁡σ(βlog⁡πθ(yw∣x)πref(yw∣x)−βlog⁡πθ(yl∣x)πref(yl∣x))]\mathcal{L}_{\mathrm{DPO}}(\pi_\theta) = -\mathbb{E}_{(x, y_w, y_l) \sim \mathcal{D}} \left[ \log \sigma \left( \beta \log \frac{\pi_\theta(y_w \mid x)}{\pi_{\mathrm{ref}}(y_w \mid x)} - \beta \log \frac{\pi_\theta(y_l \mid x)}{\pi_{\mathrm{ref}}(y_l \mid x)} \right) \right]

    where σ\sigma is the sigmoid function, πref\pi_{\mathrm{ref}} is the reference policy, and β>0\beta > 0 controls regularization strength.

    Then, ΠPPO\Pi_{\mathrm{PPO}} is a strict proper subset of ΠDPO\Pi_{\mathrm{DPO}}:

    ΠPPO⊊ΠDPO\Pi_{\mathrm{PPO}} \subsetneq \Pi_{\mathrm{DPO}}

    This result implies that any policy achievable by PPO that exploits errors in a learned reward model also minimizes the DPO objective. Conversely, DPO can achieve global loss minima by assigning arbitrary probability mass to unobserved, out-of-distribution responses (even assigning non-zero mass where πref(y∣x)=0\pi_{\mathrm{ref}}(y \mid x) = 0), whereas PPO enforces πPPO(y∣x)=0\pi_{\mathrm{PPO}}(y \mid x) = 0 whenever πref(y∣x)=0\pi_{\mathrm{ref}}(y \mid x) = 0 due to the structure of the closed-form optimum π∗(y∣x)∝πref(y∣x)exp⁡(r∗(x,y)/β)\pi^*(y \mid x) \propto \pi_{\mathrm{ref}}(y \mid x) \exp(r^*(x, y)/\beta).

  2. Knowl 2 — Critical Algorithmic Techniques for Optimizing PPO in LLM Alignment

    model/method

    To achieve optimal performance and stability when applying Proximal Policy Optimization (PPO) to large language model (LLM) fine-tuning, three specific algorithmic techniques are essential:

    1. Advantage Normalization: Standardizing advantage estimates calculated via Generalized Advantage Estimation (GAE) across rollout batches stabilizes policy gradient updates against high variance in prompt lengths and reward magnitudes.
    2. Large Batch Size Training: Scaling global rollout batch sizes (e.g., collecting 512 rollouts partitioned into mini-batches, rather than using small batch sizes such as 64) prevents policy collapse and severe performance degradation, particularly in challenging multi-step reasoning domains like code generation.
    3. Reference Model Exponential Moving Average (EMA) Updates: Updating the reference model parameters πref\pi_{\mathrm{ref}} over training using an exponential moving average of the actor policy parameters πθ\pi_\theta prevents the policy from being over-regularized against a static initial checkpoint as the policy improves.

    Key hyperparameters for effective LLM-PPO training include decoupled actor and critic networks with learning rates of 1×10−51 \times 10^{-5} for the actor and 5×10−65 \times 10^{-6} for the critic, GAE parameter λ=1.0\lambda = 1.0, discount factor γ=1.0\gamma = 1.0, KL penalty coefficient β=0.1\beta = 0.1, reward score clipping to [−20,20][-20, 20], and token-level scalar reward assignment.

  3. Knowl 3 — Comparative Alignment Performance on Competitive Code Generation

    empirical result

    On competitive programming benchmarks using CodeLlama base models, PPO consistently outperforms Supervised Fine-Tuning (SFT) and Direct Preference Optimization (DPO), achieving state-of-the-art results on APPS and CodeContest:

    • APPS Benchmark (pass@5 metric):

      • CodeLlama-7B: SFT scores 30.0% (Introductory), 7.8% (Interview), 2.8% (Competition); Iterative DPO (DPO-Iter) degrades to 20.9%, 3.4%, 1.3%; PPO scores 29.4%, 7.6%, 2.4%.
      • CodeLlama-13B: SFT scores 33.7%, 8.7%, 3.6%; DPO-Iter scores 33.0%, 8.0%, 2.8%; PPO improves to 36.4%, 11.47%, 4.6%.
      • CodeLlama-34B: SFT scores 38.6%, 10.1%, 3.9%; DPO-Iter drops to 34.2%, 9.3%, 3.7%; PPO achieves 44.4%, 18.0%, 9.1%.
    • CodeContest Benchmark (10@1k metric: 10 selections evaluated on hidden tests from 1,000 public test candidates):

      • AlphaCode-41B (with clustering) achieves 21.0% validation and 16.4% test.
      • CodeLlama-34B SFT achieves 10.3% validation and 15.2% test.
      • Standard single-round DPO fails completely, collapsing after 1 epoch to 0.0% pass rate by generating degenerate syntax snippets.
      • DPO-Iter achieves 3.5% validation and 3.2% test.
      • CodeLlama-34B fine-tuned via PPO reaches 19.7% validation and 22.4% test (using Python only), outperforming AlphaCode-41B.
  4. Knowl 4 — Alignment Performance and Preferences on Dialogue Benchmarks

    data/table

    On general dialogue (HH-RLHF) and safety alignment (SafeRLHF), PPO achieves higher rewards and win rates against chosen dataset responses, SFT baselines, and other alignment methods (RRHF, PRO, DPO, and DPO-Iter).

    Method OpenAssistant Reward Win vs Chosen (%) Win vs SFT (%)
    RRHF 0.523 28 29
    PRO 0.529 37 34
    DPO 0.611 55 53
    DPO-Iter 0.678 55 54
    PPO 0.718 57 58

    On HH-RLHF (evaluated with Llama-2-7B), pairwise GPT-4 evaluation shows PPO defeats standard DPO with 42% wins to 30% wins (28% ties), and defeats DPO-Iter with 36% wins to 28% wins (36% ties). Human evaluations show 60% to 61% agreement with GPT-4 and favor PPO over DPO (45% vs 29%) and over DPO-Iter (38% vs 33%).

    On SafeRLHF evaluated via official Beaver reward and cost models (helpfulness relative to Beaver ΔHelp\Delta\text{Help}, harmfulness Harm\text{Harm}, and safety rate S.R.\text{S.R.}):

    • Llama-1-7B: SFT (ΔHelp=−2.26,Harm=0.78,S.R.=46.5%\Delta\text{Help} = -2.26, \text{Harm} = 0.78, \text{S.R.} = 46.5\%); DPO (−2.70,−6.38,93.1%-2.70, -6.38, 93.1\%); DPO-Iter (−2.79,−11.86,100.0%-2.79, -11.86, 100.0\%); PPO (+0.66,−10.22,98.6%+0.66, -10.22, 98.6\%).
    • Llama-2-7B: SFT (−2.12,0.00,52.1%-2.12, 0.00, 52.1\%); DPO (−2.86,−6.82,95.8%-2.86, -6.82, 95.8\%); DPO-Iter (−2.96,−11.07,99.9%-2.96, -11.07, 99.9\%); PPO (+1.69,−12.08,99.5%+1.69, -12.08, 99.5\%).
  5. Knowl 5 — Ablation of PPO Training Components on Dialogue and Code Benchmarks

    data/table

    An ablation study isolating the effects of Advantage Normalization (Adv.Norm.), Large Batch Size (Large.Batch., size 512 vs baseline 64), and Exponential Moving Average updates for the Reference Model (Ref.EMA) demonstrates cumulative gains across HH-RLHF (Llama-2-7B), APPS (CodeLlama-34B), and CodeContest (CodeLlama-34B).

    Setting HH-RLHF APPS (pass@5) CodeContest
    Reward Intro Inter Comp pass@10 pass@1k
    SFT 0.532 38.6% 10.1% 3.9% 0.9% 12.0%
    Baseline PPO (bs=64) 0.706 18.0% 2.4% 1.1% 4.3% 7.7%
    + Adv.Norm. 0.716 38.1% 11.4% 4.6% 6.8% 15.4%
    + Large.Batch. (bs=512) 0.716 42.3% 14.6% 7.5% 5.1% 19.6%
    + Ref.EMA 0.718 44.4% 18.0% 9.1% 6.8% 21.4%

    Without advantage normalization and large batch size, baseline PPO exhibits severe degradation on code generation relative to SFT (e.g., dropping from 38.6% to 18.0% on APPS Introductory). Large batch training provides the largest single gain for code synthesis, and reference model EMA prevents the policy from being prematurely over-constrained by the initial SFT anchor.

  6. Knowl 6 — Vulnerability of DPO to Distribution Shift and Mitigation via Iterative DPO

    empirical result

    Direct Preference Optimization (DPO) is sensitive to distribution mismatches between the base/reference model and the preference dataset:

    1. Base Model Mismatch: When fine-tuning Llama-2-7B on SafeRLHF using an Alpaca SFT base model (SFT (Alpaca)), DPO achieves only a 55.4% safety rate and relative helpfulness of −4.19-4.19. Aligning the base model with safe demonstrations prior to DPO (SFT (Safe)) increases the safety rate by 16.4% (to 71.8%) and increases helpfulness to −1.62-1.62. On APPS, using Codellama-13B-Pretrain as the reference for DPO yields 0.24% pass@5, whereas using Codellama-13B-SFT yields 12.8% pass@5.
    2. Iterative Preference Training (DPO-Iter): Sampling responses dynamically from the latest policy, annotating them via a learned reward model or unit-test verifier, and updating the reference model iteratively over four successive rounds (Iter. 1 to Iter. 4) increases SafeRLHF safety rate from 86.7% to 99.9% and lowers harmfulness from −5.23-5.23 to −11.07-11.07. However, DPO-Iter helpfulness remains significantly lower than PPO (−2.96-2.96 vs +1.69+1.69).
  7. Knowl 7 — Direct Preference Optimization Formulation and Implicit Reward Reparameterization

    equation

    Direct Preference Optimization (DPO) eliminates explicit reward modeling by expressing the reward function rϕ(x,y)r_\phi(x, y) directly in terms of the language model policy πθ\pi_\theta, the reference policy πref\pi_{\mathrm{ref}}, and a prompt-dependent scalar function C(x)C(x):

    rϕ(x,y)=βlog⁡πθ(y∣x)πref(y∣x)+C(x)r_\phi(x, y) = \beta \log \frac{\pi_\theta(y \mid x)}{\pi_{\mathrm{ref}}(y \mid x)} + C(x)

    where β>0\beta > 0 is the KL regularization coefficient. Substituting this reparameterization into the Bradley-Terry preference negative log-likelihood over preference pairs (x,yw,yl)∼D(x, y_w, y_l) \sim \mathcal{D} (where ywy_w is preferred to yly_l) gives the closed-form DPO objective:

    LDPO(πθ)=−E(x,yw,yl)∼D[log⁡σ(βlog⁡πθ(yw∣x)πref(yw∣x)−βlog⁡πθ(yl∣x)πref(yl∣x))]\mathcal{L}_{\mathrm{DPO}}(\pi_\theta) = -\mathbb{E}_{(x, y_w, y_l) \sim \mathcal{D}} \left[ \log \sigma \left( \beta \log \frac{\pi_\theta(y_w \mid x)}{\pi_{\mathrm{ref}}(y_w \mid x)} - \beta \log \frac{\pi_\theta(y_l \mid x)}{\pi_{\mathrm{ref}}(y_l \mid x)} \right) \right]

    where σ(u)=(1+e−u)−1\sigma(u) = (1 + e^{-u})^{-1} denotes the standard sigmoid function.

  8. Knowl 8 — Impact of KL Regularization Weight Beta on Alignment Performance

    empirical result

    Evaluating the KL penalty parameter β∈{0,0.05,0.1,0.2}\beta \in \{0, 0.05, 0.1, 0.2\} indicates that an intermediate value of β=0.1\beta = 0.1 consistently produces optimal alignment outcomes across tasks and models:

    • Anthropic HH-RLHF (Llama-7B, OpenAssistant Reward):
      • PPO: β=0→0.705\beta=0 \to 0.705, β=0.05→0.720\beta=0.05 \to 0.720, β=0.1→0.718\beta=0.1 \to 0.718, β=0.2→0.629\beta=0.2 \to 0.629.
      • DPO: β=0.05→0.609\beta=0.05 \to 0.609, β=0.1→0.611\beta=0.1 \to 0.611, β=0.2→0.597\beta=0.2 \to 0.597.
    • APPS Code Generation (CodeLlama-13B, Average pass@5):
      • PPO: β=0→13.0%\beta=0 \to 13.0\%, β=0.05→14.1%\beta=0.05 \to 14.1\%, β=0.1→15.1%\beta=0.1 \to 15.1\%, β=0.2→14.9%\beta=0.2 \to 14.9\%.
      • DPO: β=0.05→12.56%\beta=0.05 \to 12.56\%, β=0.1→12.0%\beta=0.1 \to 12.0\%, β=0.2→12.32%\beta=0.2 \to 12.32\%.

    Setting β\beta too large restricts policy adaptation by over-penalizing divergence from the reference policy, degrading both dialogue and code generation performance.

  9. Knowl 9 — Impact of Preference Data Filtering on Safety-Helpfulness Trade-offs

    empirical result

    On the SafeRLHF dataset (which contains explicit dual-safety and helpfulness annotations), selective filtering of preference pairs impacts safety and helpfulness trade-offs:

    • Filtering Dual-Unsafe Pairs: Removing pairs where both responses are unsafe enables the reward model to distinguish helpfulness more effectively, increasing PPO helpfulness reward from 1.691.69 to 5.885.88 (and DPO helpfulness from −1.62-1.62 to 2.462.46) while preserving high safety rates (92.6% for PPO, 80.8% for DPO).
    • Filtering Dual-Safe and Dual-Unsafe Pairs: Removing pairs where both responses share identical safety labels forces the reward model to focus exclusively on safety, resulting in overly conservative policies that frequently refuse benign prompts. Relative helpfulness drops sharply (to −8.04-8.04 for PPO and −2.86-2.86 for DPO) while maintaining safety (94.9% for PPO, 95.8% for DPO).
    • Across all dataset filtering configurations, PPO maintains safety rates above 92%, showing greater resilience to training data compositions than DPO.
  10. Knowl 10 — Methodological Limitations Regarding Reward Modeling and Ground-Truth Feedback

    limitation

    The study identifies two explicit scope boundaries:

    1. The investigation focuses on policy optimization dynamics given a reward signal or preference pairs, and does not study methods for training robust reward models or mitigating reward model overoptimization / reward hacking.
    2. In competitive code generation tasks (APPS and CodeContest), ground-truth unit tests and execution verifiers are used to provide reward signals for PPO and preference labels for Iterative DPO, eliminating reward model estimation error in the coding experiments.

Coverage note — No substantial contributed material was omitted; all theoretical claims, empirical evaluations across dialogue and code benchmarks, ablations, and hyperparameter analyses are represented.

References

  1. 1.Agrawal, A., Mackey, L., and Kalai, A. T. Do language models know when they're hallucinating references? CoRR, abs/2305.18248, 2023. doi: 10.48550/ARXIV.2305.18248. URL https://doi.org/10.48550/arXiv.2305.18248.
  2. 2.Andrychowicz, M., Raichuk, A., Stanczyk, P., Orsini, M., Girgin, S., Marinier, R., Hussenot, L., Geist, M., Pietquin, O., Michalski, M., Gelly, S., and Bachem, O. What matters for on-policy deep actor-critic methods? A large-scale study. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net, 2021. URL https://openreview.net/forum?id=nIAxjsniDzg.
  3. 3.Anil, R., Dai, A. M., Firat, O., Johnson, M., Lepikhin, D., Passos, A., Shakeri, S., Taropa, E., Bailey, P., Chen, Z., Chu, E., Clark, J. H., Shafey, L. E., Huang, Y., Meier-Hellstern, K., Mishra, G., Moreira, E., Omernick, M., Robinson, K., Ruder, S., Tay, Y., Xiao, K., Xu, Y., Zhang, Y., Abrego, G. H., Ahn, J., Austin, J., Barham, P., Botha, J. A., Bradbury, J., Brahma, S., Brooks, K., Catasta, M., Cheng, Y., Cherry, C., Choquette-Choo, C. A., Chowdhery, A., Crepy, C., Dave, S., Dehghani, M., Dev, S., Devlin, J., Dıaz, M., Du, N., Dyer, E., Feinberg, V., Feng, F., Fienber, V., Freitag, M., Garcia, X., Gehrmann, S., Gonzalez, L., and et al. Palm 2 technical report. CoRR, abs/2305.10403, 2023. doi: 10.48550/ARXIV.2305.10403. URL https://doi.org/10.48550/arXiv.2305.10403.
  4. 4.Antropic. Claude, Jul 2023. URL https://claude.ai/chats.
  5. 5.Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., Das-Sarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., et al. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862, 2022.
  6. 6.Bradley, R. A. and Terry, M. E. Rank analysis of incomplete block designs: I. the method of paired comparisons. Biometrika, 39(3/4):324–345, 1952.
  7. 7.Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. Language models are few-shot learners. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., and Lin, H. (eds.), Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual, 2020. URL https://proceedings.neurips.cc/paper/2020/hash/1457c0d6bfcb4967418bfb8ac142f64a-Abstract.html.
  8. 8.Casper, S., Davies, X., Shi, C., Gilbert, T. K., Scheurer, J., Rando, J., Freedman, R., Korbak, T., Lindner, D., Freire, P., Wang, T., Marks, S., Segerie, C., Carroll, M., Peng, A., Christoffersen, P. J. K., Damani, M., Slocum, S., Anwar, U., Siththaranjan, A., Nadeau, M., Michaud, E. J., Pfau, J., Krasheninnikov, D., Chen, X., Langosco, L., Hase, P., Biyik, E., Dragan, A. D., Krueger, D., Sadigh, D., and Hadfield-Menell, D. Open problems and fundamental limitations of reinforcement learning from human feedback. CoRR, abs/2307.15217, 2023. doi: 10.48550/ARXIV.2307.15217. URL https://doi.org/10.48550/arXiv.2307.15217.
  9. 9.Chen, Z., Deng, Y., Yuan, H., Ji, K., and Gu, Q. Self-play fine-tuning converts weak language models to strong language models. arXiv preprint arXiv:2401.01335, 2024.
  10. 10.Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., Schuh, P., Shi, K., Tsvyashchenko, S., Maynez, J., Rao, A., Barnes, P., Tay, Y., Shazeer, N., Prabhakaran, V., Reif, E., Du, N., Hutchinson, B., Pope, R., Bradbury, J., Austin, J., Isard, M., Gur-Ari, G., Yin, P., Duke, T., Levskaya, A., Ghemawat, S., Dev, S., Michalewski, H., Garcia, X., Misra, V., Robinson, K., Fedus, L., Zhou, D., Ippolito, D., Luan, D., Lim, H., Zoph, B., Spiridonov, A., Sepassi, R., Dohan, D., Agrawal, S., Omernick, M., Dai, A. M., Pillai, T. S., Pellat, M., Lewkowycz, A., Moreira, E., Child, R., Polozov, O., Lee, K., Zhou, Z., Wang, X., Saeta, B., Diaz, M., Firat, O., Catasta, M., Wei, J., Meier-Hellstern, K., Eck, D., Dean, J., Petrov, S., and Fiedel, N. Palm: Scaling language modeling with pathways. J. Mach. Learn. Res., 24:240:1–240:113, 2023. URL http://jmlr.org/papers/v24/22-1144.html.
  11. 11.Christiano, P. F., Leike, J., Brown, T. B., Martic, M., Legg, S., and Amodei, D. Deep reinforcement learning from human preferences. In Guyon, I., von Luxburg, U., Bengio, S., Wallach, H. M., Fergus, R., Vishwanathan, S. V. N., and Garnett, R. (eds.), Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, pp. 4299–4307, 2017. URL https://proceedings.neurips.cc/paper/2017/hash/d5e2c0adad503c91f91df240d0cd4e49-Abstract.html.
  12. 12.Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, E., Wang, X., Dehghani, M., Brahma, S., Webson, A., Gu, S. S., Dai, Z., Suzgun, M., Chen, X., Chowdhery, A., Narang, S., Mishra, G., Yu, A., Zhao, V. Y., Huang, Y., Dai, A. M., Yu, H., Petrov, S., Chi, E. H., Dean, J., Devlin, J., Roberts, A., Zhou, D., Le, Q. V., and Wei, J. Scaling instruction-finetuned language models. CoRR, abs/2210.11416, 2022. doi: 10.48550/ARXIV.2210.11416. URL https://doi.org/10.48550/arXiv.2210.11416.
  13. 13.Dai, J., Pan, X., Sun, R., Ji, J., Xu, X., Liu, M., Wang, Y., and Yang, Y. Safe rlhf: Safe reinforcement learning from human feedback. arXiv preprint arXiv:2310.12773, 2023.
  14. 14.Dong, H., Xiong, W., Goyal, D., Pan, R., Diao, S., Zhang, J., Shum, K., and Zhang, T. RAFT: reward ranked finetuning for generative foundation model alignment. CoRR, abs/2304.06767, 2023. doi: 10.48550/ARXIV.2304.06767. URL https://doi.org/10.48550/arXiv.2304.06767.
  15. 15.Engstrom, L., Ilyas, A., Santurkar, S., Tsipras, D., Janoos, F., Rudolph, L., and Madry, A. Implementation matters in deep RL: A case study on PPO and TRPO. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net, 2020. URL https://openreview.net/forum?id=r1etN1rtPB.
  16. 16.Gao, L., Schulman, J., and Hilton, J. Scaling laws for reward model overoptimization. In Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., and Scarlett, J. (eds.), International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA, volume 202 of Proceedings of Machine Learning Research, pp. 10835–10866. PMLR, 2023. URL https://proceedings.mlr.press/v202/gao23h.html.
  17. 17.Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In Dy, J. G. and Krause, A. (eds.), Proceedings of the 35th International Conference on Machine Learning, ICML 2018, Stockholmsmassan, Stockholm, Sweden, July 10-15, 2018, volume 80 of Proceedings of Machine Learning Research, pp. 1856–1865. PMLR, 2018. URL http://proceedings.mlr.press/v80/haarnoja18b.html.
  18. 18.Hendrycks, D., Basart, S., Kadavath, S., Mazeika, M., Arora, A., Guo, E., Burns, C., Puranik, S., He, H., Song, D., et al. Measuring coding challenge competence with apps. arXiv preprint arXiv:2105.09938, 2021.
  19. 19.Hong, J., Lee, N., and Thorne, J. Reference-free monolithic preference optimization with odds ratio. arXiv preprint arXiv:2403.07691, 2024.
  20. 20.Kadavath, S., Conerly, T., Askell, A., Henighan, T., Drain, D., Perez, E., Schiefer, N., Hatfield-Dodds, Z., DasSarma, N., Tran-Johnson, E., Johnston, S., Showk, S. E., Jones, A., Elhage, N., Hume, T., Chen, A., Bai, Y., Bowman, S., Fort, S., Ganguli, D., Hernandez, D., Jacobson, J., Kernion, J., Kravec, S., Lovitt, L., Ndousse, K., Olsson, C., Ringer, S., Amodei, D., Brown, T., Clark, J., Joseph, N., Mann, B., McCandlish, S., Olah, C., and Kaplan, J. Language models (mostly) know what they know. CoRR, abs/2207.05221, 2022. doi: 10.48550/ARXIV.2207.05221. URL https://doi.org/10.48550/arXiv.2207.05221.
  21. 21.Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D. Scaling laws for neural language models. CoRR, abs/2001.08361, 2020. URL https://arxiv.org/abs/2001.08361.
  22. 22.Langley, P. Crafting papers on machine learning. In Langley, P. (ed.), Proceedings of the 17th International Conference on Machine Learning (ICML 2000), pp. 1207–1216, Stanford, CA, 2000. Morgan Kaufmann.
  23. 23.Lee, H., Phatale, S., Mansoor, H., Lu, K., Mesnard, T., Bishop, C., Carbune, V., and Rastogi, A. RLAIF: scaling reinforcement learning from human feedback with AI feedback. CoRR, abs/2309.00267, 2023. doi: 10.48550/ARXIV.2309.00267. URL https://doi.org/10.48550/arXiv.2309.00267.
  24. 24.Lewis, M., Yarats, D., Dauphin, Y. N., Parikh, D., and Batra, D. Deal or no deal? end-to-end learning for negotiation dialogues. arXiv preprint arXiv:1706.05125, 2017.
  25. 25.Li, Y., Choi, D., Chung, J., Kushman, N., Schrittwieser, J., Leblond, R., Eccles, T., Keeling, J., Gimeno, F., Dal Lago, A., et al. Competition-level code generation with alphacode. Science, 378(6624):1092–1097, 2022.
  26. 26.Li, Z., Xu, T., and Yu, Y. Policy optimization in rlhf: The impact of out-of-preference data. arXiv preprint arXiv:2312.10584, 2023.
  27. 27.Liang, P. P., Wu, C., Morency, L., and Salakhutdinov, R. Towards understanding and mitigating social biases in language models. In Meila, M. and Zhang, T. (eds.), Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event, volume 139 of Proceedings of Machine Learning Research, pp. 6565–6576. PMLR, 2021. URL http://proceedings.mlr.press/v139/liang21a.html.
  28. 28.Liu, T., Zhao, Y., Joshi, R., Khalman, M., Saleh, M., Liu, P. J., and Liu, J. Statistical rejection sampling improves preference optimization. CoRR, abs/2309.06657, 2023. doi: 10.48550/ARXIV.2309.06657. URL https://doi.org/10.48550/arXiv.2309.06657.
  29. 29.MistralAI. Mistral 7B — mistral.ai. https://mistral.ai/news/announcing-mistral-7b/, 2023. [Accessed 18-01-2024].
  30. 30.Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T. P., Harley, T., Silver, D., and Kavukcuoglu, K. Asynchronous methods for deep reinforcement learning. In Balcan, M. and Weinberger, K. Q. (eds.), Proceedings of the 33nd International Conference on Machine Learning, ICML 2016, New York City, NY, USA, June 19-24, 2016, volume 48 of JMLR Workshop and Conference Proceedings, pp. 1928–1937. JMLR.org, 2016. URL http://proceedings.mlr.press/v48/mniha16.html.
  31. 31.OpenAI. Introducing chatgpt, Nov 2022. URL https://openai.com/blog/chatgpt.
  32. 32.OpenAI. GPT-4 technical report. CoRR, abs/2303.08774, 2023. doi: 10.48550/ARXIV.2303.08774. URL https://doi.org/10.48550/arXiv.2303.08774.
  33. 33.Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P. F., Leike, J., and Lowe, R. Training language models to follow instructions with human feedback. In Koyejo, S., Mohamed, S., Agarwal, A., Belgrave, D., Cho, K., and Oh, A. (eds.), Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 - December 9, 2022, 2022. URL http://papers.nips.cc/paper_files/paper/2022/hash/b1efde53be364a73914f58805a001731-Abstract-Conference.html.
  34. 34.Peng, B., Li, C., He, P., Galley, M., and Gao, J. Instruction tuning with GPT-4. CoRR, abs/2304.03277, 2023. doi: 10.48550/ARXIV.2304.03277. URL https://doi.org/10.48550/arXiv.2304.03277.
  35. 35.Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019.
  36. 36.Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C. Direct preference optimization: Your language model is secretly a reward model. CoRR, abs/2305.18290, 2023. doi: 10.48550/ARXIV.2305.18290. URL https://doi.org/10.48550/arXiv.2305.18290.
  37. 37.Raffin, A., Hill, A., Gleave, A., Kanervisto, A., Ernestus, M., and Dormann, N. Stable-baselines3: Reliable reinforcement learning implementations. Journal of Machine Learning Research, 22(268):1–8, 2021. URL http://jmlr.org/papers/v22/20-1364.html.
  38. 38.Ramamurthy, R., Ammanabrolu, P., Brantley, K., Hessel, J., Sifa, R., Bauckhage, C., Hajishirzi, H., and Choi, Y. Is reinforcement learning (not) for natural language processing: Benchmarks, baselines, and building blocks for natural language policy optimization. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net, 2023. URL https://openreview.net/pdf?id=8aHzds2uUyB.
  39. 39.Rame, A., Vieillard, N., Hussenot, L., Dadashi, R., Cideron, G., Bachem, O., and Ferret, J. Warm: On the benefits of weight averaged reward models. arXiv preprint arXiv:2401.12187, 2024.
  40. 40.Russell, S. Human-compatible artificial intelligence. In Muggleton, S. H. and Chater, N. (eds.), Human-Like Machine Intelligence, pp. 3–23. Oxford University Press, 2022. doi: 10.1093/OSO/9780198862536.003.0001. URL https://doi.org/10.1093/oso/9780198862536.003.0001.
  41. 41.Russell, S. and Norvig, P. Artificial Intelligence: A Modern Approach (4th Edition). Pearson, 2020. ISBN 9780134610993. URL http://aima.cs.berkeley.edu/.
  42. 42.Santacroce, M., Lu, Y., Yu, H., Li, Y., and Shen, Y. Efficient RLHF: reducing the memory usage of PPO. CoRR, abs/2309.00754, 2023. doi: 10.48550/ARXIV.2309.00754. URL https://doi.org/10.48550/arXiv.2309.00754.
  43. 43.Schulman, J., Moritz, P., Levine, S., Jordan, M. I., and Abbeel, P. High-dimensional continuous control using generalized advantage estimation. In Bengio, Y. and LeCun, Y. (eds.), 4th International Conference on Learning Representations, ICLR 2016, San Juan, Puerto Rico, May 2-4, 2016, Conference Track Proceedings, 2016. URL http://arxiv.org/abs/1506.02438.
  44. 44.Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. Proximal policy optimization algorithms. CoRR, abs/1707.06347, 2017. URL http://arxiv.org/abs/1707.06347.
  45. 45.Sellam, T., Das, D., and Parikh, A. P. BLEURT: learning robust metrics for text generation. In Jurafsky, D., Chai, J., Schluter, N., and Tetreault, J. R. (eds.), Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, pp. 7881–7892. Association for Computational Linguistics, 2020. doi: 10.18653/V1/2020.ACL-MAIN.704. URL https://doi.org/10.18653/v1/2020.acl-main.704.
  46. 46.Sheng, E., Chang, K., Natarajan, P., and Peng, N. The woman worked as a babysitter: On biases in language generation. In Inui, K., Jiang, J., Ng, V., and Wan, X. (eds.), Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019, pp. 3405–3410. Association for Computational Linguistics, 2019. doi: 10.18653/V1/D19-1339. URL https://doi.org/10.18653/v1/D19-1339.
  47. 47.Shi, F., Chen, X., Misra, K., Scales, N., Dohan, D., Chi, E. H., Scharli, N., and Zhou, D. Large language models can be easily distracted by irrelevant context. In Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., and Scarlett, J. (eds.), International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA, volume 202 of Proceedings of Machine Learning Research, pp. 31210–31227. PMLR, 2023. URL https://proceedings.mlr.press/v202/shi23a.html.
  48. 48.Singhal, P., Goyal, T., Xu, J., and Durrett, G. A long way to go: Investigating length correlations in rlhf. arXiv preprint arXiv:2310.03716, 2023.
  49. 49.Song, F., Yu, B., Li, M., Yu, H., Huang, F., Li, Y., and Wang, H. Preference ranking optimization for human alignment. CoRR, abs/2306.17492, 2023. doi: 10.48550/ARXIV.2306.17492. URL https://doi.org/10.48550/arXiv.2306.17492.
  50. 50.Stiennon, N., Ouyang, L., Wu, J., Ziegler, D. M., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F. Learning to summarize with human feedback. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., and Lin, H. (eds.), Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual, 2020. URL https://proceedings.neurips.cc/paper/2020/hash/1f89885d556929e98d3ef9b86448f951-Abstract.html.
  51. 51.Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B. Stanford alpaca: An instruction-following llama model, 2023.
  52. 52.Tay, Y., Dehghani, M., Tran, V. Q., Garcia, X., Wei, J., Wang, X., Chung, H. W., Bahri, D., Schuster, T., Zheng, H. S., Zhou, D., Houlsby, N., and Metzler, D. UL2: unifying language learning paradigms. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net, 2023. URL https://openreview.net/pdf?id=6ruVLB727MC.
  53. 53.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Canton-Ferrer, C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., Fuller, B., Gao, C., Goswami, V., Goyal, N., Hartshorn, A., Hosseini, S., Hou, R., Inan, H., Kardas, M., Kerkez, V., Khabsa, M., Kloumann, I., Korenev, A., Koura, P. S., Lachaux, M., Lavril, T., Lee, J., Liskovich, D., Lu, Y., Mao, Y., Martinet, X., Mihaylov, T., Mishra, P., Molybog, I., Nie, Y., Poulton, A., Reizenstein, J., Rungta, R., Saladi, K., Schelten, A., Silva, R., Smith, E. M., Subramanian, R., Tan, X. E., Tang, B., Taylor, R., Williams, A., Kuan, J. X., Xu, P., Yan, Z., Zarov, I., Zhang, Y., Fan, A., Kambadur, M., Narang, S., Rodriguez, A., Stojnic, R., Edunov, S., and Scialom, T. Llama 2: Open foundation and fine-tuned chat models. CoRR, abs/2307.09288, 2023. doi: 10.48550/ARXIV.2307.09288. URL https://doi.org/10.48550/arXiv.2307.09288.
  54. 54.Xiong, W., Dong, H., Ye, C., Wang, Z., Zhong, H., Ji, H., Jiang, N., and Zhang, T. Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint. In ICLR 2024 Workshop on Mathematical and Empirical Understanding of Foundation Models, 2023.
  55. 55.Yang, K., Klein, D., Celikyilmaz, A., Peng, N., and Tian, Y. RLCD: reinforcement learning from contrast distillation for language model alignment. CoRR, abs/2307.12950, 2023. doi: 10.48550/ARXIV.2307.12950. URL https://doi.org/10.48550/arXiv.2307.12950.
  56. 56.Yao, Z., Aminabadi, R. Y., Ruwase, O., Rajbhandari, S., Wu, X., Awan, A. A., Rasley, J., Zhang, M., Li, C., Holmes, C., Zhou, Z., Wyatt, M., Smith, M., Kurilenko, L., Qin, H., Tanaka, M., Che, S., Song, S. L., and He, Y. Deepspeed-chat: Easy, fast and affordable RLHF training of chatgpt-like models at all scales. CoRR, abs/2308.01320, 2023. doi: 10.48550/ARXIV.2308.01320. URL https://doi.org/10.48550/arXiv.2308.01320.
  57. 57.Yu, C., Velu, A., Vinitsky, E., Gao, J., Wang, Y., Bayen, A. M., and Wu, Y. The surprising effectiveness of PPO in cooperative multi-agent games. In Koyejo, S., Mohamed, S., Agarwal, A., Belgrave, D., Cho, K., and Oh, A. (eds.), Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 - December 9, 2022, 2022. URL http://papers.nips.cc/paper_files/paper/2022/hash/9c1535a02f0ce079433344e14d910597-Abstract-Datasets_and_Benchmarks.html.
  58. 58.Yuan, W., Pang, R. Y., Cho, K., Sukhbaatar, S., Xu, J., and Weston, J. Self-rewarding language models. CoRR, abs/2401.10020, 2024. URL https://arxiv.org/pdf/2401.10020.pdf.
  59. 59.Yuan, Z., Yuan, H., Tan, C., Wang, W., Huang, S., and Huang, F. RRHF: rank responses to align language models with human feedback without tears. CoRR, abs/2304.05302, 2023. doi: 10.48550/ARXIV.2304.05302. URL https://doi.org/10.48550/arXiv.2304.05302.
  60. 60.Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., and Artzi, Y. Bertscore: Evaluating text generation with BERT. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net, 2020. URL https://openreview.net/forum?id=SkeHuCVFDr.
  61. 61.Zheng, R., Dou, S., Gao, S., Hua, Y., Shen, W., Wang, B., Liu, Y., Jin, S., Liu, Q., Zhou, Y., Xiong, L., Chen, L., Xi, Z., Xu, N., Lai, W., Zhu, M., Chang, C., Yin, Z., Weng, R., Cheng, W., Huang, H., Sun, T., Yan, H., Gui, T., Zhang, Q., Qiu, X., and Huang, X. Secrets of RLHF in large language models part I: PPO. CoRR, abs/2307.04964, 2023. doi: 10.48550/ARXIV.2307.04964. URL https://doi.org/10.48550/arXiv.2307.04964.
  62. 62.Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P. F., and Irving, G. Fine-tuning language models from human preferences. CoRR, abs/1909.08593, 2019. URL http://arxiv.org/abs/1909.08593.

Citation

MLA
Xu, S., et al. “Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study”. ICML 2024, 2024, http://arxiv.org/abs/2404.10719v3.
APA
Xu, S., Fu, W., Gao, J., Ye, W., Liu, W., Mei, Z., Wang, G., Yu, C., & Wu, Y. (2024). Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study. ICML 2024. http://arxiv.org/abs/2404.10719v3
Chicago
Xu, S., W. Fu, J. Gao, et al. 2024. “Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study”. ICML 2024. http://arxiv.org/abs/2404.10719v3.
Harvard
Xu, S. et al. (2024) “Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study”, ICML 2024 [Preprint]. Available at: http://arxiv.org/abs/2404.10719v3.
Vancouver
1. Xu S, Fu W, Gao J, Ye W, Liu W, Mei Z, Wang G, Yu C, Wu Y (2024) Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study. ICML 2024

BibTeX

@article{xu2024dpo,
  title = {Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study},
  author = {Xu, Shusheng and Fu, Wei and Gao, Jiaxuan and Ye, Wenjie and Liu, Weilin and Mei, Zhiyu and Wang, Guangju and Yu, Chao and Wu, Yi},
  year = {2024},
  journal = {ICML 2024},
  url = {http://arxiv.org/abs/2404.10719v3},
  eprint = {2404.10719}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/