trlX: A Framework for Large Scale Reinforcement Learning from Human Feedback

Alexander HavrillaMaksym ZhuravinskyiDuy PhungAman TiwariJonathan TowStella BidermanQuentin AnthonyLouis Castricato

article2023EMNLP59 citations

Presents trlX, an open-source framework that scales reinforcement learning from human feedback to language models exceeding 70 billion parameters by integrating advanced distributed parallelism schemes with memory-efficient online and offline training methods.

Listen

Large language models require reinforcement learning from human feedback to better align their outputs with human preferences and produce helpful, safe responses. However, fine-tuning large models with standard online reinforcement learning algorithms like Proximal Policy Optimization is computationally expensive, memory-intensive, and difficult to scale, which has historically restricted research to organizations with massive compute infrastructure.

The article demonstrates the capabilities of trlX, a feature-complete open-source framework designed for large-scale reinforcement learning fine-tuning. It evaluates how distributed parallelization, memory-saving optimizations, and alternative offline learning algorithms can scale fine-tuning to models exceeding 70 billion parameters across various resource tiers.

The authors evaluated the framework through empirical experiments on standard text summarization and helpful dialogue tasks using public benchmark datasets. The approach tested models ranging from 125 million to 20 billion parameters using both online and offline algorithms. To lower hardware barriers, the researchers integrated parameter-efficient fine-tuning techniques, layer-freezing architectures, and advanced parallel distributed training frameworks, validating output quality through standard academic language benchmarks and blind human preference evaluations.

The analysis yielded several critical findings. First, online models fine-tuned with the framework achieved over a 70% win-rate against supervised baselines in summarization and at least a 60% win-rate across dialogue benchmarks. Second, offline reinforcement learning through Implicit Language Q-Learning delivered competitive performance—also reaching win-rates above 60%—while using only a fraction of the compute time and proving significantly more resistant to reward model overfitting. Third, combining parameter-efficient adapters and partial layer freezing reduced memory overhead by up to 75% on medium-sized models while maintaining maximum attainable task performance. Finally, the evaluation revealed that the typical drop in general benchmark capability, known as the alignment tax, stems primarily from the initial supervised fine-tuning stage rather than reinforcement learning itself.

These findings indicate that alignment fine-tuning can be made substantially more cost-effective and accessible without sacrificing output quality. Organizations can achieve strong performance gains even on smaller models or limited hardware setups. Furthermore, because offline reinforcement learning circumvents the heavy infrastructure needed to maintain multiple concurrent models in memory, it provides a viable, budget-friendly alternative for preference alignment pipelines.

Teams developing language models should adopt modular memory-saving techniques like low-rank adapters and layer freezing during alignment to reduce infrastructure costs. For resource-constrained deployments, decision-makers should evaluate offline reinforcement learning as a low-overhead alternative to standard online methods. When designing alignment workflows, engineering teams must prioritize the quality and diversity of the supervised fine-tuning data, as this phase largely dictates whether broader model knowledge is preserved.

Readers should note that while offline methods are simpler to implement and more compute-efficient, they still slightly trail online optimization in absolute performance. Additionally, reinforcement learning alignment does not entirely eliminate model hallucinations or biases. The framework provides high confidence in scalable alignment execution, but organizations deploying these models must maintain ongoing monitoring and mitigation strategies during live inference.

Cover for trlX: A Framework for Large Scale Reinforcement Learning from Human Feedback

Abstract

Reinforcement learning from human feedback (RLHF) utilizes human feedback to better align large language models with human preferences via online optimization against a learned reward model. Current RLHF paradigms rely on Proximal Policy Optimization (PPO), which quickly becomes a challenge to implement and scale up to large architectures. To address this difficulty we present the trlX library as a feature-complete open-source framework for RLHF fine-tuning of models up to and exceeding 70 billion parameters. We implement support for multiple types of distributed training including distributed data parallel, model sharded, as well as tensor, sequential, and pipeline parallelism.

To increase the accessibility of RLHF to researchers, we implement compute- and memory-saving features that give trlX the flexibility to support users with a wide range of compute resources. This includes offline RL methods like Implicit Language Q Learning (IQL), low-rank adapters, and the Hydra architecture. We find offline fine-tuning offers competitive performance relative to online algorithms while being easier to implement, train, and scale. To evaluate our framework we train RLHF models on two separate well-known tasks using publicly available human preference data. Models trained with trlX achieve preference win-rates over baselines at rates comparable to the original works.

Table of Contents

  • 1 Introduction
  • 2 Background
  • Reinforcement Learning from Human Feedback
  • 3 Training with trlX
  • 3.1 Memory and Compute Saving Features
  • 3.2 Comparison with other Frameworks
  • 4 Benchmarks and Results
  • 4.1 Summarization
  • 4.2 Helpful QA Dialogue
  • 5 Conclusion
  • References
  • A Model Hyperparameters
  • B LMEval Results
  • C Annotator Instructions
  • Annotation Guidelines:
  • Helpfulness
  • Correctness
  • Harmful
  • D Synthetic Reward Modeling

Knowls

  1. Knowl 1 — trlX unifies online and offline RLHF fine-tuning

    model/method

    trlX is an open-source framework for fine-tuning language models with reinforcement learning from human feedback. It implements PPO and A2C for online reinforcement learning and Implicit Language Q-Learning (ILQL) for offline reinforcement learning, with shared trainer infrastructure and integrations for multiple distributed-training backends. The framework supports common Hugging Face encoder–decoder and decoder-only models, and includes parameter-efficient fine-tuning and hyperparameter-sweep tools. Its distinguishing scope is the combination of online and offline RLHF with large-model training support, including pipeline, sequence, and tensor parallelism.

  2. Knowl 2 — Distributed training extends trlX to models above 70 billion parameters

    experimental setup

    trlX provides different training paths for different hardware budgets: native PyTorch with memory-saving features for single-GPU users; Hugging Face Accelerate with DeepSpeed for multi-GPU training, which the authors used for models up to 20 billion parameters on one node; and GPT-NeoX or NeMo-Megatron integrations for multi-node training with tensor, sequence, and pipeline parallelism, used for models up to 70 billion parameters. In a PPO throughput comparison on OPT models, reported speeds in samples per second per GPU were 2.1 for DeepSpeed-Chat versus 2.0 for trlX at 1.3B parameters, 0.44 versus 0.41 at 6.7B, and 0.14 versus 0.12 at 30B. At the largest scale, DeepSpeed-Chat reported 0.076 for OPT 60B using LoRA, while trlX reported 0.043 for OPT 66B using Hydra with 50% of parameters trainable. The 30B and 60B DeepSpeed-Chat values were converted from a published four-hour run on 131.9k samples using 64 GPUs; the compared configurations therefore differ in model size and, at large sizes, fine-tuning method.

  3. Knowl 3 — Hydra sharing and LoRA reduce RLHF memory requirements

    empirical result

    trlX combines Hydra, which shares frozen layers among policy, value, and reference networks, with LoRA adapters to reduce memory and computation during fine-tuning. In a sentiment-task study using Pythia models from 125M to 20B parameters, models were trained for 6,000 steps with global batch size 32 on eight 80-GB A100 GPUs. The reward-versus-unfrozen-layer chart on page 4 shows that roughly half the layers could be frozen while retaining each model's maximum observed reward; freezing all but two layers was more damaging for larger models. The memory chart on page 4 and accompanying measurements report savings approaching 50% for large models with extensive freezing. On the sentiment benchmark, rank-1 LoRA reached maximum reward; for GPT-J at 6.9B parameters it trained only 0.03% of parameters and reduced memory use by a factor of three. The authors report that combining LoRA and layer freezing can enable medium-scale RLHF on a consumer-grade GPU, and that the memory-saving benefits also apply to ILQL.

  4. Knowl 4 — An orchestrator separates generation and PPO optimization batch sizes

    model/method

    trlX uses an orchestrator to decouple the inference batch size used to generate policy rollouts from the PPO batch size used for optimization. The orchestrator batches rollout generation independently, allowing users to choose rollout and update batch sizes suited to their respective workloads. This addresses a major online-training bottleneck: the authors report that rollout generation can take up to ten times as long as a combined forward and backward pass. The design is intended to reduce time spent in sequential model inference while allowing efficient batch sizes for both generation and PPO updates.

  5. Knowl 5 — PPO training uses large batches and normalized rewards and advantages

    model/method

    The paper's PPO training recipe emphasizes large global batches, reward scaling, and advantage normalization. The appendix reports a PPO learning rate of 5 × 10⁻⁶, batch size 256, four PPO epochs, target KL of 6, generalized-advantage-estimation λ of 0.95, discount factor γ of 1, and batch-level advantage normalization. Rewards are scaled by a running standard-deviation estimate without subtracting the running mean; advantages are then normalized at the batch level. The authors report that global batches of at least 128 samples helped stabilize performance and reduce run-to-run variance. The appendix gives a KL coefficient of 0.01 as a general setting, while the summarization and Helpful QA experiments described in the main text use 0.005.

  6. Knowl 6 — ILQL reaches competitive reward with substantially less training compute

    empirical result

    On Anthropic's Helpful QA dialogue task, the authors compared ILQL training time and peak reward with and without LoRA. GPT-NeoX 20B reached a maximum reward of -1.88 in 156 minutes on 32 GPUs; its LoRA variant reached -1.89 in 28 minutes on 16 GPUs, with LoRA applied to the last eight layers. Pythia 6.9B reached -1.77 in 286 minutes on 16 GPUs, while its LoRA variant reached -1.68 in 58 minutes on 16 GPUs, with LoRA applied to all layers. The non-LoRA hyperparameters were held fixed apart from a learning rate of 2.0 × 10⁻⁴. These results show that LoRA substantially reduced the reported time and GPU count while retaining comparable maximum reward in these runs. For offline preference training, the paper assigns +1 to preferred trajectories and -1 to rejected trajectories rather than relying on learned reward-model scores.

  7. Knowl 7 — Summarization evaluations favor PPO while showing promise for offline ILQL

    empirical result

    For TL;DR summarization, the authors trained reward models using 92,534 preference comparisons and evaluated on human judgments of summaries generated for test prompts. The PPO models were initialized from supervised fine-tuned (SFT) checkpoints, used a 6.9B reward model, and kept eight layers unfrozen. ILQL used posts and their associated summaries labeled -1 and +1, respectively; the authors found this labeling performed better than using learned reward-model scores and saw little benefit from initializing ILQL with the SFT checkpoint. In same-size comparisons against SFT, the human-preference chart on page 6 reports PPO win rates above 70% for the 6B and 20B models; ILQL was also competitive across model sizes despite using less compute. Annotators additionally rated coverage, clarity, and inconsistency. The authors observed that ILQL summaries were shorter than PPO and SFT summaries, yet were often preferred to SFT summaries for better coverage of key points.

  8. Knowl 8 — Helpful QA training combines human preferences with synthetic prompts

    experimental setup

    The Helpful QA experiments use Anthropic's Helpful and Harmless dataset, which contains about 118,000 interactions; its static portion includes approximately 42,000 prompted-model examples and 40,000 reranked examples. SFT baselines at 160M, 1.4B, 6.9B, and 20B were trained for one epoch on chosen responses, masking loss on dialogue history and using a learning rate of 5 × 10⁻⁵. Reward models were initialized from SFT models and trained for one epoch at 5 × 10⁻⁶; the best model, GPT-J at about 6B parameters, achieved 0.72 accuracy on the static test set and was used for RL training. The PPO prompt set contained 200,000 prompts, including synthetic multi-turn data; runs used 9,000 steps, effective batch size 128, model-dependent learning rates from 1 × 10⁻⁶ to 8 × 10⁻⁶, eight unfrozen layers, and KL coefficient 0.005. ILQL assigned +1 to chosen trajectories and -1 to rejected ones. The authors also trained Vanilla-PPO without SFT initialization, finding this feasible only for the 6B and 20B models.

  9. Knowl 9 — Helpful QA results suggest SFT drives much of the benchmark-performance drop

    empirical result

    In human evaluations against same-size SFT baselines, both PPO and ILQL achieved at least 60% win rates at the tested model sizes, according to the chart on page 8. The authors also observed that ILQL was more robust to reward-model overfitting; PPO required large batches and early stopping to limit it. On zero-shot HellaSwag and LAMBADA, SFT reduced the Pythia 6.9B scores from 0.488 and 0.564 (vanilla) to 0.432 and 0.398; PPO after SFT scored 0.421 and 0.409, while Vanilla-PPO scored 0.495 and 0.605. For GPT-NeoX 20B, the corresponding vanilla scores were 0.535 and 0.720, SFT scores 0.462 and 0.505, PPO scores 0.463 and 0.529, and Vanilla-PPO scores 0.548 and 0.618. Across the broader set of six academic benchmarks, the authors characterize RL fine-tuning on top of SFT as producing small changes, while Vanilla-PPO incurred less of the performance penalty and sometimes improved scores. They interpret these results as evidence that much of the observed benchmark reduction may arise from SFT rather than RL fine-tuning itself, while noting the need for high-quality SFT data.

  10. Knowl 10 — Synthetic preference models can fit synthetic rankings but transfer poorly to human preferences

    empirical result

    The authors studied two ways to create synthetic preference data for Helpful QA: have text-davinci-003 judge competing responses, or assume responses from larger supervised models are preferable to responses from smaller models. For the size-ordered approach, the assumed ranking was 125M < 1.4B < 6.9B < 20B < text-davinci-002 < text-davinci-003. Reward models trained on these synthetic comparisons achieved over 90% accuracy on held-out synthetic comparisons in the reported best cases; the page 16 accuracy plot shows the 20B reward model was most sample-efficient up to 120,000 comparisons, after which the 6B model did slightly better. However, the best size-ordered reward model achieved only 0.61 accuracy on the human-labeled Helpful QA test split, whereas the human-preference GPT-J reward model scored 0.78 on the synthetic test data. As a separate check, text-davinci-003 reached 0.64 one-shot accuracy as a Helpful QA preference classifier, compared with 0.71 for the GPT-J reward model. The results caution that accuracy on synthetic comparisons does not establish transfer to human preferences.

Coverage note — The detailed annotator instructions, full scores for every model on all six language-model benchmarks, and per-setting freezing curves are omitted because their main implications are captured in the evaluation and memory knowls.

References

  1. 1.Andonian, A., Anthony, Q., Biderman, S., Black, S., Gali, P., Gao, L., Hallahan, E., Levy-Kramer, J., Leahy, C., Nestler, L., et al. Gpt-neox: Large scale autoregressive language modeling in pytorch. GitHub Repo, 2021.
  2. 2.Askell, A., Bai, Y., Chen, A., Drain, D., Ganguli, D., Henighan, T., Jones, A., Joseph, N., Mann, B., DasSarma, N., et al. A general language assistant as a laboratory for alignment. arXiv preprint arXiv:2112.00861, 2021.
  3. 3.Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T. J., Joseph, N., Kadavath, S., Kernion, J., Conerly, T., El-Showk, S., Elhage, N., Hatfield-Dodds, Z., Hernandez, D., Hume, T., Johnston, S., Kravec, S., Lovitt, L., Nanda, N., Olsson, C., Amodei, D., Brown, T. B., Clark, J., McCandlish, S., Olah, C., Mann, B., and Kaplan, J. Training a helpful and harmless assistant with reinforcement learning from human feedback. ArXiv, abs/2204.05862, 2022a.
  4. 4.Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., et al. Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073, 2022b.
  5. 5.Biderman, S., Bicheno, K., and Gao, L. Datasheet for the pile. arXiv preprint arXiv:2201.07311, 2022.
  6. 6.Biderman, S. R., Schoelkopf, H., Anthony, Q. G., Bradley, H., O’Brien, K., Hallahan, E., Khan, M. A., Purohit, S., Prashanth, U. S., Raff, E., Skowron, A., Sutawika, L., and van der Wal, O. Pythia: A suite for analyzing large language models across training and scaling. ArXiv, abs/2304.01373, 2023.
  7. 7.Black, S., Biderman, S., Hallahan, E., Anthony, Q., Gao, L., Golding, L., He, H., Leahy, C., McDonell, K., Phang, J., Pieler, M., Prashanth, U. S., Purohit, S., Reynolds, L., Tow, J., Wang, B., and Weinbach, S. GPT-NeoX-20B: An open-source autoregressive language model. In Proceedings of the ACL Workshop on Challenges & Perspectives in Creating Large Language Models, 2022. URL https://arxiv.org/abs/2204.06745.
  8. 8.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. Language models are few-shot learners. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M. F., and Lin, H. (eds.), Advances in Neural Information Processing Systems, volume 33, pp. 1877–1901. Curran Associates, Inc., 2020. URL https://proceedings.neurips.cc/paper/2020/file/1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf.
  9. 9.Christiano, P. F., Leike, J., Brown, T. B., Martic, M., Legg, S., and Amodei, D. Deep reinforcement learning from human preferences. ArXiv, abs/1706.03741, 2017.
  10. 10.Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, E., Wang, X., Dehghani, M., Brahma, S., Webson, A., Gu, S. S., Dai, Z., Suzgun, M., Chen, X., Chowdhery, A., Valter, D., Narang, S., Mishra, G., Yu, A. W., Zhao, V., Huang, Y., Dai, A. M., Yu, H., Petrov, S., hsin Chi, E. H., Dean, J., Devlin, J., Roberts, A., Zhou, D., Le, Q. V., and Wei, J. Scaling instruction-finetuned language models. ArXiv, abs/2210.11416, 2022.
  11. 11.Dettmers, T., Lewis, M., Shleifer, S., and Zettlemoyer, L. 8-bit optimizers via block-wise quantization. ArXiv, abs/2110.02861, 2021.
  12. 12.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding. ArXiv, abs/1810.04805, 2019.
  13. 13.Ganguli, D., Lovitt, L., Kernion, J., Askell, A., Bai, Y., Kadavath, S., Mann, B., Perez, E., Schiefer, N., Ndousse, K., et al. Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned. arXiv preprint arXiv:2209.07858, 2022.
  14. 14.Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., et al. The pile: An 800gb dataset of diverse text for language modeling. arXiv preprint arXiv:2101.00027, 2020.
  15. 15.Gao, L., Tow, J., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., McDonell, K., Muennighoff, N., et al. A framework for few-shot language model evaluation. Version v0. 0.1. Sept, 2021.
  16. 16.Glaese, A., McAleese, N., Trkebacz, M., Aslanides, J., Firoiu, V., Ewalds, T., Rauh, M., Weidinger, L., Chadwick, M., Thacker, P., Campbell-Gillingham, L., Uesato, J., Huang, P.-S., Comanescu, R., Yang, F., See, A., Dathathri, S., Greig, R., Chen, C., Fritz, D., Elias, J. S., Green, R., Mokr’a, S., Fernando, N., Wu, B., Foley, R., Young, S., Gabriel, I., Isaac, W. S., Mellor, J. F. J., Hassabis, D., Kavukcuoglu, K., Hendricks, L. A., and Irving, G. Improving alignment of dialogue agents via targeted human judgements. ArXiv, abs/2209.14375, 2022.
  17. 17.Gugger, S., Debut, L., Thomas Wolf, T., Schmid, P., Mueller, Z., and Mangrulkar, S. Accelerate: Training and inference at scale made simple, efficient and adaptable. https://github.com/huggingface/accelerate, 2022.
  18. 18.Honovich, O., Scialom, T., Levy, O., and Schick, T. Unnatural instructions: Tuning language models with (almost) no human labor. arXiv preprint arXiv:2212.09689, 2022.
  19. 19.Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., and Chen, W. Lora: Low-rank adaptation of large language models. ArXiv, abs/2106.09685, 2021.
  20. 20.Knox, W. B. and Stone, P. Interactively shaping agents via human reinforcement: The tamer framework. In The Fifth International Conference on Knowledge Capture, September 2009. URL http://www.cs.utexas.edu/users/ai-lab?KCAP09-knox.
  21. 21.Kuchaiev, O., Li, J., Nguyen, H., Hrinchuk, O., Leary, R., Ginsburg, B., Kriman, S., Beliaev, S., Lavrukhin, V., Cook, J., et al. Nemo: a toolkit for building ai applications using neural modules. arXiv preprint arXiv:1909.09577, 2019.
  22. 22.Leandro, V. W. Transformer reinforcement learning. https://github.com/lvwerra/trl, 2019.
  23. 23.Nakano, R., Hilton, J., Balaji, S. A., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., Jiang, X., Cobbe, K., Eloundou, T., Krueger, G., Button, K., Knight, M., Chess, B., and Schulman, J. Webgpt: Browser-assisted question-answering with human feedback. ArXiv, abs/2112.09332, 2021.
  24. 24.Ott, M., Edunov, S., Baevski, A., Fan, A., Gross, S., Ng, N., Grangier, D., and Auli, M. fairseq: A fast, extensible toolkit for sequence modeling. In Proceedings of NAACL-HLT 2019: Demonstrations, 2019.
  25. 25.Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L. E., Simens, M., Askell, A., Welinder, P., Christiano, P. F., Leike, J., and Lowe, R. J. Training language models to follow instructions with human feedback. ArXiv, abs/2203.02155, 2022.
  26. 26.Raffel, C., Shazeer, N. M., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. Exploring the limits of transfer learning with a unified text-to-text transformer. ArXiv, abs/1910.10683, 2019.
  27. 27.Rajbhandari, S., Rasley, J., Ruwase, O., and He, Y. Zero: Memory optimizations toward training trillion parameter models. SC20: International Conference for High Performance Computing, Networking, Storage and Analysis, pp. 1–16, 2019.
  28. 28.Ramamurthy, R., Ammanabrolu, P., Brantley, K., Hessel, J., Sifa, R., Bauckhage, C., Hajishirzi, H., and Choi, Y. Is reinforcement learning (not) for natural language processing?: Benchmarks, baselines, and building blocks for natural language policy optimization. ArXiv, abs/2210.01241, 2022.
  29. 29.Roberts, A., Chung, H. W., Levskaya, A., Mishra, G., Bradbury, J., Andor, D., Narang, S., Lester, B., Gaffney, C., Mohiuddin, A., Hawthorne, C., Lewkowycz, A., Salcianu, A., van Zee, M., Austin, J., Goodman, S., Soares, L. B., Hu, H., Tsvyashchenko, S., Chowdhery, A., Bastings, J., Bulian, J., Garcia, X., Ni, J., Chen, A., Kenealy, K., Clark, J. H., Lee, S., Garrette, D., Lee-Thorp, J., Raffel, C., Shazeer, N., Ritter, M., Bosma, M., Passos, A., Maitin-Shepard, J., Fiedel, N., Omernick, M., Saeta, B., Sepassi, R., Spiridonov, A., Newlan, J., and Gesmundo, A. Scaling up models and data with t5x and seqio. arXiv preprint arXiv:2203.17189, 2022. URL https://arxiv.org/abs/2203.17189.
  30. 30.Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. Proximal policy optimization algorithms. ArXiv, abs/1707.06347, 2017.
  31. 31.Shoeybi, M., Patwary, M., Puri, R., LeGresley, P., Casper, J., and Catanzaro, B. Megatron-lm: Training multi-billion parameter language models using model parallelism. ArXiv, abs/1909.08053, 2019.
  32. 32.Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F. Learning to summarize with human feedback. Advances in Neural Information Processing Systems, 33:3008–3021, 2020.
  33. 33.Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G. Llama: Open and efficient foundation language models. ArXiv, abs/2302.13971, 2023a.
  34. 34.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023b.
  35. 35.Wang, B. and Komatsuzaki, A. GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model. https://github.com/kingoflolz/mesh-transformer-jax, May 2021.
  36. 36.Wang, Y., Kordi, Y., Mishra, S., Liu, A., Smith, N. A., Khashabi, D., and Hajishirzi, H. Self-instruct: Aligning language model with self generated instructions. arXiv preprint arXiv:2212.10560, 2022a.
  37. 37.Wang, Y., Mishra, S., Alipoormolabashi, P., Kordi, Y., Mirzaei, A., Arunkumar, A., Ashok, A., Dhanasekaran, A. S., Naik, A., Stap, D., et al. Supernaturalinstructions: Generalization via declarative instructions on 1600+ nlp tasks. URL https://arxiv.org/abs/2204.07705, 2022b.
  38. 38.Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., et al. Huggingface’s transformers: State-of-the-art natural language processing. arXiv preprint arXiv:1910.03771, 2019.
  39. 39.Yao, Z., Aminabadi, R. Y., Ruwase, O., Rajbhandari, S., Wu, X., Awan, A. A., Rasley, J., Zhang, M., Li, C., Holmes, C., Zhou, Z., Wyatt, M., Smith, M., Kurilenko, L., Qin, H., Tanaka, M., Che, S., Song, S. L., and He, Y. DeepSpeed-Chat: Easy, Fast and Affordable RLHF Training of ChatGPT-like Models at All Scales. arXiv preprint arXiv:2308.01320, 2023.
  40. 40.Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., Mihaylov, T., Ott, M., Shleifer, S., Shuster, K., Simig, D., Koura, P. S., Sridhar, A., Wang, T., and Zettlemoyer, L. Opt: Open pre-trained transformer language models. ArXiv, abs/2205.01068, 2022.
  41. 41.Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G. Fine-tuning language models from human preferences. ArXiv, abs/1909.08593, 2019.

Citation

MLA
Havrilla, A., et al. “trlX: A Framework for Large Scale Reinforcement Learning from Human Feedback”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 8578–95, https://doi.org/10.18653/v1/2023.emnlp-main.530.
APA
Havrilla, A., Zhuravinskyi, M., Phung, D., Tiwari, A., Tow, J., Biderman, S., Anthony, Q., & Castricato, L. (2023). trlX: A Framework for Large Scale Reinforcement Learning from Human Feedback. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 8578–8595. https://doi.org/10.18653/v1/2023.emnlp-main.530
Chicago
Havrilla, A., M. Zhuravinskyi, D. Phung, et al. 2023. “trlX: A Framework for Large Scale Reinforcement Learning from Human Feedback”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 8578–95. https://doi.org/10.18653/v1/2023.emnlp-main.530.
Harvard
Havrilla, A. et al. (2023) “trlX: A Framework for Large Scale Reinforcement Learning from Human Feedback”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 8578–8595. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.530.
Vancouver
1. Havrilla A, Zhuravinskyi M, Phung D, Tiwari A, Tow J, Biderman S, Anthony Q, Castricato L (2023) trlX: A Framework for Large Scale Reinforcement Learning from Human Feedback. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 8578–8595

BibTeX

@inproceedings{havrilla-etal-2023-trlx,
    title = "trl{X}: A Framework for Large Scale Reinforcement Learning from Human Feedback",
    author = "Havrilla, Alexander  and
      Zhuravinskyi, Maksym  and
      Phung, Duy  and
      Tiwari, Aman  and
      Tow, Jonathan  and
      Biderman, Stella  and
      Anthony, Quentin  and
      Castricato, Louis",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.530/",
    doi = "10.18653/v1/2023.emnlp-main.530",
    pages = "8578--8595"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/