Built independently by an author, for readers. Read the story and support ChapterPal

keyword

RL fine-tuning

Reinforcement learning fine-tuning is a post-training method that optimizes a pre-trained or supervised model by using reinforcement learning objectives to maximize a reward signal rather than imitating fixed target demonstrations. In this process, the model acts as a policy that generates actions or text sequences, which are subsequently evaluated and assigned scalar rewards based on human preferences, learned reward models, or automated verification systems. Optimization algorithms, such as policy gradient techniques, adjust the model parameters to increase the probability of high-reward responses while penalizing undesirable outputs. This technique is widely utilized in foundation model alignment to guide model behavior toward safety, helpfulness, and human values, as well as to enhance performance on multi-step reasoning and problem-solving tasks.

4 items

trlX: A Framework for Large Scale Reinforcement Learning from Human Feedback

trlX: A Framework for Large Scale Reinforcement Learning from Human Feedback

Alexander Havrilla, Maksym Zhuravinskyi, Duy Phung, Aman Tiwari, Jonathan Tow, Stella Biderman, Quentin Anthony, Louis Castricato

OrganizationsBooz Allen HamiltonCarperAIEleutherAIGeorgia Institute of TechnologyStability AIThe Ohio State UniversityVectorShift

Why you should read this

Presents trlX, an open-source framework that scales reinforcement learning from human feedback to language models exceeding 70 billion parameters by integrating advanced distributed parallelism schemes with memory-efficient online and offline training methods.

Reinforcement learning from human feedback (RLHF) utilizes human feedback to better align large language models with human preferences via online optimization against a learned reward model. Current RLHF paradigms rely on Proximal Policy Optimization (PPO), which quickly becomes a challenge to implement and scale up to large architectures. To address this difficulty we present the trlX library as a feature-complete open-source framework for RLHF fine-tuning of models up to and exceeding 70 billion parameters. We implement support for multiple types of distributed training including distributed data parallel, model sharded, as well as tensor, sequential, and pipeline parallelism. To increase the accessibility of RLHF to researchers, we implement compute- and memory-saving features that give trlX the flexibility to support users with a wide range of compute resources. This includes offline RL methods like Implicit Language Q Learning (IQL), low-rank adapters, and the Hydra architecture. We find offline fine-tuning offers competitive performance relative to online algorithms while being easier to implement, train, and scale. To evaluate our framework we train RLHF models on two separate well-known tasks using publicly available human preference data. Models trained with trlX achieve preference win-rates over baselines at rates comparable to the original works.

Added

2026-10-03

Do We Need Adam? Surprisingly Strong and Sparse Reinforcement Learning with SGD in LLMs

Do We Need Adam? Surprisingly Strong and Sparse Reinforcement Learning with SGD in LLMs

Sagnik Mukherjee, Lifan Yuan, Pavan Jayasinha, Dilek Hakkani-Tür, Hao Peng

OrganizationsUniversity of Illinois Urbana-ChampaignUniversity of Waterloo

Why you should read this

Demonstrates that memory-efficient SGD matches or outperforms AdamW in reinforcement learning for large language models while naturally updating fewer than 0.02% of model parameters without explicit regularization.

Reinforcement learning (RL), particularly RL from verifiable reward (RLVR), has become a crucial phase of training large language models (LLMs) and a key focus of current scaling efforts. However, optimization practices in RL largely follow those of next-token prediction stages (e.g., pretraining and supervised fine-tuning), despite fundamental differences between RL and these stages highlighted by recent work. One such practice is the use of the AdamW optimizer, which is widely adopted for training large-scale transformers despite its high memory overhead. Our analysis shows that both momentum and adaptive learning rates in AdamW are less influential in RL than in SFT, leading us to hypothesize that RL benefits less from Adam-style per-parameter adaptive learning rates and momentum. Confirming this hypothesis, our experiments demonstrate that the substantially more memory-efficient SGD, which is known to perform poorly in supervised learning of large-scale transformers, matches or even outperforms AdamW in RL for LLMs. Remarkably, full fine-tuning with SGD updates fewer than 0.02% of model parameters without any sparsity-promoting regularization, more than 1000 times fewer than AdamW. Our analysis offers potential reasons for this update sparsity. These findings provide new insights into the optimization dynamics of RL in LLMs and show that RL can be substantially more parameter-efficient than previously recognized.

Added

2026-09-14

Creative Commons License
On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes

On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes

Rishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk, Sabela Ramos, Matthieu Geist, Olivier Bachem

OrganizationsGoogleMilaUniversity of Toronto

Why you should read this

Proposes Generalized Knowledge Distillation, an on-policy framework that resolves distribution mismatch during language model compression by training students on their self-generated sequences with teacher feedback.

Knowledge distillation (KD) is widely used for compressing a teacher model to reduce its inference cost and memory footprint, by training a smaller student model. However, current KD methods for auto-regressive sequence models suffer from distribution mismatch between output sequences seen during training and those generated by the student during inference. To address this issue, we introduce Generalized Knowledge Distillation (GKD). Instead of solely relying on a fixed set of output sequences, GKD trains the student on its self-generated output sequences by leveraging feedback from the teacher on such sequences. Unlike supervised KD approaches, GKD also offers the flexibility to employ alternative loss functions between the student and teacher, which can be useful when the student lacks the expressivity to mimic the teacher's distribution. Furthermore, GKD facilitates the seamless integration of distillation with RL fine-tuning (RLHF). We demonstrate the efficacy of GKD for distilling auto-regressive language models on summarization, translation, and arithmetic reasoning tasks, and task-agnostic distillation for instruction-tuning.

Added

2026-09-03

Creative Commons License