Aligning Language Models with Preferences through f-divergence Minimization
Dongyoung GoTomasz KorbakGermán KruszewskiJos RozenNahyeon RyuMarc Dymetman
Presents f-DPG, a generalized framework that unifies RLHF and Distributional Policy Gradients by minimizing arbitrary f-divergences against target energy-based models, revealing that objectives like Jensen-Shannon divergence achieve superior alignment and diversity trade-offs compared to traditional KL divergence.
Aligning large language models with human preferences—such as harmlessness, truthfulness, and stylistic guidelines—is a central requirement for deploying artificial intelligence safely and effectively. Existing alignment techniques typically frame this challenge as steering a generative model toward a desired target distribution of text. However, current practices remain fragmented across disparate algorithms and mathematical objectives, obscuring how the choice of optimization objective affects the final behavior, diversity, and reliability of the aligned model.
The article introduces and evaluates f-DPG, a generalized framework that unifies existing alignment paradigms under the mathematical family of f-divergence minimization. The primary objective is to demonstrate that aligning language models can be decoupled from specific legacy algorithms, allowing practitioners to optimize any evaluating target distribution using diverse mathematical objectives to achieve superior trade-offs between preference adherence and text diversity.
To establish this framework, the authors derived a universal policy gradient formula that supports variance reduction and conditioned generation, encompassing methods like standard reinforcement learning with divergence penalties and generative distributional control. They empirically benchmarked four distinct divergence objectives across thirteen alignment tasks. These evaluations encompassed scalar sentiment rewards, strict lexical keyword constraints, demographic and religious debiasing, factual consistency in abstractive summarization, and syntactically valid code generation, spanning model scales from 127 million to 1.5 billion parameters.
The investigation yielded several critical findings. First, there is no universally optimal divergence objective; different objectives establish distinct operational trade-offs between alignment precision and output diversity. Second, the Jensen-Shannon divergence objective consistently strikes the best balance, outperforming the widely used forward Kullback-Leibler divergence by a substantial margin across diverse tasks, even when evaluated against forward divergence metrics. Third, the reverse Kullback-Leibler formulation common in standard reinforcement learning frequently induces mode collapse, sharply reducing sample diversity. Finally, empirical scaling trends show that while increasing model size steadily improves alignment scores, it does not close the performance gap between optimal and suboptimal divergence objectives.
These findings indicate that the mathematical objective chosen to guide model fine-tuning dictates system behavior just as strongly as the training data or model scale. Relying on default alignment objectives can lead to brittle models prone to repetitive outputs or high-variance training dynamics. Selecting appropriate divergence measures allows development teams to mitigate deployment risks, enhance factual reliability in summarization, and preserve creative diversity without incurring additional architectural compute costs.
Organizations developing aligned language models should adopt flexible divergence frameworks and utilize the Jensen-Shannon objective as a robust, high-performing default for general alignment pipelines. When specific deployment contexts demand absolute compliance over variety, reverse divergence variants can be selected deliberately with appropriate baseline adjustments. Future efforts should extend these evaluations to larger frontier architectures exceeding ten billion parameters and explore dynamically adaptive divergence objectives during training.
Confidence in these findings is high across the tested architectures, task suites, and model scales. However, practitioners should note that experiments were conducted on models up to 1.5 billion parameters within simulated preference environments. Further validation in production pipelines featuring complex, multi-turn human feedback loops is recommended before full-scale operational rollout.
- Paper: f-GAN: Training Generative Neural Samplers using Variational Divergence Minimization, Sebastian Nowozin et al. (2016). This paper establishes how variational training can minimize arbitrary f-divergences, the mathematical framework that f-DPG adapts to language-model alignment.
- Paper: Deep reinforcement learning from human preferences, Paul F. Christiano et al. (2017). Its preference-trained reward models and RL policies provide the foundational human-feedback alignment setup that the source recasts through divergence minimization.
- Paper: Fine-Tuning Language Models from Human Preferences, Daniel M. Ziegler et al. (2019). This work develops the KL-constrained PPO approach to language-model preference tuning that the source identifies as RLHF's reverse-KL formulation.
- Paper: Learning to summarize from human feedback, Nisan Stiennon et al. (2020). Its KL-regularized PPO method for optimizing human-preference rewards offers a concrete language-generation example of the RLHF objective analyzed in the source.
No sufficiently relevant recommendations were found.
