Aligning Language Models with Preferences through f-divergence Minimization

Dongyoung GoTomasz KorbakGermán KruszewskiJos RozenNahyeon RyuMarc Dymetman

article2023ICML132 citations

Presents f-DPG, a generalized framework that unifies RLHF and Distributional Policy Gradients by minimizing arbitrary f-divergences against target energy-based models, revealing that objectives like Jensen-Shannon divergence achieve superior alignment and diversity trade-offs compared to traditional KL divergence.

Listen

Aligning large language models with human preferences—such as harmlessness, truthfulness, and stylistic guidelines—is a central requirement for deploying artificial intelligence safely and effectively. Existing alignment techniques typically frame this challenge as steering a generative model toward a desired target distribution of text. However, current practices remain fragmented across disparate algorithms and mathematical objectives, obscuring how the choice of optimization objective affects the final behavior, diversity, and reliability of the aligned model.

The article introduces and evaluates f-DPG, a generalized framework that unifies existing alignment paradigms under the mathematical family of f-divergence minimization. The primary objective is to demonstrate that aligning language models can be decoupled from specific legacy algorithms, allowing practitioners to optimize any evaluating target distribution using diverse mathematical objectives to achieve superior trade-offs between preference adherence and text diversity.

To establish this framework, the authors derived a universal policy gradient formula that supports variance reduction and conditioned generation, encompassing methods like standard reinforcement learning with divergence penalties and generative distributional control. They empirically benchmarked four distinct divergence objectives across thirteen alignment tasks. These evaluations encompassed scalar sentiment rewards, strict lexical keyword constraints, demographic and religious debiasing, factual consistency in abstractive summarization, and syntactically valid code generation, spanning model scales from 127 million to 1.5 billion parameters.

The investigation yielded several critical findings. First, there is no universally optimal divergence objective; different objectives establish distinct operational trade-offs between alignment precision and output diversity. Second, the Jensen-Shannon divergence objective consistently strikes the best balance, outperforming the widely used forward Kullback-Leibler divergence by a substantial margin across diverse tasks, even when evaluated against forward divergence metrics. Third, the reverse Kullback-Leibler formulation common in standard reinforcement learning frequently induces mode collapse, sharply reducing sample diversity. Finally, empirical scaling trends show that while increasing model size steadily improves alignment scores, it does not close the performance gap between optimal and suboptimal divergence objectives.

These findings indicate that the mathematical objective chosen to guide model fine-tuning dictates system behavior just as strongly as the training data or model scale. Relying on default alignment objectives can lead to brittle models prone to repetitive outputs or high-variance training dynamics. Selecting appropriate divergence measures allows development teams to mitigate deployment risks, enhance factual reliability in summarization, and preserve creative diversity without incurring additional architectural compute costs.

Organizations developing aligned language models should adopt flexible divergence frameworks and utilize the Jensen-Shannon objective as a robust, high-performing default for general alignment pipelines. When specific deployment contexts demand absolute compliance over variety, reverse divergence variants can be selected deliberately with appropriate baseline adjustments. Future efforts should extend these evaluations to larger frontier architectures exceeding ten billion parameters and explore dynamically adaptive divergence objectives during training.

Confidence in these findings is high across the tested architectures, task suites, and model scales. However, practitioners should note that experiments were conducted on models up to 1.5 billion parameters within simulated preference environments. Further validation in production pipelines featuring complex, multi-turn human feedback loops is recommended before full-scale operational rollout.

arXiv: 2302.08215

No sufficiently relevant recommendations were found.

Cover for Aligning Language Models with Preferences through f-divergence Minimization

Abstract

Aligning language models with preferences can be posed as approximating a target distribution representing some desired behavior. Existing approaches differ both in the functional form of the target distribution and the algorithm used to approximate it. For instance, Reinforcement Learning from Human Feedback (RLHF) corresponds to minimizing a reverse KL from an implicit target distribution arising from a KL penalty in the objective. On the other hand, Generative Distributional Control (GDC) has an explicit target distribution and minimizes a forward KL from it using the Distributional Policy Gradient (DPG) algorithm. In this paper, we propose a new approach, f-DPG, which allows the use of any f-divergence to approximate any target distribution that can be evaluated. f-DPG unifies both frameworks (RLHF, GDC) and the approximation methods (DPG, RL with KL penalties). We show the practical benefits of various choices of divergence objectives and demonstrate that there is no universally optimal objective but that different divergences present different alignment and diversity trade-offs. We show that Jensen-Shannon divergence strikes a good balance between these objectives, and frequently outperforms forward KL divergence by a wide margin, leading to significant improvements over prior work. These distinguishing characteristics between divergences persist as the model size increases, highlighting the importance of selecting appropriate divergence objectives.

Table of Contents

  • 1. Introduction
  • 2. Background
  • 2.1. Defining a Target Distribution
  • 2.2. Approximating the target distribution
  • 3. Formal Aspects
  • 3.1. f-divergences
  • 3.2. Distributional alignment with f-divergences
  • 3.3. Adding a baseline
  • 3.4. Recovering Some Existing Methods
  • 3.5. Estimating Z
  • 3.6. Conditional Target Distributions
  • 4. Experiments
  • 4.1. Alignment with Scalar Preferences
  • 4.2. Alignment with Lexical Constraints
  • 4.3. Alignment with Distributional Constraints
  • 4.4. Alignment with Conditional Constraints
  • 4.5. Scaling Trends of f-DPG
  • 4.6. Ablation Study
  • 5. Discussion and Conclusion
  • References
  • A. Complements on Formal Aspects and Proofs
  • A.1. Equivalent definitions for f-divergences
  • A.2. Illustrations of a few f-divergences
  • A.3. Proof of Theorem 1
  • A.4. About non-differentiability of f
  • A.5. f-DPG algorithm
  • A.6. Baseline: alternative derivation
  • B. Extended Related Work
  • C. Implementation Details
  • D. Additional Experiments
  • D.1. Generation Quality
  • E. f-DPG on Conditional Target Distributions
  • E.1. Additional Conditional Preferences Experiments and Details
  • F. Optimal Reward Model for a Decision Maker with a Categorical Distribution
  • G. Additional Figures
  • H. Ablation Studies
  • H.1. Matching Other Language Model within Parameter Family
  • H.2. Checking Fluency in Unseen Downstream Task
  • H.3. Ablation Studies on Training Scheme
  • I. Samples

Knowls

  1. Knowl 1 — A sampleable policy-gradient formula for arbitrary f-divergences

    theoretical result

    Let pp be a target distribution and let πθ\pi_\theta be a differentiable, parameterized distribution over a discrete set X\mathcal X, with parameters θ∈Θ\theta\in\Theta. Let f:(0,∞)→Rf:(0,\infty)\to\mathbb R be convex, satisfy f(1)=0f(1)=0, and have derivative f′f'. If either Supp⁡(p)⊆Supp⁡(πθ)\operatorname{Supp}(p)\subseteq\operatorname{Supp}(\pi_\theta) for every θ\theta, or the support of πθ\pi_\theta is independent of θ\theta, then

    ∇θDf(πθ∥p)=Ex∼πθ[f′ ⁣(πθ(x)p(x))∇θlog⁡πθ(x)].\nabla_\theta D_f(\pi_\theta\|p)=\mathbb E_{x\sim\pi_\theta}\left[f'\!\left(\frac{\pi_\theta(x)}{p(x)}\right)\nabla_\theta\log\pi_\theta(x)\right].

    Here DfD_f is the f-divergence generated by ff, and xx is a sample from the model. If p(x)=0p(x)=0 while πθ(x)>0\pi_\theta(x)>0, the derivative term is interpreted as f′(∞)f'(\infty). At points where ff is not differentiable, a subgradient may be used. The formula gives an unbiased, model-sampled gradient estimator without requiring samples from the target.

  2. Knowl 2 — f-DPG minimizes a chosen divergence within a parameterized model family

    model/method

    Given a target probability distribution p(x)p(x) over a discrete text space X\mathcal X and a generative model πθ(x)\pi_\theta(x) with parameters θ∈Θ\theta\in\Theta, f-DPG fits the model by minimizing Df(πθ∥p)D_f(\pi_\theta\|p) over Θ\Theta. For distributions with appropriate support, this divergence is Ex∼p[f(πθ(x)/p(x))]\mathbb E_{x\sim p}[f(\pi_\theta(x)/p(x))]; when πθ\pi_\theta assigns mass where pp is zero, the definition includes the corresponding support correction. The generator ff is convex and satisfies f(1)=0f(1)=0. The model-sampled gradient uses the pseudo-reward f′(πθ(x)/p(x))f'(\pi_\theta(x)/p(x)) multiplying ∇θlog⁡πθ(x)\nabla_\theta\log\pi_\theta(x).

    The paper evaluates four choices: forward KL, Df(πθ∥p)=KL(p∥πθ)D_f(\pi_\theta\|p)=\mathrm{KL}(p\|\pi_\theta), with f(t)=−log⁡tf(t)=-\log t; reverse KL, KL(πθ∥p)\mathrm{KL}(\pi_\theta\|p), with f(t)=tlog⁡tf(t)=t\log t; total variation, with f(t)=12∣1−t∣f(t)=\tfrac12|1-t|; and Jensen–Shannon divergence, with f(t)=tlog⁡2tt+1+log⁡2t+1f(t)=t\log\frac{2t}{t+1}+\log\frac{2}{t+1}. Here t>0t>0 is the ratio of model probability to target probability, and logarithms are natural. If the model family contains pp, all these objectives have the same zero-divergence optimum; if it does not, the best-fitting model can depend on the divergence.

  3. Knowl 3 — f-DPG recovers distributional policy gradients and KL-penalized RL

    model/method

    The f-DPG gradient formulation contains two established language-model alignment procedures as special cases. For a target pp and model πθ\pi_\theta, choosing f(t)=−log⁡tf(t)=-\log t makes Df(πθ∥p)=KL(p∥πθ)D_f(\pi_\theta\|p)=\mathrm{KL}(p\|\pi_\theta), the forward-KL objective used by distributional policy gradients (DPG) for fitting an evaluable target distribution.

    For reward-based alignment, let a(x)a(x) be a reference language model, r(x)r(x) a scalar reward, and β>0\beta>0 a coefficient. Define p(x)∝a(x)exp⁡(r(x)/β)p(x)\propto a(x)\exp(r(x)/\beta). Choosing f(t)=tlog⁡tf(t)=t\log t makes f-DPG minimize KL(πθ∥p)\mathrm{KL}(\pi_\theta\|p), equivalent to maximizing Ex∼πθ[r(x)−βlog⁡(πθ(x)/a(x))]\mathbb E_{x\sim\pi_\theta}[r(x)-\beta\log(\pi_\theta(x)/a(x))]. Thus the framework separates specification of the target from the choice of divergence used to approximate it.

  4. Knowl 4 — No single divergence wins across alignment tasks; Jensen–Shannon often balances alignment and diversity

    empirical result

    Across the paper’s language-model alignment experiments, the divergence selected for optimization materially affected both approximation and generation behavior; the best choice varied with the target, and minimizing one divergence did not guarantee the lowest values of the others. Reverse-KL optimization generally favored strong alignment at the expense of diversity, whereas forward-KL optimization tended toward the contrasting combination of weaker alignment and greater diversity. Jensen–Shannon optimization frequently provided a better balance and often improved on forward-KL DPG across measured divergences. Its pseudo-rewards are smooth in both directions of mismatch, which the authors offer as a plausible explanation for its empirical behavior. Total variation’s thresholded pseudo-reward can be robust to outliers but may have high variance when model and target are already close. These observations are empirical tendencies, not a claim that Jensen–Shannon is universally optimal.

  5. Knowl 5 — f-DPG extends to targets conditioned on contexts

    theoretical result

    For conditional generation, let cc be a context drawn from a distribution τ(c)\tau(c), let pc(x)p_c(x) be the target distribution for that context, and let πθ(x∣c)\pi_\theta(x\mid c) be a parameterized conditional model. f-DPG minimizes the expected conditional divergence

    Ec∼τ[Df(πθ(⋅∣c)∥pc)].\mathbb E_{c\sim\tau}\left[D_f\bigl(\pi_\theta(\cdot\mid c)\|p_c\bigr)\right].

    Under the support conditions for the f-DPG gradient formula, its gradient is

    Ec∼τEx∼πθ(⋅∣c)[f′ ⁣(πθ(x∣c)pc(x))∇θlog⁡πθ(x∣c)].\mathbb E_{c\sim\tau}\mathbb E_{x\sim\pi_\theta(\cdot\mid c)}\left[f'\!\left(\frac{\pi_\theta(x\mid c)}{p_c(x)}\right)\nabla_\theta\log\pi_\theta(x\mid c)\right].

    Here xx is an output, cc is a context, and ff is a convex f-divergence generator with f(1)=0f(1)=0. The extension allows training on contexts sampled from τ\tau without requiring direct sampling from each target pcp_c.

  6. Knowl 6 — Subtracting a constant baseline preserves the f-DPG gradient

    theoretical result

    For a fixed parameter value θ\theta, define the f-DPG pseudo-reward rθ(x)=−f′(πθ(x)/p(x))r_\theta(x)=-f'(\pi_\theta(x)/p(x)). For any constant baseline BB that does not depend on the sampled output xx,

    Ex∼πθ[(rθ(x)−B)∇θlog⁡πθ(x)]=Ex∼πθ[rθ(x)∇θlog⁡πθ(x)].\mathbb E_{x\sim\pi_\theta}\left[(r_\theta(x)-B)\nabla_\theta\log\pi_\theta(x)\right] = \mathbb E_{x\sim\pi_\theta}\left[r_\theta(x)\nabla_\theta\log\pi_\theta(x)\right].

    The equality follows from the zero expectation of the model’s score, so subtracting BB does not bias the gradient estimate. The paper generally uses an exponential-moving-average estimate of the mean pseudo-reward, with weight 0.990.99; for forward-KL DPG it instead uses the analytically computed expectation, equal to 11.

  7. Knowl 7 — Importance sampling estimates the target normalizer from model samples

    model/method

    If an evaluable target is specified by a nonnegative unnormalized function P(x)P(x), its normalized distribution is p(x)=P(x)/Zp(x)=P(x)/Z, where Z=∑x∈XP(x)Z=\sum_{x\in\mathcal X}P(x) is the partition function. When ZZ is unknown, a sample x∼πθx\sim\pi_\theta gives the importance-sampling estimate P(x)/πθ(x)P(x)/\pi_\theta(x), whose expectation under πθ\pi_\theta is ZZ. Averaging these estimates over samples provides an estimate of the normalizer, allowing f-DPG to evaluate the target-to-model probability ratio needed for its pseudo-reward. The paper’s online implementation updates the running mean using each sample and its probability under the model that generated it. In a binary-constraint ablation, the estimate converged to the known normalizer, and using the estimate rather than the true normalizer produced no significant difference in the learned model.

  8. Knowl 8 — Lexical-constraint experiments favor alternatives to forward-KL DPG

    empirical result

    The lexical experiments required generated text to contain one specified word: “amazing” (initial frequency 1×10−31\times10^{-3}), “restaurant” (6×10−46\times10^{-4}), “amusing” (6×10−56\times10^{-5}), or “Wikileaks” (8×10−68\times10^{-6}). The authors tested both a binary-constraint target, proportional to a pretrained model’s probability times an indicator for satisfying the constraint, and a reward-based target in which the word-presence indicator acts as the reward. For the binary target, the target assigns zero probability to texts that omit the word, so reverse KL is infinite and cannot be used.

    For both target constructions, the tested f-DPG methods reduced the measured divergences as training progressed. On the binary-target tasks, total-variation and Jensen–Shannon DPG outperformed forward-KL DPG even when quality was assessed by forward KL; for the reward-based targets, the non-forward-KL objectives also generally improved on forward-KL DPG, with Jensen–Shannon often giving strong word-presence performance. Reverse-KL optimization tended to yield lower normalized entropy, although the authors found no significant difference in sentence-level diversity on the lexical generation metrics.

  9. Knowl 9 — Scalar sentiment alignment exposes a reverse-KL quality–diversity trade-off

    empirical result

    For scalar-preference alignment, the target was p(x)∝a(x)exp⁡(r(x)/β)p(x)\propto a(x)\exp(r(x)/\beta), where aa was GPT-2 small, r(x)=log⁡ϕ(x)r(x)=\log\phi(x), ϕ(x)\phi(x) was the positive-sentiment probability from a DistilBERT classifier, and β=0.1\beta=0.1. Reverse-KL DPG achieved the best reverse-KL value and the highest expected sentiment score among the compared objectives, but it also departed more from the reference model and produced lower entropy. The authors associate this behavior with strong penalties on samples where the target assigns low probability, which can concentrate generation on a subset of the target distribution. Forward-KL, total-variation, and Jensen–Shannon DPG consistently reduced all four tracked divergences in this experiment, with Jensen–Shannon performing best overall. Additional generation metrics likewise indicated lower distribution-level diversity but better individual-sample perplexity for reverse-KL DPG.

  10. Knowl 10 — Distributional debiasing improves regard balance, with objective-specific trade-offs

    empirical result

    The distributional-constraint experiments targeted gender-related prevalence and regard toward Muslims. For gender debiasing, the target moments were 0.50.5 for the feature indicating more female than male pronouns and 11 for the feature indicating the presence of at least one word from a science vocabulary. For regard balancing, the target regard score for Muslim-prompted text was 0.5680.568, matching the score observed for Christians; the initial Muslim-prompted score was 0.3850.385. After training, the Christians-to-Muslims mean regard-score ratio changed from 1:0.6771:0.677 to 1:0.8011:0.801 on average.

    On the gender task, the tested objectives other than forward-KL DPG outperformed that baseline, and reverse-KL DPG best matched the pointwise constraint while showing lower entropy. On the regard task, total-variation DPG had difficulty converging when the starting distribution was already close to the target; the authors attribute this to high variance from its thresholded pseudo-reward. These results show that objective choice affects both constraint satisfaction and diversity or optimization stability.

  11. Knowl 11 — Alignment gains with model scale do not erase divergence-objective differences

    empirical result

    On the scalar sentiment task, the authors increased model size from GPT-2 small (117 million parameters) to GPT-2 XL (1.5 billion parameters) and tracked expected sentiment reward and entropy. Expected reward increased gradually with model size, but the relative performance differences among the divergence objectives persisted rather than disappearing at larger scale. The study therefore found that added model capacity improved alignment but did not by itself close the gap between better- and worse-performing objectives. The authors describe the gradual trend as suggestive that the findings may extend to larger models, rather than as direct evidence from models larger than 1.5 billion parameters.

  12. Knowl 12 — Conditional factuality and compilability constraints improve with f-DPG

    empirical result

    For conditional summarization, the model was trained to produce summaries whose named entities were all present in the source document and that contained at least four named entities. f-DPG increased the fraction of consistent named entities and also improved ROUGE, despite not using reference summaries for training; Jensen–Shannon DPG converged to the target better than forward-KL DPG. In a separate conditional code-generation task, outputs were checked for Python compilability. f-DPG increased the fraction of compilable functions and reduced average PEP8 violations, with Jensen–Shannon DPG again converging better than forward-KL DPG. These experiments used context-conditioned targets and therefore test the conditional extension rather than only unconditional text generation.

Coverage note — Detailed generated-text examples, full training-curve plots, and secondary training-scheme and parameter-capacity ablations were omitted because they are supporting diagnostics rather than additional core methods or findings.

References

  1. 1.Arjovsky, M., Chintala, S., and Bottou, L. Wasserstein generative adversarial networks. In Precup, D. and Teh, Y. W. (eds.), Proc. of ICML, volume 70 of Proceedings of Machine Learning Research, pp. 214–223. PMLR, 2017. URL http://proceedings.mlr.press/v70/arjovsky17a.html.
  2. 2.Askell, A., Bai, Y., Chen, A., Drain, D., Ganguli, D., Henighan, T., Jones, A., Joseph, N., Mann, B., DasSarma, N., et al. A general language assistant as a laboratory for alignment. arXiv preprint arXiv:2112.00861, 2021. URL https://arxiv.org/abs/2112.00861.
  3. 3.Bahdanau, D., Brakel, P., Xu, K., Goyal, A., Lowe, R., Pineau, J., Courville, A. C., and Bengio, Y. An actor-critic algorithm for sequence prediction. In Proc. of ICLR. OpenReview.net, 2017. URL https://openreview.net/forum?id=SJDaqqveg.
  4. 4.Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., Das-Sarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., Joseph, N., Kadavath, S., Kernion, J., Conerly, T., El-Showk, S., Elhage, N., Hatfield-Dodds, Z., Hernandez, D., Hume, T., Johnston, S., Kravec, S., Lovitt, L., Nanda, N., Olsson, C., Amodei, D., Brown, T., Clark, J., McCandlish, S., Olah, C., Mann, B., and Kaplan, J. Training a helpful and harmless assistant with reinforcement learning from human feedback, 2022a. URL https://arxiv.org/abs/2204.05862.
  5. 5.Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., Das-Sarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., et al. Training a helpful and harmless assistant with reinforcement learning from human feedback. ArXiv preprint, abs/2204.05862, 2022b. URL https://arxiv.org/abs/2204.05862.
  6. 6.Baxter, J. and Bartlett, P. L. Infinite-horizon policy-gradient estimation. Journal of Artificial Intelligence Research, 15:319–350, 2001.
  7. 7.Berger, A. L., Della Pietra, S. A., and Della Pietra, V. J. A maximum entropy approach to natural language processing. Computational Linguistics, 22(1):39–71, 1996. URL https://aclanthology.org/J96-1002.
  8. 8.Black, S., Gao, L., Wang, P., Leahy, C., and Biderman, S. GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow, 2021. URL https://doi.org/10.5281/zenodo.5297715.
  9. 9.Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. Language models are few-shot learners. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., and Lin, H. (eds.), Proc. of NeurIPS, 2020. URL https://proceedings.neurips.cc/paper/2020/hash/1457c0d6bfcb4967418bfb8ac142f64a-Abstract.html.
  10. 10.Cao, Y., Sotnikova, A., Daume III, H., Rudinger, R., and Zou, L. Theory-grounded measurement of U.S. social stereotypes in English language models. In Proc. of NAACL-HLT, pp. 1276–1295, Seattle, United States, 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.naacl-main.92. URL https://aclanthology.org/2022.naacl-main.92.
  11. 11.Che, T., Li, Y., Jacob, A. P., Bengio, Y., and Li, W. Mode regularized generative adversarial networks. In Proc. of ICLR. OpenReview.net, 2017. URL https://openreview.net/forum?id=HJKkY35le.
  12. 12.Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, E., Wang, X., Dehghani, M., Brahma, S., et al. Scaling instruction-finetuned language models. ArXiv preprint, abs/2210.11416, 2022. URL https://arxiv.org/abs/2210.11416.
  13. 13.Dathathri, S., Madotto, A., Lan, J., Hung, J., Frank, E., Molino, P., Yosinski, J., and Liu, R. Plug and play language models: A simple approach to controlled text generation. In Proc. of ICLR. OpenReview.net, 2020. URL https://openreview.net/forum?id=H1edEyBKDS.
  14. 14.Dohan, D., Xu, W., Lewkowycz, A., Austin, J., Bieber, D., Lopes, R. G., Wu, Y., Michalewski, H., Saurous, R. A., Sohl-Dickstein, J., et al. Language model cascades. ArXiv preprint, abs/2207.10342, 2022. URL https://arxiv.org/abs/2207.10342.
  15. 15.Eikema, B., Kruszewski, G., Dance, C. R., Elsahar, H., and Dymetman, M. An approximate sampler for energy-based models with divergence diagnostics. Transactions of Machine Learning Research, 2022. URL https://openreview.net/forum?id=VW4IrC0n0M.
  16. 16.Fan, A., Lewis, M., and Dauphin, Y. Hierarchical neural story generation. In Proc. of ACL, pp. 889–898, Melbourne, Australia, 2018. Association for Computational Linguistics. doi: 10.18653/v1/P18-1082. URL https://aclanthology.org/P18-1082.
  17. 17.Gehman, S., Gururangan, S., Sap, M., Choi, Y., and Smith, N. A. RealToxicityPrompts: Evaluating neural toxic degeneration in language models. In Findings of EMNLP, pp. 3356–3369, Online, 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.findings-emnlp.301. URL https://aclanthology.org/2020.findings-emnlp.301.
  18. 18.Ghasemipour, S. K. S., Zemel, R., and Gu, S. A divergence minimization perspective on imitation learning methods. In Kaelbling, L. P., Kragic, D., and Sugiura, K. (eds.), Proceedings of the Conference on Robot Learning, volume 100 of Proceedings of Machine Learning Research, pp. 1259–1277. PMLR, 2020. URL https://proceedings.mlr.press/v100/ghasemipour20a.html.
  19. 19.Glaese, A., McAleese, N., Trebacz, M., Aslanides, J., Firoiu, V., Ewalds, T., Rauh, M., Weidinger, L., Chadwick, M., Thacker, P., Campbell-Gillingham, L., Uesato, J., Huang, P., Comanescu, R., Yang, F., See, A., Dathathri, S., Greig, R., Chen, C., Fritz, D., Elias, J. S., Green, R., Mokra, S., Fernando, N., Wu, B., Foley, R., Young, S., Gabriel, I., Isaac, W., Mellor, J., Hassabis, D., Kavukcuoglu, K., Hendricks, L. A., and Irving, G. Improving alignment of dialogue agents via targeted human judgements. CoRR, abs/2209.14375, 2022. doi: 10.48550/arXiv.2209.14375. URL https://doi.org/10.48550/arXiv.2209.14375.
  20. 20.Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. Generative adversarial networks. Commun. ACM, 63(11):139–144, 2020. ISSN 0001-0782. doi: 10.1145/3422622. URL https://doi.org/10.1145/3422622.
  21. 21.Goyal, K., Dyer, C., and Berg-Kirkpatrick, T. Exposing the implicit energy networks behind masked language models via metropolis–hastings. In Proc. of ICLR. OpenReview.net, 2022. URL https://openreview.net/forum?id=6PvWo1kEvlT.
  22. 22.HF Canonical Model Maintainers. distilbert-base-uncased-finetuned-sst-2-english (revision bfdd146), 2022. URL https://huggingface.co/distilbert-base-uncased-finetuned-sst-2-english.
  23. 23.Hiriart-Urruty, J.-B. and Lemarechal, C. Convex analysis and minimization algorithms I: Fundamentals, volume 305. Springer science & business media, 2013.
  24. 24.Huszar, F. How (not) to train your generative model: Scheduled sampling, likelihood, adversary? ArXiv preprint, abs/1511.05101, 2015. URL https://arxiv.org/abs/1511.05101.
  25. 25.Jaques, N., Gu, S., Bahdanau, D., Hernandez-Lobato, J. M., Turner, R. E., and Eck, D. Sequence tutor: Conservative fine-tuning of sequence generation models with kl-control. In Precup, D. and Teh, Y. W. (eds.), Proc. of ICML, volume 70 of Proceedings of Machine Learning Research, pp. 1645–1654. PMLR, 2017. URL http://proceedings.mlr.press/v70/jaques17a.html.
  26. 26.Jaques, N., Ghandeharioun, A., Shen, J. H., Ferguson, C., Lapedriza, A., Jones, N., Gu, S., and Picard, R. W. Way off-policy batch deep reinforcement learning of implicit human preferences in dialog. ArXiv preprint, abs/1907.00456, 2019. URL https://arxiv.org/abs/1907.00456.
  27. 27.Kappen, H. J., Gomez, V., and Opper, M. Optimal control as a graphical model inference problem. Machine learning, 87(2):159–182, 2012.
  28. 28.Kappen, H. J., Gomez, V., and Opper, M. Optimal control as a graphical model inference problem. In Borrajo, D., Kambhampati, S., Oddi, A., and Fratini, S. (eds.), Proceedings of the Twenty-Third International Conference on Automated Planning and Scheduling, ICAPS 2013, Rome, Italy, June 10-14, 2013. AAAI, 2013. URL http://www.aaai.org/ocs/index.php/ICAPS/ICAPS13/paper/view/6012.
  29. 29.Ke, L., Choudhury, S., Barnes, M., Sun, W., Lee, G., and Srinivasa, S. S. Imitation learning as f-divergence minimization. In LaValle, S. M., Lin, M., Ojala, T., Shell, D. A., and Yu, J. (eds.), Algorithmic Foundations of Robotics XIV, Proceedings of the Fourteenth Workshop on the Algorithmic Foundations of Robotics, WAFR 2021, Oulu, Finland, June 21-23, 2021, volume 17 of Springer Proceedings in Advanced Robotics, pp. 313–329. Springer, 2021. doi: 10.1007/978-3-030-66723-8_19. URL https://doi.org/10.1007/978-3-030-66723-8_19.
  30. 30.Khalifa, M., Elsahar, H., and Dymetman, M. A distributional approach to controlled text generation. In Proc. of ICLR. OpenReview.net, 2021. URL https://openreview.net/forum?id=jWkw45-9AbL.
  31. 31.Kingma, D. P. and Ba, J. Adam: A method for stochastic optimization. In Bengio, Y. and LeCun, Y. (eds.), Proc. of ICLR, 2015. URL http://arxiv.org/abs/1412.6980.
  32. 32.Korbak, T., Elsahar, H., Kruszewski, G., and Dymetman, M. Controlling conditional language models without catastrophic forgetting. In Chaudhuri, K., Jegelka, S., Song, L., Szepesvari, C., Niu, G., and Sabato, S. (eds.), Proc. of ICML, volume 162 of Proceedings of Machine Learning Research, pp. 11499–11528. PMLR, 2022a. URL https://proceedings.mlr.press/v162/korbak22a.html.
  33. 33.Korbak, T., Elsahar, H., Kruszewski, G., and Dymetman, M. On reinforcement learning and distribution matching for fine-tuning language models with no catastrophic forgetting. In Oh, A. H., Agarwal, A., Belgrave, D., and Cho, K. (eds.), Proc. of NeurIPS, 2022b. URL https://openreview.net/forum?id=XvI6h-s4un.
  34. 34.Korbak, T., Perez, E., and Buckley, C. L. RL with KL penalties is better viewed as bayesian inference. CoRR, abs/2205.11275, 2022c. doi: 10.48550/arXiv.2205.11275. URL https://doi.org/10.48550/arXiv.2205.11275.
  35. 35.Lebret, R., Grangier, D., and Auli, M. Neural text generation from structured data with application to the biography domain. In Proc. of EMNLP, pp. 1203–1213, Austin, Texas, 2016. Association for Computational Linguistics. doi: 10.18653/v1/D16-1128. URL https://aclanthology.org/D16-1128.
  36. 36.LeCun, Y., Chopra, S., Hadsell, R., Ranzato, M., and Huang, F. A tutorial on energy-based learning. Predicting structured data, 1(0), 2006.
  37. 37.Levine, S. Reinforcement learning and control as probabilistic inference: Tutorial and review. ArXiv preprint, abs/1805.00909, 2018. URL https://arxiv.org/abs/1805.00909.
  38. 38.Li, J., Galley, M., Brockett, C., Gao, J., and Dolan, B. A diversity-promoting objective function for neural conversation models. In Proc. of NAACL-HLT, pp. 110–119, San Diego, California, 2016a. Association for Computational Linguistics. doi: 10.18653/v1/N16-1014. URL https://aclanthology.org/N16-1014.
  39. 39.Li, J., Monroe, W., Ritter, A., Jurafsky, D., Galley, M., and Gao, J. Deep reinforcement learning for dialogue generation. In Proc. of EMNLP, pp. 1192–1202, Austin, Texas, 2016b. Association for Computational Linguistics. doi: 10.18653/v1/D16-1127. URL https://aclanthology.org/D16-1127.
  40. 40.Liese, F. and Vajda, I. On divergences and informations in statistics and information theory. IEEE Transactions on Information Theory, 52(10):4394–4412, 2006.
  41. 41.Lin, C.-Y. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out, pp. 74–81, Barcelona, Spain, 2004. Association for Computational Linguistics. URL https://aclanthology.org/W04-1013.
  42. 42.Lin, S., Hilton, J., and Evans, O. TruthfulQA: Measuring how models mimic human falsehoods. In Proc. of ACL, pp. 3214–3252, Dublin, Ireland, 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.acl-long.229. URL https://aclanthology.org/2022.acl-long.229.
  43. 43.Liu, C.-W., Lowe, R., Serban, I., Noseworthy, M., Charlin, L., and Pineau, J. How NOT to evaluate your dialogue system: An empirical study of unsupervised evaluation metrics for dialogue response generation. In Proc. of EMNLP, pp. 2122–2132, Austin, Texas, 2016. Association for Computational Linguistics. doi: 10.18653/v1/D16-1230. URL https://aclanthology.org/D16-1230.
  44. 44.Maas, A. L., Daly, R. E., Pham, P. T., Huang, D., Ng, A. Y., and Potts, C. Learning word vectors for sentiment analysis. In Proc. of ACL, pp. 142–150, Portland, Oregon, USA, 2011. Association for Computational Linguistics. URL https://aclanthology.org/P11-1015.
  45. 45.Maynez, J., Narayan, S., Bohnet, B., and McDonald, R. On faithfulness and factuality in abstractive summarization. In Proc. of ACL, pp. 1906–1919, Online, 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.acl-main.173. URL https://aclanthology.org/2020.acl-main.173.
  46. 46.Menick, J., Trebacz, M., Mikulik, V., Aslanides, J., Song, F., Chadwick, M., Glaese, M., Young, S., Campbell-Gillingham, L., Irving, G., and McAleese, N. Teaching language models to support answers with verified quotes, 2022. URL https://arxiv.org/abs/2203.11147.
  47. 47.Mescheder, L. M., Geiger, A., and Nowozin, S. Which training methods for gans do actually converge? In Dy, J. G. and Krause, A. (eds.), Proc. of ICML, volume 80 of Proceedings of Machine Learning Research, pp. 3478–3487. PMLR, 2018. URL http://proceedings.mlr.press/v80/mescheder18a.html.
  48. 48.Miao, N., Zhou, H., Mou, L., Yan, R., and Li, L. CGMH: constrained sentence generation by metropolis-hastings sampling. In The Thirty-Third AAAI Conference on Artificial Intelligence, AAAI 2019, The Thirty-First Innovative Applications of Artificial Intelligence Conference, IAAI 2019, The Ninth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2019, Honolulu, Hawaii, USA, January 27 - February 1, 2019, pp. 6834–6842. AAAI Press, 2019. doi: 10.1609/aaai.v33i01.33016834. URL https://doi.org/10.1609/aaai.v33i01.33016834.
  49. 49.Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M. A. Playing atari with deep reinforcement learning. CoRR, abs/1312.5602, 2013. URL http://arxiv.org/abs/1312.5602.
  50. 50.Nallapati, R., Zhou, B., dos Santos, C., Gulcehre, C., and Xiang, B. Abstractive text summarization using sequence-to-sequence RNNs and beyond. In Proceedings of the 20th SIGNLL Conference on Computational Natural Language Learning, pp. 280–290, Berlin, Germany, 2016. Association for Computational Linguistics. doi: 10.18653/v1/K16-1028. URL https://aclanthology.org/K16-1028.
  51. 51.Nan, F., Nallapati, R., Wang, Z., Nogueira dos Santos, C., Zhu, H., Zhang, D., McKeown, K., and Xiang, B. Entity-level factual consistency of abstractive text summarization. In Proc. of EACL, pp. 2727–2733, Online, 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.eacl-main.235. URL https://aclanthology.org/2021.eacl-main.235.
  52. 52.Ngo, H., Raterink, C., Araujo, J. G., Zhang, I., Chen, C., Morisot, A., and Frosst, N. Mitigating harm in language models with conditional-likelihood filtration. ArXiv preprint, abs/2108.07790, 2021. URL https://arxiv.org/abs/2108.07790.
  53. 53.Norouzi, M., Bengio, S., Chen, Z., Jaitly, N., Schuster, M., Wu, Y., and Schuurmans, D. Reward augmented maximum likelihood for neural structured prediction. In Lee, D. D., Sugiyama, M., von Luxburg, U., Guyon, I., and Garnett, R. (eds.), Proc. of NeurIPS, pp. 1723–1731, 2016. URL https://proceedings.neurips.cc/paper/2016/hash/2f885d0fbe2e131bfc9d98363e55d1d4-Abstract.html.
  54. 54.Nowozin, S., Cseke, B., and Tomioka, R. f-gan: Training generative neural samplers using variational divergence minimization. In Lee, D. D., Sugiyama, M., von Luxburg, U., Guyon, I., and Garnett, R. (eds.), Proc. of NeurIPS, pp. 271–279, 2016. URL https://proceedings.neurips.cc/paper/2016/hash/cedebb6e872f539bef8c3f919874e9d7-Abstract.html.
  55. 55.Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Gray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R. Training language models to follow instructions with human feedback. In Oh, A. H., Agarwal, A., Belgrave, D., and Cho, K. (eds.), Proc. of NeurIPS, 2022. URL https://openreview.net/forum?id=TG8KACxEON.
  56. 56.Parshakova, T., Andreoli, J.-M., and Dymetman, M. Distributional reinforcement learning for energy-based sequential models. ArXiv preprint, abs/1912.08517, 2019. URL https://arxiv.org/abs/1912.08517.
  57. 57.Pasunuru, R. and Bansal, M. Reinforced video captioning with entailment rewards. In Proc. of EMNLP, pp. 979–985, Copenhagen, Denmark, 2017. Association for Computational Linguistics. doi: 10.18653/v1/D17-1103. URL https://aclanthology.org/D17-1103.
  58. 58.Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S. Pytorch: An imperative style, high-performance deep learning library. In Wallach, H. M., Larochelle, H., Beygelzimer, A., d’Alche-Buc, F., Fox, E. B., and Garnett, R. (eds.), Proc. of NeurIPS, pp. 8024–8035, 2019. URL https://proceedings.neurips.cc/paper/2019/hash/bdbca288fee7f92f2bfa9f7012727740-Abstract.html.
  59. 59.Paulus, R., Xiong, C., and Socher, R. A deep reinforced model for abstractive summarization. In Proc. of ICLR. OpenReview.net, 2018. URL https://openreview.net/forum?id=HkAClQgA-.
  60. 60.Peters, J. and Schaal, S. Reinforcement learning by reward-weighted regression for operational space control. In Ghahramani, Z. (ed.), Proc. of ICML, volume 227 of ACM International Conference Proceeding Series, pp. 745–750. ACM, 2007. doi: 10.1145/1273496.1273590. URL https://doi.org/10.1145/1273496.1273590.
  61. 61.Polyanskiy, Y. f-divergences, 2019. URL https://people.lids.mit.edu/yp/homepage/data/LN_fdiv.pdf.
  62. 62.Qin, L., Welleck, S., Khashabi, D., and Choi, Y. Cold decoding: Energy-based constrained text generation with langevin dynamics. ArXiv preprint, abs/2202.11705, 2022. URL https://arxiv.org/abs/2202.11705.
  63. 63.Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019.
  64. 64.Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., Liu, P. J., et al. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21(140):1–67, 2020.
  65. 65.Ranzato, M., Chopra, S., Auli, M., and Zaremba, W. Sequence level training with recurrent neural networks. In Bengio, Y. and LeCun, Y. (eds.), Proc. of ICLR, 2016. URL http://arxiv.org/abs/1511.06732.
  66. 66.Raychev, V., Bielik, P., and Vechev, M. Probabilistic model for code with decision trees. ACM SIGPLAN Notices, 51(10):731–747, 2016.
  67. 67.Rockafellar, R. T. Convex analysis, volume 18. Princeton university press, 1970.
  68. 68.Roller, S., Dinan, E., Goyal, N., Ju, D., Williamson, M., Liu, Y., Xu, J., Ott, M., Smith, E. M., Boureau, Y.-L., and Weston, J. Recipes for building an open-domain chatbot. In Proc. of EACL, pp. 300–325, Online, 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.eacl-main.24. URL https://aclanthology.org/2021.eacl-main.24.
  69. 69.Sason, I. On f-divergences: Integral representations, local behavior, and inequalities. Entropy, 20(5):383, 2018.
  70. 70.Sason, I. and Verdu, S. f-divergence inequalities. IEEE Transactions on Information Theory, 62(11):5973–6006, 2016.
  71. 71.Saunders, W., Yeh, C., Wu, J., Bills, S., Ouyang, L., Ward, J., and Leike, J. Self-critiquing models for assisting human evaluators. ArXiv preprint, abs/2206.05802, 2022. URL https://arxiv.org/abs/2206.05802.
  72. 72.Scheurer, J., Campos, J. A., Chan, J. S., Chen, A., Cho, K., and Perez, E. Training language models with natural language feedback. ArXiv preprint, abs/2204.14146, 2022. URL https://arxiv.org/abs/2204.14146.
  73. 73.Schulman, J., Moritz, P., Levine, S., Jordan, M. I., and Abbeel, P. High-dimensional continuous control using generalized advantage estimation. In Bengio, Y. and LeCun, Y. (eds.), Proc. of ICLR, 2016. URL http://arxiv.org/abs/1506.02438.
  74. 74.Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. Proximal policy optimization algorithms. ArXiv preprint, abs/1707.06347, 2017. URL https://arxiv.org/abs/1707.06347.
  75. 75.Sheng, E., Chang, K.-W., Natarajan, P., and Peng, N. The woman worked as a babysitter: On biases in language generation. In Proc. of EMNLP, pp. 3407–3412, Hong Kong, China, 2019. Association for Computational Linguistics. doi: 10.18653/v1/D19-1339. URL https://aclanthology.org/D19-1339.
  76. 76.Solaiman, I. and Dennison, C. Process for adapting language models to society (palms) with values-targeted datasets. Proc. of NeurIPS, 34:5861–5873, 2021.
  77. 77.Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., Kluska, A., Lewkowycz, A., Agarwal, A., Power, A., Ray, A., Warstadt, A., Kocurek, A. W., Safaya, A., Tazarv, A., Xiang, A., Parrish, A., Nie, A., Hussain, A., Askell, A., Dsouza, A., Rahane, A., Iyer, A. S., Andreassen, A., Santilli, A., Stuhlmuller, A., Dai, A. M., La, A., Lampinen, A. K., Zou, A., Jiang, A., Chen, A., Vuong, A., Gupta, A., Gottardi, A., Norelli, A., Venkatesh, A., Gholamidavoodi, A., Tabassum, A., Menezes, A., Kirubarajan, A., Mullokandov, A., Sabharwal, A., Herrick, A., Efrat, A., Erdem, A., Karakas, A., and et al. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. CoRR, abs/2206.04615, 2022. doi: 10.48550/arXiv.2206.04615. URL https://doi.org/10.48550/arXiv.2206.04615.
  78. 78.Stiennon, N., Ouyang, L., Wu, J., Ziegler, D. M., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. Learning to summarize from human feedback. In Proc. of NeurIPS, NIPS’20, Red Hook, NY, USA, 2020. Curran Associates Inc. ISBN 9781713829546.
  79. 79.Tambwekar, P., Dhuliawala, M., Martin, L. J., Mehta, A., Harrison, B., and Riedl, M. O. Controllable neural story plot generation via reward shaping. In Kraus, S. (ed.), Proc. of IJCAI, pp. 5982–5988. ijcai.org, 2019. doi: 10.24963/ijcai.2019/829. URL https://doi.org/10.24963/ijcai.2019/829.
  80. 80.Theis, L., van den Oord, A., and Bethge, M. A note on the evaluation of generative models. In Bengio, Y. and LeCun, Y. (eds.), Proc. of ICLR, 2016. URL http://arxiv.org/abs/1511.01844.
  81. 81.Thoppilan, R., Freitas, D. D., Hall, J., Shazeer, N., Kulshreshtha, A., Cheng, H., Jin, A., Bos, T., Baker, L., Du, Y., Li, Y., Lee, H., Zheng, H. S., Ghafouri, A., Menegali, M., Huang, Y., Krikun, M., Lepikhin, D., Qin, J., Chen, D., Xu, Y., Chen, Z., Roberts, A., Bosma, M., Zhou, Y., Chang, C., Krivokon, I., Rusch, W., Pickett, M., Meier-Hellstern, K. S., Morris, M. R., Doshi, T., Santos, R. D., Duke, T., Soraker, J., Zevenbergen, B., Prabhakaran, V., Diaz, M., Hutchinson, B., Olson, K., Molina, A., Hoffman-John, E., Lee, J., Aroyo, L., Rajakumar, R., Butryna, A., Lamm, M., Kuzmina, V., Fenton, J., Cohen, A., Bernstein, R., Kurzweil, R., Aguera-Arcas, B., Cui, C., Croak, M., Chi, E. H., and Le, Q. Lamda: Language models for dialog applications. ArXiv preprint, abs/2201.08239, 2022. URL https://arxiv.org/abs/2201.08239.
  82. 82.Todorov, E. Linearly-solvable markov decision problems. In Scholkopf, B., Platt, J. C., and Hofmann, T. (eds.), Proc. of NeurIPS, pp. 1369–1376. MIT Press, 2006a. URL https://proceedings.neurips.cc/paper/2006/hash/d806ca13ca3449af72a1ea5aedbed26a-Abstract.html.
  83. 83.Todorov, E. Linearly-solvable markov decision problems. In Scholkopf, B., Platt, J. C., and Hofmann, T. (eds.), Proc. of NeurIPS, pp. 1369–1376. MIT Press, 2006b. URL https://proceedings.neurips.cc/paper/2006/hash/d806ca13ca3449af72a1ea5aedbed26a-Abstract.html.
  84. 84.Wang, D., Liu, H., and Liu, Q. Variational inference with tail-adaptive f-divergence. In Bengio, S., Wallach, H. M., Larochelle, H., Grauman, K., Cesa-Bianchi, N., and Garnett, R. (eds.), Proc. of NeurIPS, pp. 5742–5752, 2018. URL https://proceedings.neurips.cc/paper/2018/hash/1cd138d0499a68f4bb72bee04bbec2d7-Abstract.html.
  85. 85.Weidinger, L., Mellor, J., Rauh, M., Griffin, C., Uesato, J., Huang, P., Cheng, M., Glaese, M., Balle, B., Kasirzadeh, A., Kenton, Z., Brown, S., Hawkins, W., Stepleton, T., Biles, C., Birhane, A., Haas, J., Rimell, L., Hendricks, L. A., Isaac, W. S., Legassick, S., Irving, G., and Gabriel, I. Ethical and social risks of harm from language models. ArXiv preprint, abs/2112.04359, 2021. URL https://arxiv.org/abs/2112.04359.
  86. 86.Welbl, J., Glaese, A., Uesato, J., Dathathri, S., Mellor, J., Hendricks, L. A., Anderson, K., Kohli, P., Coppin, B., and Huang, P.-S. Challenges in detoxifying language models. In Findings of EMNLP, pp. 2447–2469, Punta Cana, Dominican Republic, 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.findings-emnlp.210. URL https://aclanthology.org/2021.findings-emnlp.210.
  87. 87.Williams, R. J. Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine learning, 8(3):229–256, 1992.
  88. 88.Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Davison, J., Shleifer, S., von Platen, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Le Scao, T., Gugger, S., Drame, M., Lhoest, Q., and Rush, A. Transformers: State-of-the-art natural language processing. In Proc. of EMNLP, pp. 38–45, Online, 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.emnlp-demos.6. URL https://aclanthology.org/2020.emnlp-demos.6.
  89. 89.Xu, J., Ju, D., Li, M., Boureau, Y.-L., Weston, J., and Dinan, E. Bot-adversarial dialogue for safe conversational agents. In Proc. of NAACL-HLT, pp. 2950–2968, Online, 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.naacl-main.235. URL https://aclanthology.org/2021.naacl-main.235.
  90. 90.Zelikman, E., Wu, Y., Mu, J., and Goodman, N. STar: Bootstrapping reasoning with reasoning. In Oh, A. H., Agarwal, A., Belgrave, D., and Cho, K. (eds.), Proc. of NeurIPS, 2022. URL https://openreview.net/forum?id=_3ELRdg2sgI.
  91. 91.Zhao, J., Khashabi, D., Khot, T., Sabharwal, A., and Chang, K.-W. Ethical-advice taker: Do language models understand natural language interventions? In Findings of ACL, pp. 4158–4164, Online, 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.findings-acl.364. URL https://aclanthology.org/2021.findings-acl.364.
  92. 92.Zhu, Y., Lu, S., Zheng, L., Guo, J., Zhang, W., Wang, J., and Yu, Y. Texygen: A benchmarking platform for text generation models. In Collins-Thompson, K., Mei, Q., Davison, B. D., Liu, Y., and Yilmaz, E. (eds.), Proc. of SIGIR, pp. 1097–1100. ACM, 2018. doi: 10.1145/3209978.3210080. URL https://doi.org/10.1145/3209978.3210080.
  93. 93.Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G. Fine-tuning language models from human preferences. ArXiv preprint, abs/1909.08593, 2019. URL https://arxiv.org/abs/1909.08593.
  94. 94.Ziegler, D. M., Nix, S., Chan, L., Bauman, T., Schmidt-Nielsen, P., Lin, T., Scherlis, A., Nabeshima, N., Weinstein-Raun, B., de Haas, D., Shlegeris, B., and Thomas, N. Adversarial training for high-stakes reliability. CoRR, abs/2205.01663, 2022. doi: 10.48550/arXiv.2205.01663. URL https://doi.org/10.48550/arXiv.2205.01663.

Citation

MLA
Go, D., et al. “Aligning Language Models with Preferences Through $f$-divergence Minimization”. International Conference on Machine Learning, vol. 202, 2023, pp. 11546–83, https://proceedings.mlr.press/v202/go23a.html.
APA
Go, D., Korbak, T., Kruszewski, G., Rozen, J., Ryu, N., & Dymetman, M. (2023). Aligning Language Models with Preferences through $f$-divergence Minimization. International Conference on Machine Learning, 202, 11546–11583. https://proceedings.mlr.press/v202/go23a.html
Chicago
Go, D., T. Korbak, G. Kruszewski, J. Rozen, N. Ryu, and M. Dymetman. 2023. “Aligning Language Models with Preferences Through $f$-divergence Minimization”. International Conference on Machine Learning 202: 11546–83. https://proceedings.mlr.press/v202/go23a.html.
Harvard
Go, D. et al. (2023) “Aligning Language Models with Preferences through $f$-divergence Minimization”, International Conference on Machine Learning. PMLR, pp. 11546–11583. Available at: https://proceedings.mlr.press/v202/go23a.html.
Vancouver
1. Go D, Korbak T, Kruszewski G, Rozen J, Ryu N, Dymetman M (2023) Aligning Language Models with Preferences through $f$-divergence Minimization. In: International Conference on Machine Learning. PMLR, pp 11546–11583

BibTeX

@InProceedings{pmlr-v202-go23a,
  title = 	 {Aligning Language Models with Preferences through $f$-divergence Minimization},
  author =       {Go, Dongyoung and Korbak, Tomasz and Kruszewski, Germ\`{a}n and Rozen, Jos and Ryu, Nahyeon and Dymetman, Marc},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {11546--11583},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/go23a/go23a.pdf},
  url = 	 {https://proceedings.mlr.press/v202/go23a.html},
  abstract = 	 {Aligning language models with preferences can be posed as approximating a target distribution representing some desired behavior. Existing approaches differ both in the functional form of the target distribution and the algorithm used to approximate it. For instance, Reinforcement Learning from Human Feedback (RLHF) corresponds to minimizing a reverse KL from an implicit target distribution arising from a KL penalty in the objective. On the other hand, Generative Distributional Control (GDC) has an explicit target distribution and minimizes a forward KL from it using the Distributional Policy Gradient (DPG) algorithm. In this paper, we propose a new approach, $f$-DPG, which allows the use of any $f$-divergence to approximate any target distribution that can be evaluated. $f$-DPG unifies both frameworks (RLHF, GDC) and the approximation methods (DPG, RL with KL penalties). We show the practical benefits of various choices of divergence objectives and demonstrate that there is no universally optimal objective but that different divergences present different alignment and diversity trade-offs. We show that Jensen-Shannon divergence strikes a good balance between these objectives, and frequently outperforms forward KL divergence by a wide margin, leading to significant improvements over prior work. These distinguishing characteristics between divergences persist as the model size increases, highlighting the importance of selecting appropriate divergence objectives.}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/