Built independently by an author, for readers. Read the story and support ChapterPal

keyword

sequence likelihood calibration

Sequence likelihood calibration is a training and alignment method for language models designed to ensure that the probability assigned to a generated sequence directly reflects its relative quality or human preference ranking. While standard maximum likelihood estimation trains models to predict text on a token-by-token basis, the resulting full-sequence probabilities often fail to rank higher-quality outputs consistently above lower-quality alternatives. Sequence likelihood calibration addresses this discrepancy by applying a ranking or margin-based loss over candidate generations, explicitly training the model to assign higher likelihood to preferred or reference sequences relative to dispreferred ones. By directly aligning sequence-level probabilities with desired outputs, this technique serves as an effective post-training and preference optimization framework without requiring complex reinforcement learning loops or separate reward models.

2 items

Human Alignment of Large Language Models through Online Preference Optimisation

Human Alignment of Large Language Models through Online Preference Optimisation

Daniele Calandriello, Zhaohan Daniel Guo, Rémi Munos, Mark Rowland, Yunhao Tang, Bernardo Ávila Pires, Pierre Harvey Richemond, Charline Le Lan, Michal Valko, Tianqi Liu, Rishabh Joshi, Zeyu Zheng, Bilal Piot

OrganizationsGoogle

Why you should read this

Establishes a theoretical equivalence between offline Identity Preference Optimisation and online Nash Mirror Descent, introducing a unified algorithm that combines online self-play with regularized mixture sampling to improve language model alignment.

Ensuring alignment of language models’ outputs with human preferences is critical to guarantee a useful, safe, and pleasant user experience. Thus, human alignment has been extensively studied recently and several methods such as Reinforcement Learning from Human Feedback (RLHF), Direct Preference Optimisation (DPO) and Sequence Likelihood Calibration (SLiC) have emerged. In this paper, our contribution is two-fold. First, we show the equivalence between two recent alignment methods, namely Identity Preference Optimisation (IPO) and Nash Mirror Descent (Nash-MD). Second, we introduce a generalisation of IPO, named IPO-MD, that leverages the regularised sampling approach proposed by Nash-MD. This equivalence may seem surprising at first sight, since IPO is an offline method whereas Nash-MD is an online method using a preference model. However, this equivalence can be proven when we consider the online version of IPO, that is when both generations are sampled by the online policy and annotated by a trained preference model. Optimising the IPO loss with such a stream of data becomes then equivalent to finding the Nash equilibrium of the preference model through self-play. Building on this equivalence, we introduce the IPO-MD algorithm that generates data with a mixture policy (between the online and reference policy) similarly as the general Nash-MD algorithm. We compare online-IPO and IPO-MD to different online versions of existing losses on preference data such as DPO and SLiC on a summarisation task.

Added

2026-09-26