keyword
Generalized preference optimization
Generalized preference optimization is a unified algorithmic framework for aligning generative models, such as large language models, with human preferences directly from offline comparison datasets. Instead of relying on separate reward modeling and reinforcement learning loops, it formulates alignment through a broad family of loss functions parameterized by convex functions. This mathematical structure encompasses various direct preference learning algorithms, including direct preference optimization and identity preference optimization, as specific instances within a single theoretical foundation. By altering the chosen convex function, generalized preference optimization characterizes how different offline objectives implicitly enforce regularization relative to a reference policy, enabling systematic analysis and development of diverse offline alignment strategies.
1 item

