keyword
KKT conditions
The Karush-Kuhn-Tucker conditions, commonly abbreviated as KKT conditions, are first-order necessary mathematical conditions for a solution to be optimal in a constrained nonlinear optimization problem involving both equality and inequality constraints. Generalizing the classical method of Lagrange multipliers, they establish four fundamental criteria: stationarity, where the gradient of the Lagrangian function equals zero; primal feasibility, ensuring the candidate solution satisfies all original constraints; dual feasibility, requiring non-negative multipliers for inequality constraints; and complementary slackness, which dictates that an inequality multiplier must be zero whenever its corresponding constraint is strictly inactive. Under standard regularity conditions, these criteria are necessary for local optimality, and in the case of convex optimization problems, they serve as both necessary and sufficient conditions for global optimality.
2 items

A Multi-objective / Multi-task Learning Framework Induced by Pareto Stationarity
Michinari Momma, Chaosheng Dong, Jia Liu
Why you should read this
Develops a generic multi-objective learning framework based on Pareto stationarity that incorporates user preferences and extends weighted Chebyshev optimization to discover models outperforming existing baselines in a single training run.
Multi-objective optimization (MOO) and multi-task learning (MTL) have gained much popularity with prevalent use cases such as production model development of regression / classification / ranking models with MOO, and training deep learning models with MTL. Despite the long history of research in MOO, its application to machine learning requires development of solution strategy, and algorithms have recently been developed to solve specific problems such as discovery of any Pareto optimal (PO) solution, and that with a particular form of preference. In this paper, we develop a novel and generic framework to discover a PO solution with multiple forms of preferences. It allows us to formulate a generic MOO / MTL problem to express a preference, which is solved to achieve both alignment with the preference and PO, at the same time. Specifically, we apply the framework to solve the weighted Chebyshev problem and an extension of that. The former is known as a method to discover the Pareto front, the latter helps to find a model that outperforms an existing model with only one run. Experimental results demonstrate not only the method achieves competitive performance with existing methods, but also it allows us to achieve the performance from different forms of preferences.
Added
2026-10-03

On the Algorithmic Implementation of Multiclass Kernel-based Vector Machines
Koby Crammer, Yoram Singer
Why you should read this
Develops a direct multiclass support vector machine framework based on a generalized margin that reduces large-scale quadratic optimization into small, single-example subproblems solved via a provably convergent fixed-point algorithm.
In this paper we describe the algorithmic implementation of multiclass kernel-based vector machines. Our starting point is a generalized notion of the margin to multiclass problems. Using this notion we cast multiclass categorization problems as a constrained optimization problem with a quadratic objective function. Unlike most of previous approaches which typically decompose a multiclass problem into multiple independent binary classification tasks, our notion of margin yields a direct method for training multiclass predictors. By using the dual of the optimization problem we are able to incorporate kernels with a compact set of constraints and decompose the dual problem into multiple optimization problems of reduced size. We describe an efficient fixed-point algorithm for solving the reduced optimization problems and prove its convergence. We then discuss technical details that yield significant running time improvements for large datasets. Finally, we describe various experiments with our approach comparing it to previously studied kernel-based methods. Our experiments indicate that for multiclass problems we attain state-of-the-art accuracy.
Added
2026-09-15
