Multiple Kernel Learning Algorithms

Mehmet GönenEthem Alpaydin

article2011JMLR2,004 citations

Presents a comprehensive taxonomy and empirical comparison of multiple kernel learning algorithms across six key dimensions to guide the selection of combination methods based on computational complexity, solution sparsity, and kernel types.

Listen

Organizations often deploy machine learning models that must integrate disparate data streams, such as varying feature representations, distinct modalities, or differing similarity metrics. Selecting a single kernel function for a support vector machine can introduce substantial bias and fail to capture complementary information across heterogeneous sources. Multiple kernel learning addresses this challenge by automating the selection and combination of multiple similarity measures.

The article establishes a systematic framework to classify existing multiple kernel learning methods and evaluates their real-world trade-offs in classification accuracy, computational complexity, and model size. To achieve this, the authors categorized algorithms across six architectural dimensions and conducted standardized empirical benchmarks across four real-world datasets spanning bioinformatics, digit recognition, and web advertisement classification.

The empirical analysis yielded several critical findings regarding algorithm design. First, combining multiple kernels systematically outperformed single-kernel baselines across the board. Second, when simple linear kernels were combined, nonlinear and data-dependent methods consistently achieved the highest accuracies—reaching over 99 percent on handwritten digit benchmarks and significantly outperforming simple averaging. Third, when combining complex Gaussian kernels, trained combinations failed to provide statistically significant accuracy gains over a basic unweighted average of kernels. Fourth, while nonlinear kernel combinations improved predictive accuracy, they significantly increased model complexity by storing up to 90 to 100 percent of instances as support vectors. Finally, iterative two-step algorithms, particularly localized models, required dozens to over one hundred optimization solver calls, creating significant computational overhead compared to direct one-step methods.

These findings indicate that system architects should tailor their kernel combination strategy to the complexity of the underlying base representations. When baseline features rely on simple linear similarities, investing computational budget into localized or nonlinear combiners delivers marked accuracy improvements. Conversely, when features already incorporate complex nonlinear mappings, simple unweighted averaging is often sufficient and avoids unnecessary training overhead and deployment latency. Decision-makers must balance accuracy against operational costs, as highly nonlinear combinations increase inference memory and storage footprints due to larger support vector retention.

Teams implementing multiple kernel learning should select algorithms based on deployment constraints: use unweighted averaging as a fast, robust baseline; deploy localized methods when minimizing stored support vectors is essential; and utilize group Lasso-based formulations when training time must remain low. Future research should prioritize developing faster training routines for localized models and refining automated kernel pruning to reduce memory usage in resource-constrained environments.

Confidence in these conclusions is high for binary classification tasks on moderate-sized tabular, image, and biological datasets. However, practitioners should exercise caution when scaling these findings to massive streaming datasets, complex multiclass domains, or real-time systems where high support vector retention and lengthy solver convergence could cause performance bottlenecks.

  • Paper: Learning the Kernel Matrix with Semidefinite Programming, Gert R. G. Lanckriet et al. (2004). This foundational paper establishes the semidefinite programming framework for learning kernel matrices and linear combinations from data, serving as the primary baseline and theoretical starting point for multiple kernel learning algorithms.
  • Paper: Choosing Multiple Parameters for Support Vector Machines, OLIVIER CHAPELLE et al. (2002). This work introduces gradient-descent optimization for tuning multiple kernel and scaling parameters in support vector machines, establishing the continuous parameter-tuning paradigm surveyed in the source.
  • Paper: Support-vector networks, Corinna Cortes et al. (1995). This classic paper formalizes support vector networks, providing the core large-margin classification architecture that multiple kernel learning methods generalize and combine.
  • Paper: A training algorithm for optimal margin classifiers, Bernhard E. Boser et al. (1992). This work formulates the optimal margin hyperplane and dual quadratic optimization problem foundational to kernel machines.
  • Paper: On Combining Classifiers, Josef Kittler et al. (1998). This paper develops the theoretical framework for combining classifiers across distinct data representations, providing the decision-fusion principles behind kernel combination strategies.
  • Paper: Statistical Comparisons of Classifiers over Multiple Data Sets, Janez Demšar (2006). This text provides standard non-parametric statistical methodologies for rigorously benchmarking and comparing multiple machine learning algorithms across diverse benchmark datasets.
Cover for Multiple Kernel Learning Algorithms

Abstract

In recent years, several methods have been proposed to combine multiple kernels instead of using a single one. These different kernels may correspond to using different notions of similarity or may be using information coming from multiple sources (different representations or different feature subsets). In trying to organize and highlight the similarities and differences between them, we give a taxonomy of and review several multiple kernel learning algorithms. We perform experiments on real data sets for better illustration and comparison of existing algorithms. We see that though there may not be large differences in terms of accuracy, there is difference between them in complexity as given by the number of stored support vectors, the sparsity of the solution as given by the number of used kernels, and training time complexity. We see that overall, using multiple kernels instead of a single one is useful and believe that combining kernels in a nonlinear or data-dependent way seems more promising than linear combination in fusing information provided by simple linear kernels, whereas linear methods are more reasonable when combining complex Gaussian kernels.

Table of Contents

  • 1. Introduction
  • 2. Key Properties of Multiple Kernel Learning
  • 2.1 The Learning Method
  • 2.2 The Functional Form
  • 2.3 The Target Function
  • 2.4 The Training Method
  • 2.5 The Base Learner
  • 2.6 The Computational Complexity
  • 3. Multiple Kernel Learning Algorithms
  • 3.1 Fixed Rules
  • 3.2 Heuristic Approaches
  • 3.3 Similarity Optimizing Linear Approaches with Arbitrary Kernel Weights
  • 3.4 Similarity Optimizing Linear Approaches with Nonnegative Kernel Weights
  • 3.5 Similarity Optimizing Linear Approaches with Kernel Weights on a Simplex
  • 3.6 Structural Risk Optimizing Linear Approaches with Arbitrary Kernel Weights
  • 3.7 Structural Risk Optimizing Linear Approaches with Nonnegative Kernel Weights
  • 3.8 Structural Risk Optimizing Linear Approaches with Kernel Weights on a Simplex
  • 3.9 Structural Risk Optimizing Nonlinear Approaches
  • 3.10 Structural Risk Optimizing Data-Dependent Approaches
  • 3.11 Bayesian Approaches
  • 3.12 Boosting Approaches
  • 4. Experiments
  • 4.1 Compared Algorithms
  • 4.2 Experimental Methodology
  • 4.3 Protein Fold Prediction Experiments
  • 4.4 Pendigits Digit Recognition Experiments
  • 4.5 Multiple Features Digit Recognition Experiments
  • 4.6 Internet Advertisements Experiments
  • 4.7 Overall Comparison
  • 4.8 Overall Comparison Using Gaussian Kernel
  • 5. Conclusions
  • Acknowledgments
  • Appendix A. List of Acronyms
  • Appendix B. List of Notation
  • References

Knowls

  1. Knowl 1 — Six-Dimensional Taxonomy for Categorizing Multiple Kernel Learning Algorithms

    definition

    Multiple kernel learning (MKL) algorithms combine a set of PP candidate kernel functions, {km(xim,xjm)}m=1P\{k_m(x_i^m, x_j^m)\}_{m=1}^P, into a composite kernel function kη(xi,xj)=fη({km(xim,xjm)}m=1P)k_\eta(x_i, x_j) = f_\eta(\{k_m(x_i^m, x_j^m)\}_{m=1}^P) parameterized by η\eta. Existing MKL methods are systematically categorized across six defining dimensions:

    1. Learning Method: The conceptual paradigm used to determine combination parameters η\eta:

      • Fixed rules: Parameter-free algebraic combinations (e.g., unweighted addition or multiplication).
      • Heuristics: Parameterized combinations set by individual kernel quality measures or performance evaluations evaluated independently.
      • Optimization: Parameters learned by optimizing a joint objective function via mathematical programming.
      • Bayesian approaches: Kernel weights treated as random variables with prior distributions (e.g., Dirichlet priors).
      • Boosting approaches: Kernels iteratively added to the ensemble until a stopping criterion is met.
    2. Functional Form: The algebraic structure of fη(⋅)f_\eta(\cdot):

      • Linear combination: kη(xi,xj)=∑m=1Pηmkm(xim,xjm)k_\eta(x_i, x_j) = \sum_{m=1}^P \eta_m k_m(x_i^m, x_j^m), subdivided into linear sum (arbitrary η∈RP\eta \in \mathbb{R}^P), conic sum (nonnegative η∈R+P\eta \in \mathbb{R}_+^P), and convex sum (η∈R+P\eta \in \mathbb{R}_+^P with ∑m=1Pηm=1\sum_{m=1}^P \eta_m = 1).
      • Nonlinear combination: Multiplicative, polynomial, or exponential combinations of kernel functions.
      • Data-dependent combination: Combination weights ηm(x)\eta_m(x) conditioned on the input instance xx, allowing spatially localized kernel mixing.
    3. Target Function: The objective optimized to select η\eta:

      • Similarity-based: Maximizing an alignment or metric (e.g., kernel-target alignment, centered alignment, Euclidean distance, Kullback-Leibler divergence) between the combined kernel and an ideal target kernel yy⊤y y^\top.
      • Structural risk: Minimizing empirical loss plus a regularizer over base learner weights and kernel parameters.
      • Bayesian evidence: Maximizing likelihood, marginal likelihood, or posterior distributions.
    4. Training Method:

      • One-step methods: Computing base learner parameters and kernel weights in a single step, either sequentially (kernel weights first, base learner second) or simultaneously.
      • Two-step methods: Alternating iterative optimization, alternating between updating base learner parameters with fixed kernel weights and updating kernel weights with fixed base learner parameters.
    5. Base Learner: The underlying machine learning algorithm (e.g., Support Vector Machines (SVM), Support Vector Regression (SVR), Kernel Fisher Discriminant Analysis (KFDA), Regularized Kernel Discriminant Analysis (RKDA), Kernel Ridge Regression (KRR), Gaussian Processes (GP)).

    6. Computational Complexity: The mathematical optimization class solved during training (e.g., Semidefinite Programming (SDP), Quadratically Constrained Quadratic Programming (QCQP), Second-Order Cone Programming (SOCP), Semi-Infinite Linear Programming (SILP), Quadratic Programming (QP)).

  2. Knowl 2 — Unified Structural Risk Optimization Framework for Linear Multiple Kernel Learning

    model/method

    Linear multiple kernel learning with structural risk minimization can be formulated within a unified optimization framework parameterized by the ℓp\ell_p-norm (p≥1p \ge 1) on the kernel combination coefficients η∈R+P\eta \in \mathbb{R}_+^P.

    In the Tikhonov regularization formulation, the primal problem is: min⁡{wm}m=1P,b,η12∑m=1P∥wm∥22ηm+C∑i=1NL(∑m=1P⟨wm,Φm(xim)⟩+b,yi)+μ∥η∥pp\min_{\{w_m\}_{m=1}^P, b, \eta} \frac{1}{2} \sum_{m=1}^P \frac{\|w_m\|_2^2}{\eta_m} + C \sum_{i=1}^N L\left(\sum_{m=1}^P \langle w_m, \Phi_m(x_i^m)\rangle + b, y_i\right) + \mu \|\eta\|_p^p subject to wm∈RSm, b∈R, η∈R+P\text{subject to } w_m \in \mathbb{R}^{S_m}, \ b \in \mathbb{R}, \ \eta \in \mathbb{R}_+^P where xim∈RDmx_i^m \in \mathbb{R}^{D_m} is the mm-th feature representation of the ii-th instance, Φm:RDm→RSm\Phi_m: \mathbb{R}^{D_m} \to \mathbb{R}^{S_m} is the corresponding feature map, wmw_m is the weight vector associated with feature space mm, bb is the shared bias, yi∈{−1,+1}y_i \in \{-1, +1\} is the target label, L(⋅,⋅)L(\cdot, \cdot) is a loss function (e.g., hinge loss L(f(x),y)=max⁡(0,1−yf(x))L(f(x), y) = \max(0, 1 - y f(x))), C>0C > 0 is the regularization trade-off parameter, and μ>0\mu > 0 regularizes the kernel weights.

    In the equivalent Ivanov regularization formulation, the norm constraint is placed explicitly on η\eta: min⁡{wm}m=1P,b,η12∑m=1P∥wm∥22ηm+C∑i=1NL(∑m=1P⟨wm,Φm(xim)⟩+b,yi)\min_{\{w_m\}_{m=1}^P, b, \eta} \frac{1}{2} \sum_{m=1}^P \frac{\|w_m\|_2^2}{\eta_m} + C \sum_{i=1}^N L\left(\sum_{m=1}^P \langle w_m, \Phi_m(x_i^m)\rangle + b, y_i\right) subject to wm∈RSm, b∈R, η∈R+P, ∥η∥pp≤1\text{subject to } w_m \in \mathbb{R}^{S_m}, \ b \in \mathbb{R}, \ \eta \in \mathbb{R}_+^P, \ \|\eta\|_p^p \le 1

    When p=1p = 1, the Ivanov formulation corresponds to convex combination linear MKL, which enforces sparsity across kernels (acting as kernel/feature selection). For p>1p > 1 (such as p=2p=2), the formulation enforces non-sparse combinations, allowing complementary information across all candidate kernels to be retained.

  3. Knowl 3 — Closed-Form Kernel Weight Update for Arbitrary $\ell_p$-Norm Multiple Kernel Learning

    equation

    In two-step alternating optimization algorithms for ℓp\ell_p-norm multiple kernel learning (p≥1p \ge 1), the base learner solves a standard canonical SVM for fixed kernel weights η=(η1,…,ηP)⊤\eta = (\eta_1, \dots, \eta_P)^\top, yielding dual Lagrange multipliers α∈R+N\alpha \in \mathbb{R}_+^N. The ℓ2\ell_2-norm of the hyperplane weight vector wmw_m in the mm-th feature space satisfies: ∥wm∥2=ηm∑i=1N∑j=1Nαiαjyiyjkm(xim,xjm)\|w_m\|_2 = \eta_m \sqrt{\sum_{i=1}^N \sum_{j=1}^N \alpha_i \alpha_j y_i y_j k_m(x_i^m, x_j^m)} where yi∈{−1,+1}y_i \in \{-1, +1\} are class labels and km(xim,xjm)k_m(x_i^m, x_j^m) is the mm-th base kernel.

    For a fixed α\alpha, the exact closed-form solution that updates each kernel weight ηm\eta_m under the constraint ∥η∥p≤1\|\eta\|_p \le 1 (p≥1p \ge 1) is given by: ηm=∥wm∥22p+1(∑h=1P∥wh∥22pp+1)1/p\eta_m = \frac{\|w_m\|_2^{\frac{2}{p+1}}}{\left(\sum_{h=1}^P \|w_h\|_2^{\frac{2p}{p+1}}\right)^{1/p}}

    When p=1p = 1, this closed-form update reduces to the convex sum normalization: ηm=∥wm∥2∑h=1P∥wh∥2\eta_m = \frac{\|w_m\|_2}{\sum_{h=1}^P \|w_h\|_2}

  4. Knowl 4 — Centered-Kernel Alignment and Exact Solution for Kernel Weights

    theoretical result

    For two kernel matrices K1,K2∈RN×NK_1, K_2 \in \mathbb{R}^{N \times N}, the empirical centered-kernel alignment CA(K1,K2)CA(K_1, K_2) is defined as: CA(K1,K2)=⟨K1c,K2c⟩F⟨K1c,K1c⟩F⟨K2c,K2c⟩FCA(K_1, K_2) = \frac{\langle K_1^c, K_2^c \rangle_F}{\sqrt{\langle K_1^c, K_1^c \rangle_F \langle K_2^c, K_2^c \rangle_F}} where ⟨A,B⟩F=tr⁡(A⊤B)\langle A, B \rangle_F = \operatorname{tr}(A^\top B) is the Frobenius inner product, and KcK^c represents the centered kernel matrix calculated as: Kc=K−1N11⊤K−1NK11⊤+1N2(1⊤K1)11⊤K^c = K - \frac{1}{N} \mathbf{1}\mathbf{1}^\top K - \frac{1}{N} K \mathbf{1}\mathbf{1}^\top + \frac{1}{N^2} (\mathbf{1}^\top K \mathbf{1}) \mathbf{1}\mathbf{1}^\top with 1∈RN\mathbf{1} \in \mathbb{R}^N being the vector of all ones.

    When combining PP candidate kernel matrices linearly as Kη=∑m=1PηmKmK_\eta = \sum_{m=1}^P \eta_m K_m, maximizing the centered alignment to the ideal target kernel yy⊤y y^\top over the ℓ2\ell_2-sphere M={η∈RP:∥η∥2=1}\mathcal{M} = \{\eta \in \mathbb{R}^P : \|\eta\|_2 = 1\}: max⁡η∈MCA(Kη,yy⊤)\max_{\eta \in \mathcal{M}} CA(K_\eta, y y^\top) has the unique analytical closed-form solution: η=M−1a∥M−1a∥2\eta = \frac{M^{-1}a}{\|M^{-1}a\|_2} where M∈RP×PM \in \mathbb{R}^{P \times P} is the matrix with entries Mmh=⟨Kmc,Khc⟩FM_{mh} = \langle K_m^c, K_h^c \rangle_F (m,h=1,…,Pm, h = 1, \dots, P) and a∈RPa \in \mathbb{R}^P is the vector with entries am=⟨Kmc,yy⊤⟩Fa_m = \langle K_m^c, y y^\top \rangle_F (m=1,…,Pm = 1, \dots, P).

  5. Knowl 5 — Localized Multiple Kernel Learning (LMKL) Formulation

    model/method

    Localized Multiple Kernel Learning (LMKL) combines candidate kernels via instance-specific weights computed by a parametric gating model ηm(x∣V)\eta_m(x | V). The decision function is: f(x)=∑m=1Pηm(x∣V)⟨wm,Φm(xm)⟩+bf(x) = \sum_{m=1}^P \eta_m(x | V) \langle w_m, \Phi_m(x^m) \rangle + b

    The primal optimization problem is formulated as: min⁡{wm}m=1P,ξ,b,V12∑m=1P∥wm∥22+C∑i=1Nξi\min_{\{w_m\}_{m=1}^P, \xi, b, V} \frac{1}{2} \sum_{m=1}^P \|w_m\|_2^2 + C \sum_{i=1}^N \xi_i subject to yi(∑m=1Pηm(xi∣V)⟨wm,Φm(xim)⟩+b)≥1−ξi,ξi≥0∀i=1,…,N\text{subject to } y_i \left( \sum_{m=1}^P \eta_m(x_i | V) \langle w_m, \Phi_m(x_i^m) \rangle + b \right) \ge 1 - \xi_i, \quad \xi_i \ge 0 \quad \forall i = 1, \dots, N where VV parameterizes the gating model defined on a gating representation xG∈RDGx^G \in \mathbb{R}^{D_G}.

    Two principal gating models are used:

    1. Softmax Gating (competitive kernel selection): ηm(x∣V)=exp⁡(⟨vm,xG⟩+vm0)∑h=1Pexp⁡(⟨vh,xG⟩+vh0)∀m=1,…,P\eta_m(x | V) = \frac{\exp(\langle v_m, x^G \rangle + v_{m0})}{\sum_{h=1}^P \exp(\langle v_h, x^G \rangle + v_{h0})} \quad \forall m=1, \dots, P
    2. Sigmoid Gating (cooperative kernel combination): ηm(x∣V)=11+exp⁡(−(⟨vm,xG⟩+vm0))∀m=1,…,P\eta_m(x | V) = \frac{1}{1 + \exp(-(\langle v_m, x^G \rangle + v_{m0}))} \quad \forall m=1, \dots, P where V={vm,vm0}m=1PV = \{v_m, v_{m0}\}_{m=1}^P with vm∈RDGv_m \in \mathbb{R}^{D_G} and vm0∈Rv_{m0} \in \mathbb{R}.

    For a fixed gating parameter set VV, the problem reduces to a convex canonical SVM dual with the localized combined kernel: kη(xi,xj)=∑m=1Pηm(xi∣V)km(xim,xjm)ηm(xj∣V)k_\eta(x_i, x_j) = \sum_{m=1}^P \eta_m(x_i | V) k_m(x_i^m, x_j^m) \eta_m(x_j | V) Training alternates between solving the SVM dual for α\alpha given VV, and taking a gradient-descent step on VV with respect to the dual objective J(V)J(V) given α\alpha.

  6. Knowl 6 — Nonlinear Multiple Kernel Learning with Degree-$d$ Polynomial Combinations

    model/method

    Nonlinear multiple kernel learning (NLMKL) learns a degree-dd homogeneous polynomial combination of PP base kernels: kη(xi,xj)=∑q∈Rη1q1η2q2⋯ηPqPk1(xi1,xj1)q1k2(xi2,xj2)q2⋯kP(xiP,xjP)qPk_\eta(x_i, x_j) = \sum_{q \in \mathcal{R}} \eta_1^{q_1} \eta_2^{q_2} \cdots \eta_P^{q_P} k_1(x_i^1, x_j^1)^{q_1} k_2(x_i^2, x_j^2)^{q_2} \cdots k_P(x_i^P, x_j^P)^{q_P} where R={q∈Z+P:∑m=1Pqm=d}\mathcal{R} = \left\{q \in \mathbb{Z}_+^P : \sum_{m=1}^P q_m = d\right\} and η=(η1,…,ηP)⊤∈R+P\eta = (\eta_1, \dots, \eta_P)^\top \in \mathbb{R}_+^P.

    For the quadratic case (d=2d=2), this evaluates to: kη(xi,xj)=∑m=1P∑h=1Pηmηhkm(xim,xjm)kh(xih,xjh)k_\eta(x_i, x_j) = \sum_{m=1}^P \sum_{h=1}^P \eta_m \eta_h k_m(x_i^m, x_j^m) k_h(x_i^h, x_j^h)

    The combination parameters η\eta are optimized over a bounded convex constraint set M\mathcal{M} via min-max optimization: min⁡η∈Mmax⁡α∈RN−α⊤(Kη+λI)α+2y⊤α\min_{\eta \in \mathcal{M}} \max_{\alpha \in \mathbb{R}^N} -\alpha^\top (K_\eta + \lambda I) \alpha + 2 y^\top \alpha where λ>0\lambda > 0 is a ridge regularization parameter, and M\mathcal{M} is defined either as an ℓ1\ell_1-norm bounded set M1={η∈R+P:∥η−η0∥1≤Λ}\mathcal{M}_1 = \{\eta \in \mathbb{R}_+^P : \|\eta - \eta_0\|_1 \le \Lambda\} or an ℓ2\ell_2-norm bounded set M2={η∈R+P:∥η−η0∥2≤Λ}\mathcal{M}_2 = \{\eta \in \mathbb{R}_+^P : \|\eta - \eta_0\|_2 \le \Lambda\} (with prior point η0\eta_0 and radius Λ\Lambda).

    Optimization proceeds in two alternating steps: solving for base learner dual variables α\alpha with fixed η\eta, and performing a projected gradient descent update on η\eta onto M\mathcal{M} with fixed α\alpha.

  7. Knowl 7 — Generalized Multiple Kernel Learning (GMKL)

    model/method

    Generalized Multiple Kernel Learning (GMKL) extends MKL to arbitrary differentiable kernel combination functions kη(xi,xj)k_\eta(x_i, x_j) and arbitrary differentiable regularization functions r(η)r(\eta) on the kernel weights η∈R+P\eta \in \mathbb{R}_+^P.

    The primal problem for binary classification with hinge loss is: min⁡wη,ξ,b,η12∥wη∥22+C∑i=1Nξi+r(η)\min_{w_\eta, \xi, b, \eta} \frac{1}{2} \|w_\eta\|_2^2 + C \sum_{i=1}^N \xi_i + r(\eta) subject to yi(⟨wη,Φη(xi)⟩+b)≥1−ξi,ξi≥0∀i=1,…,N,η≥0\text{subject to } y_i (\langle w_\eta, \Phi_\eta(x_i) \rangle + b) \ge 1 - \xi_i, \quad \xi_i \ge 0 \quad \forall i=1,\dots,N, \quad \eta \ge 0 where Φη(⋅)\Phi_\eta(\cdot) is the composite feature mapping corresponding to kη(⋅,⋅)k_\eta(\cdot, \cdot).

    For a fixed η\eta, solving the canonical SVM yields dual variables α∈R+N\alpha \in \mathbb{R}_+^N and optimal dual objective value J(η)J(\eta): J(η)=∑i=1Nαi−12∑i=1N∑j=1Nαiαjyiyjkη(xi,xj)+r(η)J(\eta) = \sum_{i=1}^N \alpha_i - \frac{1}{2} \sum_{i=1}^N \sum_{j=1}^N \alpha_i \alpha_j y_i y_j k_\eta(x_i, x_j) + r(\eta)

    The gradient of J(η)J(\eta) with respect to ηm\eta_m is calculated analytically as: ∂J(η)∂ηm=∂r(η)∂ηm−12∑i=1N∑j=1Nαiαjyiyj∂kη(xi,xj)∂ηm∀m=1,…,P\frac{\partial J(\eta)}{\partial \eta_m} = \frac{\partial r(\eta)}{\partial \eta_m} - \frac{1}{2} \sum_{i=1}^N \sum_{j=1}^N \alpha_i \alpha_j y_i y_j \frac{\partial k_\eta(x_i, x_j)}{\partial \eta_m} \quad \forall m=1, \dots, P This allows non-convex and product combinations, such as product Gaussian feature scaling kηP(xi,xj)=exp⁡(−∑m=1Dηm(xi[m]−xj[m])2)k_\eta^P(x_i, x_j) = \exp\left(-\sum_{m=1}^D \eta_m (x_i[m] - x_j[m])^2\right), to be trained using projected gradient descent.

  8. Knowl 8 — Empirical Superiority of Nonlinear and Data-Dependent MKL on Simple Linear Kernels

    empirical result

    On benchmark classification datasets with heterogeneous feature representations (Protein Fold, Pendigits, Multiple Features, Internet Advertisements) evaluated with simple linear base kernels kLIN(xi,xj)=⟨xi,xj⟩k_{\text{LIN}}(x_i, x_j) = \langle x_i, x_j \rangle:

    1. Combination vs Single Best: Every MKL combination method consistently outperforms the single best feature SVM, confirming that integrating heterogeneous feature representations provides significant performance gains.
    2. Linear MKL vs Unweighted Mean: Standard linear combination MKL methods (such as SimpleMKL, original SOCP MKL, and GLMKL) achieve average test accuracies that are comparable to, or only marginally different from, the parameter-free unweighted kernel average (k(xi,xj)=1P∑m=1Pkm(xim,xjm)k(x_i, x_j) = \frac{1}{P} \sum_{m=1}^P k_m(x_i^m, x_j^m)).
    3. Nonlinear and Data-Dependent Advantage: Nonlinear MKL (NLMKL with ℓ1\ell_1 and ℓ2\ell_2 bounds) and localized MKL (LMKL with softmax and sigmoid gating) achieve statistically significant accuracy improvements over single-kernel SVM, feature concatenation (SVM-all), and the unweighted kernel mean (RBMKL mean). On the Pendigits digit recognition tasks, NLMKL achieves >99%>99\% test accuracy, improving over the unweighted sum baseline by 6%6\% to 8%8\%.
  9. Knowl 9 — Empirical Inefficacy of Complex Kernel Combinations on Gaussian Kernels

    empirical result

    When evaluated on multiple feature representations using nonlinear Gaussian kernels across multiple bandwidth scales {s∈{Dm/2,Dm,2Dm}}\left\{s \in \left\{\sqrt{D_m}/2, \sqrt{D_m}, 2\sqrt{D_m}\right\}\right\}:

    1. No Significant Advantage Over Fixed Mean: No trained MKL algorithm (including alignment-based, structural risk linear, nonlinear, or localized MKL) achieves a statistically significant accuracy improvement over the simple unweighted mean of Gaussian kernels (kη(xi,xj)=1P∑m=1Pkm(xim,xjm)k_\eta(x_i, x_j) = \frac{1}{P} \sum_{m=1}^P k_m(x_i^m, x_j^m)).
    2. Degradation of Nonlinear/Data-Dependent Combinations: Unlike in the linear kernel setting, nonlinear MKL (NLMKL) and localized MKL (LMKL) fail to improve upon linear combinations and are often slightly outperformed by the unweighted mean when candidate kernels are already highly nonlinear Gaussian kernels.
    3. Kernel Pruning: While trained MKL algorithms do not boost classification accuracy over the unweighted mean with Gaussian kernels, sparsity-inducing algorithms (such as ABMKL conic/convex, CABMKL conic, SimpleMKL, and ℓ1\ell_1-norm GLMKL) successfully eliminate redundant kernel scales without degrading classification performance.
  10. Knowl 10 — Support Vector Sparsity and Optimization Complexity Trade-offs across MKL Families

    empirical result

    Benchmarking across diverse MKL formulations reveals stark trade-offs in decision function complexity (percentage of support vectors stored) and training optimization cost (number of solver iterations):

    1. Support Vector Retention:

      • Multiplicative/Nonlinear Combinations (e.g., product rule RBMKL, quadratic NLMKL) store the highest percentage of support vectors (frequently exceeding 85–95%85\text{--}95\% on benchmark tasks). By shrinking off-diagonal kernel similarities toward zero, multiplicative combinations cause the classifier to approximate a 1-nearest-neighbor rule.
      • Localized MKL (LMKL with softmax/sigmoid gating) consistently stores the smallest percentage of support vectors (often 10–30%10\text{--}30\% fewer than linear SVM baselines) by tailoring active kernels to local feature space partitions.
      • Linear MKL methods retain support vector proportions comparable to standard SVMs.
    2. Training Optimization Complexity:

      • Group Lasso-based alternating MKL (GLMKL for both p=1p=1 and p=2p=2) converges in very few outer iterations (typically 4 to 15 SVM solver calls), achieving the fastest training among two-step MKL methods with statistical equivalence to one-step methods in number of solver calls.
      • Gradient-based and localized methods (SimpleMKL, GMKL, LMKL) require substantially more outer iterations (often 20 to >100>100 solver calls), with LMKL exhibiting high iteration variance due to random initialization of the gating network parameters.

Coverage note — Detailed reviews of individual domain-specific applications from external literature (e.g., specific bioinformatics gene prioritizations or KFDA/RKDA algebraic reformulations from prior papers) were omitted to focus on the overarching taxonomy, core mathematical frameworks (GLMKL, LMKL, NLMKL, GMKL, Centered Alignment), and the experimental comparative results.

References

  1. 1.Ethem Alpaydın. Combined 5×25 \times 2 cv FF test for comparing supervised classification learning algorithms. Neural Computation, 11(8):1885–1892, 1999.
  2. 2.Andreas Argyriou, Charles A. Micchelli, and Massimiliano Pontil. Learning convex combinations of continuously parameterized basic kernels. In Proceeding of the 18th Conference on Learning Theory, 2005.
  3. 3.Andreas Argyriou, Raphael Hauser, Charles A. Micchelli, and Massimiliano Pontil. A DC-programming algorithm for kernel selection. In Proceedings of the 23rd International Conference on Machine Learning, 2006.
  4. 4.Francis R. Bach. Consistency of the group Lasso and multiple kernel learning. Journal of Machine Learning Research, 9:1179–1225, 2008.
  5. 5.Francis R. Bach. Exploring large feature spaces with hierarchical multiple kernel learning. In Advances in Neural Information Processing Systems 21, 2009.
  6. 6.Francis R. Bach, Gert R. G. Lanckriet, and Michael I. Jordan. Multiple kernel learning, conic duality, and the SMO algorithm. In Proceedings of the 21st International Conference on Machine Learning, 2004.
  7. 7.Asa Ben-Hur and William Stafford Noble. Kernel methods for predicting protein-protein interactions. Bioinformatics, 21(Suppl 1):i38–46, 2005.
  8. 8.Kristin P. Bennett, Michinari Momma, and Mark J. Embrechts. MARK: A boosting algorithm for heterogeneous kernel models. In Proceedings of the 8th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2002.
  9. 9.Jinbo Bi, Tong Zhang, and Kristin P. Bennett. Column-generation boosting methods for mixture of kernels. In Proceedings of the 10th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2004.
  10. 10.Olivier Bousquet and Daniel J. L. Herrmann. On the complexity of learning the kernel matrix. In Advances in Neural Information Processing Systems 15, 2003.
  11. 11.Olivier Chapelle and Alain Rakotomamonjy. Second order optimization of kernel parameters. In NIPS Workshop on Automatic Selection of Optimal Kernels, 2008.
  12. 12.Olivier Chapelle, Vladimir Vapnik, Olivier Bousquet, and Sayan Mukherjee. Choosing multiple parameters for support vector machines. Machine Learning, 46(1–3):131–159, 2002.
  13. 13.Mario Christoudias, Raquel Urtasun, and Trevor Darrell. Bayesian localized multiple kernel learning. Technical Report UCB/EECS-2009-96, University of California at Berkeley, 2009.
  14. 14.Domenico Conforti and Rosita Guido. Kernel based support vector machine via semidefinite programming: Application to medical diagnosis. Computers and Operations Research, 37(8):1389–1394, 2010.
  15. 15.Corinna Cortes, Mehryar Mohri, and Afshin Rostamizadeh. L2L_2 regularization for learning kernels. In Proceedings of the 25th Conference on Uncertainty in Artificial Intelligence, 2009.
  16. 16.Corinna Cortes, Mehryar Mohri, and Rostamizadeh Afshin. Two-stage learning kernel algorithms. In Proceedings of the 27th International Conference on Machine Learning, 2010a.
  17. 17.Corinna Cortes, Mehryar Mohri, and Afshin Rostamizadeh. Learning non-linear combinations of kernels. In Advances in Neural Information Processing Systems 22, 2010b.
  18. 18.Koby Crammer, Joseph Keshet, and Yoram Singer. Kernel design using boosting. In Advances in Neural Information Processing Systems 15, 2003.
  19. 19.Nello Cristianini and John Shawe-Taylor. An Introduction to Support Vector Machines and Other Kernel-Based Learning Methods. Cambridge University Press, 2000.
  20. 20.Nello Cristianini, John Shawe-Taylor, Andree Elisseef, and Jaz Kandola. On kernel-target alignment. In Advances in Neural Information Processing Systems 14, 2002.
  21. 21.Theodoros Damoulas and Mark A. Girolami. Probabilistic multi-class multi-kernel learning: On protein fold recognition and remote homology detection. Bioinformatics, 24(10):1264–1270, 2008.
  22. 22.Theodoros Damoulas and Mark A. Girolami. Combining feature spaces for classification. Pattern Recognition, 42(11):2671–2683, 2009a.
  23. 23.Theodoros Damoulas and Mark A. Girolami. Pattern recognition with a Bayesian kernel combination machine. Pattern Recognition Letters, 30(1):46–54, 2009b.
  24. 24.Tijl De Bie, Leon-Charles Tranchevent, Liesbeth M. M. van Oeffelen, and Yves Moreau. Kernel-based data fusion for gene prioritization. Bioinformatics, 23(13):i125–132, 2007.
  25. 25.Isaac Martín de Diego, Javier M. Moguerza, and Alberto Muñoz. Combining kernel information for support vector classification. In Proceedings of the 4th International Workshop Multiple Classifier Systems, 2004.
  26. 26.Isaac Martín de Diego, Alberto Muñoz, and Javier M. Moguerza. Methods for the combination of kernel matrices within a support vector framework. Machine Learning, 78(1–2):137–174, 2010a.
  27. 27.Isaac Martín de Diego, Ángel Serrano, Cristina Conde, and Enrique Cabello. Face verification with a kernel fusion method. Pattern Recognition Letters, 31:837–844, 2010b.
  28. 28.Réda Dehak, Najim Dehak, Patrick Kenny, and Pierre Dumouchel. Kernel combination for SVM speaker verification. In Proceedings of the Speaker and Language Recognition Workshop, 2008.
  29. 29.Janez Demšar. Statistical comparisons of classifiers over multiple data sets. Journal of Machine Learning Research, 7:1–30, 2006.
  30. 30.Glenn Fung, Murat Dundar, Jinbo Bi, and Bharat Rao. A fast iterative algorithm for Fisher discriminant using heterogeneous kernels. In Proceedings of the 21st International Conference on Machine Learning, 2004.
  31. 31.Peter Vincent Gehler and Sebastian Nowozin. Infinite kernel learning. Technical report, Max Planck Institute for Biological Cybernetics, 2008.
  32. 32.Mark Girolami and Simon Rogers. Hierarchic Bayesian models for kernel learning. In Proceedings of the 22nd International Conference on Machine Learning, 2005.
  33. 33.Mark Girolami and Mingjun Zhong. Data integration for classification problems employing Gaussian process priors. In Advances in Neural Processing Systems 19, 2007.
  34. 34.Mehmet Gönen and Ethem Alpaydın. Localized multiple kernel learning. In Proceedings of the 25th International Conference on Machine Learning, 2008.
  35. 35.Yves Grandvalet and Stéphane Canu. Adaptive scaling for feature selection in SVMs. In Advances in Neural Information Processing Systems 15, 2003.
  36. 36.Junfeng He, Shih-Fu Chang, and Lexing Xie. Fast kernel learning for spatial pyramid matching. In Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 2008.
  37. 37.Mingqing Hu, Yiqiang Chen, and James Tin-Yau Kwok. Building sparse multiple-kernel SVM classifiers. IEEE Transactions on Neural Networks, 20(5):827–839, 2009.
  38. 38.Christian Igel, Tobias Glasmachers, Britta Mersch, Nico Pfeifer, and Peter Meinicke. Gradient-based optimization of kernel-target alignment for sequence kernels applied to bacterial gene start detection. IEEE/ACM Transactions on Computational Biology and Bioinformatics, 4(2):216–226, 2007.
  39. 39.Thorsten Joachims, Nello Cristianini, and John Shawe-Taylor. Composite kernels for hypertext categorisation. In Proceedings of the 18th International Conference on Machine Learning, 2001.
  40. 40.Jaz Kandola, John Shawe-Taylor, and Nello Cristianini. Optimizing kernel alignment over combinations of kernels. In Proceedings of the 19th International Conference on Machine Learning, 2002.
  41. 41.Seung-Jean Kim, Alessandro Magnani, and Stephen Boyd. Optimal kernel selection in kernel Fisher discriminant analysis. In Proceedings of the 23rd International Conference on Machine Learning, 2006.
  42. 42.Marius Kloft, Ulf Brefeld, Sören Sonnenburg, Pavel Laskov, Klaus-Robert Müller, and Alexander Zien. Efficient and accurate ℓp\ell_p-norm multiple kernel learning. In Advances in Neural Information Processing Systems 22, 2010a.
  43. 43.Marius Kloft, Ulf Brefeld, Sören Sonnenburg, and Alexander Zien. Non-sparse regularization and efficient training with multiple kernels. Technical report, Electrical Engineering and Computer Sciences, University of California at Berkeley, 2010b.
  44. 44.Gert R. G. Lanckriet, Nello Cristianini, Peter Bartlett, Laurent El Ghaoui, and Michael I. Jordan. Learning the kernel matrix with semidefinite programming. In Proceedings of the 19th International Conference on Machine Learning, 2002.
  45. 45.Gert R. G. Lanckriet, Nello Cristianini, Peter Bartlett, Laurent El Ghaoui, and Michael I. Jordan. Learning the kernel matrix with semidefinite programming. Journal of Machine Learning Research, 5:27–72, 2004a.
  46. 46.Gert R. G. Lanckriet, Tijl de Bie, Nello Cristianini, Michael I. Jordan, and William Stafford Noble. A statistical framework for genomic data fusion. Bioinformatics, 20(16):2626–2635, 2004b.
  47. 47.Gert R. G. Lanckriet, Minghua Deng, Nello Cristianini, Michael I. Jordan, and William Stafford Noble. Kernel-based data fusion and its application to protein function prediction in Yeast. In Proceedings of the Pacific Symposium on Biocomputing, 2004c.
  48. 48.Wan-Jui Lee, Sergey Verzakov, and Robert P. W. Duin. Kernel combination versus classifier combination. In Proceedings of the 7th International Workshop on Multiple Classifier Systems, 2007.
  49. 49.Darrin P. Lewis, Tony Jebara, and William Stafford Noble. Support vector machine learning from heterogeneous data: An empirical analysis using protein sequence and structure. Bioinformatics, 22(22):2753–2760, 2006a.
  50. 50.Darrin P. Lewis, Tony Jebara, and William Stafford Noble. Nonstationary kernel combination. In Proceedings of the 23rd International Conference on Machine Learning, 2006b.
  51. 51.Yen-Yu Lin, Tyng-Luh Liu, and Chiou-Shann Fuh. Dimensionality reduction for data in multiple feature representations. In Advances in Neural Processing Systems 21, 2009.
  52. 52.Huma Lodhi, Craig Saunders, John Shawe-Taylor, Nello Cristianini, and Chris Watkins. Text classification using string kernels. Journal of Machine Learning Research, 2:419–444, 2002.
  53. 53.Chris Longworth and Mark J. F. Gales. Multiple kernel learning for speaker verification. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing, 2008.
  54. 54.Chris Longworth and Mark J. F. Gales. Combining derivative and parametric kernels for speaker verification. IEEE Transactions on Audio, Speech, and Language Processing, 17(4):748–757, 2009.
  55. 55.Brian McFee and Gert Lanckriet. Partial order embedding with multiple kernels. In Proceedings of the 26th International Conference on Machine Learning, 2009.
  56. 56.Charles A. Micchelli and Massimiliano Pontil. Learning the kernel function via regularization. Journal of Machine Learning Research, 6:1099–1125, 2005.
  57. 57.Javier M. Moguerza, Alberto Muñoz, and Isaac Martín de Diego. Improving support vector classification via the combination of multiple sources of information. In Proceedings of the Structural, Syntactic, and Statistical Pattern Recognition, Joint IAPR International Workshops, 2004.
  58. 58.Mosek. The MOSEK Optimization Tools Manual Version 6.0 (Revision 106). MOSEK ApS, Denmark, 2011.
  59. 59.Canh Hao Nguyen and Tu Bao Ho. An efficient kernel matrix evaluation measure. Pattern Recognition, 41(11):3366–3372, 2008.
  60. 60.William Stafford Noble. Support vector machine applications in computational biology. In Bernhard Schölkopf, Koji Tsuda, and Jean-Philippe Vert, editors, Kernel Methods in Computational Biology, chapter 3. The MIT Press, 2004.
  61. 61.Cheng Soon Ong and Alexander J. Smola. Machine learning using hyperkernels. In Proceedings of the 20th International Conference on Machine Learning, 2003.
  62. 62.Cheng Soon Ong, Alexander J. Smola, and Robert C. Williamson. Hyperkernels. In Advances in Neural Information Processing Systems 15, 2003.
  63. 63.Cheng Soon Ong, Alexander J. Smola, and Robert C. Williamson. Learning the kernel with hyperkernels. Journal of Machine Learning Research, 6:1043–1071, 2005.
  64. 64.Ayşegül Özen, Mehmet Gönen, Ethem Alpaydın, and Türkan Haliloğlu. Machine learning integration for predicting the effect of single amino acid substitutions on protein stability. BMC Structural Biology, 9(1):66, 2009.
  65. 65.Süreyya Özöğür-Akyüz and Gerhard Wilhelm Weber. Learning with infinitely many kernels via semi-infinite programming. In Proceedings of Euro Mini Conference on Continuous Optimization and Knowledge-Based Technologies, 2008.
  66. 66.Paul Pavlidis, Jason Weston, Jinsong Cai, and William Noble Grundy. Gene functional classification from heterogeneous data. In Proceedings of the 5th Annual International Conference on Computational Molecular Biology, 2001.
  67. 67.Shibin Qiu and Terran Lane. Multiple kernel learning for support vector regression. Technical report, Computer Science Department, University of New Mexico, 2005.
  68. 68.Shibin Qiu and Terran Lane. A framework for multiple kernel support vector regression and its applications to siRNA efficacy prediction. IEEE/ACM Transactions on Computational Biology and Bioinformatics, 6(2):190–199, 2009.
  69. 69.Alain Rakotomamonjy, Francis Bach, Stéphane Canu, and Yves Grandvalet. More efficiency in multiple kernel learning. In Proceedings of the 24th International Conference on Machine Learning, 2007.
  70. 70.Alain Rakotomamonjy, Francis R. Bach, Stéphane Canu, and Yves Grandvalet. SimpleMKL. Journal of Machine Learning Research, 9:2491–2521, 2008.
  71. 71.Jagarlapudi Saketha Nath, Govindaraj Dinesh, Sankaran Raman, Chiranjib Bhattacharya, Aharon Ben-Tal, and Kalpathi R. Ramakrishnan. On the algorithmics and applications of a mixed-norm based kernel learning formulation. In Advances in Neural Information Processing Systems 22, 2010.
  72. 72.Bernhard Schölkopf, Koji Tsuda, and Jean-Philippe Vert, editors. Kernel Methods in Computational Biology. The MIT Press, 2004.
  73. 73.Sören Sonnenburg, Gunnar Rätsch, and Christin Schäfer. A general and efficient multiple kernel learning algorithm. In Advances in Neural Information Processing Systems 18, 2006a.
  74. 74.Sören Sonnenburg, Gunnar Rätsch, Christin Schäfer, and Bernhard Schölkopf. Large scale multiple kernel learning. Journal of Machine Learning Research, 7:1531–1565, 2006b.
  75. 75.Niranjan Subrahmanya and Yung C. Shin. Sparse multiple kernel learning for signal processing applications. IEEE Transactions on Pattern Analysis and Machine Intelligence, 32(5):788–798, 2010.
  76. 76.Marie Szafranski, Yves Grandvalet, and Alain Rakotomamonjy. Composite kernel learning. In Proceedings of the 25th International Conference on Machine Learning, 2008.
  77. 77.Marie Szafranski, Yves Grandvalet, and Alain Rakotomamonjy. Composite kernel learning. Machine Learning, 79(1–2):73–103, 2010.
  78. 78.Ying Tan and Jun Wang. A support vector machine with a hybrid kernel and minimal Vapnik-Chervonenkis dimension. IEEE Transactions on Knowledge and Data Engineering, 16(4):385–395, 2004.
  79. 79.Hiroaki Tanabe, Tu Bao Ho, Canh Hao Nguyen, and Saori Kawasaki. Simple but effective methods for combining kernels in computational biology. In Proceedings of IEEE International Conference on Research, Innovation and Vision for the Future, 2008.
  80. 80.Ivor Wai-Hung Tsang and James Tin-Yau Kwok. Efficient hyperkernel learning using second-order cone programming. IEEE Transactions on Neural Networks, 17(1):48–58, 2006.
  81. 81.Koji Tsuda, Shinsuke Uda, Taishin Kin, and Kiyoshi Asai. Minimizing the cross validation error to mix kernel matrices of heterogeneous biological data. Neural Processing Letters, 19(1):63–72, 2004.
  82. 82.Vladimir Vapnik. The Nature of Statistical Learning Theory. John Wiley & Sons, 1998.
  83. 83.Manik Varma and Bodla Rakesh Babu. More generality in efficient multiple kernel learning. In Proceedings of the 26th International Conference on Machine Learning, 2009.
  84. 84.Manik Varma and Debajyoti Ray. Learning the discriminative power-invariance trade-off. In Proceedings of the International Conference in Computer Vision, 2007.
  85. 85.Jason Weston, Sayan Mukherjee, Olivier Chapelle, Massimiliano Pontil, Tomaso Poggio, and Vladimir Vapnik. Feature selection for SVMs. In Advances in Neural Information Processing Systems 13, 2001.
  86. 86.Frank Wilcoxon. Individual comparisons by ranking methods. Biometrics Bulletin, 1(6):80–83, 1945.
  87. 87.Mingrui Wu, Bernhard Schölkopf, and Gökhan Bakır. A direct method for building sparse kernel learning algorithms. Journal of Machine Learning Research, 7:603–624, 2006.
  88. 88.Linli Xu, James Neufeld, Bryce Larson, and Dale Schuurmans. Maximum margin clustering. In Advances in Neural Processing Systems 17, 2005.
  89. 89.Zenglin Xu, Rong Jin, Irwin King, and Michael R. Lyu. An extended level method for efficient multiple kernel learning. In Advances in Neural Information Processing Systems 21, 2009a.
  90. 90.Zenglin Xu, Rong Jin, Jieping Ye, Michael R. Lyu, and Irwin King. Non-monotonic feature selection. In Proceedings of the 26th International Conference on Machine Learning, 2009b.
  91. 91.Zenglin Xu, Rong Jin, Haiqin Yang, Irwin King, and Michael R. Lyu. Simple and efficient multiple kernel learning by group Lasso. In Proceedings of the 27th International Conference on Machine Learning, 2010a.
  92. 92.Zenglin Xu, Rong Jin, Shenghuo Zhu, Michael R. Lyu, and Irwin King. Smooth optimization for effective multiple kernel learning. In Proceedings of the 24th AAAI Conference on Artifical Intelligence, 2010b.
  93. 93.Yoshihiro Yamanishi, Francis Bach, and Jean-Philippe Vert. Glycan classification with tree kernels. Bioinformatics, 23(10):1211–1216, 2007.
  94. 94.Fei Yan, Krystian Mikolajczyk, Josef Kittler, and Muhammad Tahir. A comparison of ℓ1\ell_1 norm and ℓ2\ell_2 norm multiple kernel SVMs in image and video classification. In Proceedings of the 7th International Workshop on Content-Based Multimedia Indexing, 2009.
  95. 95.Jingjing Yang, Yuanning Li, Yonghong Tian, Ling-Yu Duan, and Wen Gao. Group-sensitive multiple kernel learning for object categorization. In Proceedings of the 12th IEEE International Conference on Computer Vision, 2009a.
  96. 96.Jingjing Yang, Yuanning Li, Yonghong Tian, Ling-Yu Duan, and Wen Gao. A new multiple kernel approach for visual concept learning. In Proceedings of the 15th International Multimedia Modeling Conference, 2009b.
  97. 97.Jingjing Yang, Yuanning Li, Yonghong Tian, Ling-Yu Duan, and Wen Gao. Per-sample multiple kernel approach for visual concept learning. EURASIP Journal on Image and Video Processing, 2010.
  98. 98.Jieping Ye, Jianhui Chen, and Shuiwang Ji. Discriminant kernel and regularization parameter learning via semidefinite programming. In Proceedings of the 24th International Conference on Machine Learning, 2007a.
  99. 99.Jieping Ye, Shuiwang Ji, and Jianhui Chen. Learning the kernel matrix in discriminant analysis via quadratically constrained quadratic programming. In Proceedings of the 13th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2007b.
  100. 100.Jieping Ye, Shuiwang Ji, and Jianhui Chen. Multi-class discriminant kernel learning via convex programming. Journal of Machine Learning Research, 9:719–758, 2008.
  101. 101.Yiming Ying, Kaizhu Huang, and Colin Campbell. Enhanced protein fold recognition through a novel data integration approach. BMC Bioinformatics, 10(1):267, 2009.
  102. 102.Bin Zhao, James T. Kwok, and Changshui Zhang. Multiple kernel clustering. In Proceedings of the 9th SIAM International Conference on Data Mining, 2009.
  103. 103.Alexander Zien and Cheng Soon Ong. Multiclass multiple kernel learning. In Proceedings of the 24th International Conference on Machine Learning, 2007.
  104. 104.Alexander Zien and Cheng Soon Ong. An automated combination of kernels for predicting protein subcellular localization. In Proceedings of the 8th International Workshop on Algorithms in Bioinformatics, 2008.

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/