What is Multiple Kernel Learning?
Multiple Kernel Learning (MKL) is a family of methods that jointly learns a predictor and the weights or structure used to combine several kernel functions.
Quick Facts
| Specification | Official Specification |
|---|
How It Works
Jointly optimize kernel weights and a predictor
MKL alternates or jointly solves for predictor parameters and kernel-combination variables under a regularized objective. Sparse formulations can drive many kernel weights to zero, while non-sparse norms spread mass across correlated or complementary kernels.
The JMLR MKL survey organizes methods by learning formulation, optimization, and application. MKL is broader than averaging kernels: constraints, regularizers, localized weights, and the downstream learner determine what is actually learned.
Normalize base kernels before interpreting weights
Kernel matrices can have very different traces, diagonals, or variances. Without training-only normalization, a large numeric scale can dominate the mixture or distort regularization. Centering and normalization must be fitted without using validation or test labels, then reproduced for new samples.
A larger learned weight does not prove causal feature importance. Redundant kernels can exchange weight, correlated modalities can mask one another, and regularization can favor sparse or diffuse solutions. Interpret weights only alongside ablations, uncertainty across folds, and predictive performance.
Nest model selection and budget Gram-matrix cost
Selecting base kernels, bandwidths, normalization, sparsity, and regularization on the same evaluation fold leaks model-selection information. Use an inner loop for all kernel and predictor choices and an outer or untouched split for final assessment; compare against each single kernel and an unweighted average.
SimpleMKL illustrates a reduced-gradient approach for a convex kernel mixture. With M dense kernels over n samples, materializing bases can require O(Mn^2) storage, so use caching, low-rank factors, block products, or explicit features when the measured budget demands them.
Key Characteristics
- Learns a combination of multiple kernel-defined similarity geometries
- Preserves positive semidefiniteness under nonnegative kernel sums
- Can impose sparse or non-sparse regularization on kernel weights
- Requires scale-compatible kernel matrices for meaningful optimization
- Couples kernel selection with the downstream predictor objective
- Can multiply Gram-matrix storage and computation by the number of bases
Common Use Cases
- Combining text, image, graph, or sensor modalities in one predictor
- Learning across predefined feature groups with separate kernels
- Selecting among several RBF bandwidths without one fixed geometry
- Integrating biological sequence, pathway, and expression similarities
- Testing whether learned kernel fusion improves over single-kernel baselines
Example
Loading code...Frequently Asked Questions
Why use Multiple Kernel Learning instead of one kernel?
MKL can combine complementary geometries from feature groups, modalities, or bandwidths and select their weights under one predictive objective. It is useful only when nested validation shows an advantage over each single kernel and a simple fixed average.
Does a weighted sum of kernels remain a valid kernel?
A nonnegative weighted sum of positive-semidefinite kernels remains positive semidefinite. Negative or data-dependent combinations require separate mathematical justification; an arbitrary mixture of similarity scores is not automatically a valid RKHS kernel.
Why must base kernels be normalized in MKL?
Different traces or diagonal scales can let numeric magnitude dominate learned weights and regularization. Choose a documented normalization using training data only, apply the matching cross-kernel transformation at inference, and include that choice in validation.
Do MKL weights measure feature importance?
Not by themselves. Weight depends on kernel scale, redundancy, constraints, regularization, and correlated information. Use held-out ablations, stability across resamples, and domain checks before treating a weight as descriptive evidence, and never call it causal importance.
How does Multiple Kernel Learning scale?
Storing M dense n-by-n Gram matrices can require O(Mn²) memory before optimization cost. Kernel caching, on-demand blocks, low-rank approximations, explicit features, and a smaller validated base set can reduce cost while introducing trade-offs to measure.