What is RBF Kernel?
RBF Kernel (Radial Basis Function Kernel) is a positive-definite kernel that assigns similarity `exp(-gamma ||x-y||^2)` according to the squared Euclidean distance between two inputs.
Quick Facts
| Full Name | Radial Basis Function Kernel |
|---|---|
| Specification | Official Specification |
How It Works
Translate distance into localized similarity
The standard form is k(x,y) = exp(-gamma ||x-y||^2). It equals one for identical vectors and remains positive for every finite distance. Because it depends only on x-y, it is shift-invariant; because it depends only on the norm, it is also isotropic. The Dive into Deep Learning Gaussian-process chapter visualizes how RBF length scale changes sampled-function geometry.
The kernel is positive definite on Euclidean space, so it defines an RKHS and can be used by compatible kernel algorithms. Local similarity is not probability: rows do not sum to one, and a value such as 0.8 has no universal semantic meaning across datasets or parameter settings.
Reconcile gamma, sigma, and length-scale conventions
Two common parameterizations are exp(-gamma r^2) and exp(-r^2 / (2 ell^2)), giving gamma = 1 / (2 ell^2). Some libraries use sigma for ell, while others define it differently. Copying a numeric value without its formula can therefore change the kernel substantially.
Current scikit-learn documentation uses the gamma form. Select the parameter inside a training-only protocol, and inspect similarity quantiles or the kernel spectrum rather than relying on a default detached from feature scale.
Audit scaling, concentration, and numerical behavior
Large-variance coordinates dominate squared distance unless preprocessing gives them an intentional scale. In high dimensions, distances may concentrate, causing most off-diagonal values to collapse into a narrow range. Missing-value imputation, categorical encoding, and duplicate records also alter the geometry.
Very small gamma makes the Gram matrix nearly constant; very large gamma makes it approach identity for distinct points. Compare validation quality across scales, monitor conditioning and effective rank, preserve the preprocessing contract, and use Nyström or Random Fourier approximations only after measuring task-level error.
Key Characteristics
- Maps squared Euclidean distance to a similarity between zero and one
- Defines a stationary and isotropic positive-definite kernel
- Induces an infinite-dimensional feature space without explicit coordinates
- Uses gamma or an equivalent length scale to set locality
- Depends critically on feature scaling and distance concentration
- Can become nearly constant or nearly diagonal at extreme bandwidths
Common Use Cases
- Creating nonlinear boundaries for kernel classification
- Fitting smooth nonlinear functions with kernel regression
- Building affinity matrices for spectral clustering
- Measuring distributions and dependence with MMD or HSIC
- Benchmarking Random Fourier and Nyström kernel approximations
Example
Loading code...Frequently Asked Questions
What does gamma control in the RBF Kernel?
Gamma controls how quickly similarity decays with squared distance. A larger gamma creates narrower neighborhoods; a smaller gamma makes distant points remain similar. Its useful range depends on feature scaling and the observed distance distribution.
How are RBF gamma and length scale related?
For `exp(-gamma r^2)` and `exp(-r^2 / (2 ell^2))`, the relationship is `gamma = 1 / (2 ell^2)`. Always verify the exact library formula because symbols such as sigma and length scale are not used consistently.
Why should features be scaled before using an RBF Kernel?
Squared Euclidean distance combines all coordinates. A feature with much larger numeric variation can dominate the kernel even when it is not more important. Scaling should reflect intended geometry and be fitted only on training data.
Is the RBF Kernel always a good default?
It is a useful nonlinear baseline for numerical vectors, not a universal choice. It assumes Euclidean, isotropic locality and can fail on structured, sparse, mixed-unit, or high-dimensional data. Compare it with linear and domain-specific kernels.
How can the RBF Kernel scale to more samples?
Use block computation when exact values are still feasible, or evaluate Nyström and Random Fourier Features as approximations. Validate kernel error and downstream quality on important slices, because faster pairwise estimates do not guarantee preserved decisions.