What is Kernel PCA?
Kernel PCA (Kernel Principal Component Analysis) is a nonlinear dimensionality-reduction method that performs principal component analysis in an implicit feature space defined by a positive-semidefinite kernel.
Quick Facts
| Full Name | Kernel Principal Component Analysis |
|---|---|
| Created | 1998 by Bernhard Schölkopf, Alexander Smola, and Klaus-Robert Müller |
| Specification | Official Specification |
How It Works
Center the kernel in feature space
The original Kernel PCA paper rewrites PCA using only inner products, replaces them with a kernel matrix K, and solves the resulting eigenproblem. Because PCA assumes centered observations, Kernel PCA must double-center the training Gram matrix with Kc = H K H, where H = I - 11^T / n.
New-sample kernels require the same training-set means, not centering against the incoming batch. Persist the preprocessing, training references or approximation basis, kernel parameters, centered eigenvectors, eigenvalues, and sign convention as one transform.
Treat the kernel as the model
An RBF bandwidth that is too small makes almost every sample isolated; one that is too large makes the Gram matrix nearly constant. Polynomial degree and offset can similarly dominate scale. Tune these choices inside the training protocol and inspect the kernel spectrum, effective rank, duplicate rows, and sensitivity across resamples.
The largest feature-space variance is not necessarily the most predictive or semantically important signal. Compare a linear PCA baseline, validate downstream metrics on held-out data, and avoid selecting parameters from the final two-dimensional picture.
Budget quadratic state and separate inverse reconstruction
A full Gram matrix uses quadratic memory in the number of training samples, while eigendecomposition can dominate runtime. Nyström and other low-rank approximations change the learned subspace and must be evaluated as separate artifacts.
Current scikit-learn KernelPCA documentation exposes kernels, eigensolvers, inverse fitting, and new-sample transformation. The inverse is a learned pre-image approximation because feature-space coordinates generally do not have an exact input-space inverse; reconstruction error therefore measures that auxiliary model as well as the embedding.
Key Characteristics
- Performs linear PCA in an implicit kernel-defined feature space
- Requires double-centering of the training Gram matrix
- Can represent nonlinear components without materializing feature coordinates
- Depends strongly on kernel family, scale, preprocessing, and sample coverage
- Usually requires quadratic kernel storage for an exact fitted sample set
- Has no general exact pre-image from retained components to input space
Common Use Cases
- Testing nonlinear structure after establishing a linear PCA baseline
- Extracting curved features from small or medium numerical datasets
- Denoising when a validated kernel captures the intended signal geometry
- Comparing kernel eigenspaces across controlled data revisions
- Teaching the relationship among kernels, Gram matrices, and PCA
Example
Loading code...Frequently Asked Questions
How is Kernel PCA different from ordinary PCA?
Ordinary PCA finds a linear subspace in the input coordinates. Kernel PCA performs the same variance analysis in an implicit feature space, so its input-to-component mapping can be nonlinear. That flexibility adds kernel selection, quadratic state, and pre-image problems.
Why must the Kernel PCA matrix be centered?
PCA assumes zero-mean observations in the space where covariance is measured. Kernel values are inner products of implicit feature vectors, so double-centering the Gram matrix subtracts their feature-space mean without explicitly constructing those vectors.
How should an RBF gamma value be chosen for Kernel PCA?
Choose gamma inside cross-validation or another training-only protocol using downstream and stability evidence. Inspect the kernel spectrum and pairwise similarities: extreme gamma values can make the matrix almost identity or almost constant, producing unhelpful components.
Can Kernel PCA transform unseen samples?
Yes when the implementation retains the fitted references or approximation basis. Compute kernels from the new sample to the training set, apply the training centering terms, and project with the frozen eigenvectors. Do not recenter against the new batch.
Can Kernel PCA reconstruct the original input exactly?
Not in general. Retained feature-space coordinates may have no exact input-space pre-image. An inverse transform is usually a separately fitted approximation whose errors depend on the kernel, retained components, regularization, and coverage of the training data.