What is Gaussian Process?
A Gaussian Process is a stochastic process whose function values at every finite collection of inputs have a joint multivariate Gaussian distribution specified by a mean function and covariance function.
Quick Facts
| Specification | Official Specification |
|---|
How It Works
Define a coherent distribution over function values
For inputs x1 ... xn, a GP requires [f(x1), ..., f(xn)] to be multivariate normal with mean entries m(xi) and covariance entries k(xi, xj). The kernel must generate a positive semidefinite covariance matrix for every finite input set. Williams and Rasmussen demonstrate function-space Gaussian priors whose predictions, conditional on fixed hyperparameters, can be computed with matrix operations.
Condition latent functions and observations separately
With Gaussian observation noise, conditioning yields a posterior mean and latent-function variance. A future observed target additionally includes observation-noise variance, so a latent credible band and a predictive interval are different objects. The Stan User's Guide separates the covariance kernel from diagonal noise and shows latent and marginalized formulations. Jitter added for numerical stability must not be silently interpreted as measured noise.
Validate kernels, computation, and uncertainty
Exact GP regression normally factorizes an n x n covariance matrix, requiring cubic time and quadratic storage in the observation count. Sparse, inducing-point, basis-function, state-space, or structured-kernel approximations trade accuracy and assumptions for scale. Inspect Cholesky failures, condition numbers, hyperparameter sensitivity, prior and posterior draws, residual structure, Log Predictive Density, interval coverage and width, and performance under representative shift; a narrow GP posterior is not automatically calibrated.
Key Characteristics
- Defines consistent Gaussian distributions for every finite input set
- Uses a mean function and positive semidefinite covariance kernel
- Produces joint posterior means, variances, and cross-covariances
- Separates latent-function uncertainty from observation noise
- Encodes strong assumptions about smoothness, distance, and stationarity
- Requires approximation or structure when exact covariance algebra is too costly
Common Use Cases
- Probabilistic regression with limited observations
- Surrogate modeling for Bayesian Optimization
- Spatial interpolation and geostatistical prediction
- Time-series or dynamical modeling with structured kernels
- Uncertainty-aware calibration of scientific simulations
Example
Loading code...Frequently Asked Questions
Why is a Gaussian Process called a distribution over functions?
A draw assigns a mutually consistent value to every input, so each draw can be viewed as one function. Computation only materializes finite collections of those values, whose joint distribution is multivariate Gaussian under the chosen mean and kernel.
Is Gaussian Process regression nonparametric?
It is commonly called nonparametric because model capacity can grow with the observations rather than being fixed by one finite coefficient vector. It still has kernel and likelihood hyperparameters, prior assumptions, and potentially many approximation parameters.
What does the kernel do in a Gaussian Process?
The kernel defines prior covariance between function values at pairs of inputs. Its geometry determines which points share information and encodes assumptions such as smoothness, periodicity, anisotropy, additivity, or stationarity. A convenient kernel can still be badly misspecified.
Why do exact Gaussian Processes scale poorly?
Exact inference usually performs a Cholesky factorization of the dense covariance matrix, which takes cubic time and quadratic memory in the number of observations. Structured matrices, sparse approximations, inducing points, or state-space representations can reduce cost.
Does Gaussian Process variance equal prediction error?
No. Posterior variance is conditional on the model, kernel, hyperparameters, likelihood, and data. Misspecification or distribution shift can make it overconfident. Check proper scores, empirical coverage, interval width, residuals, and representative out-of-sample slices.