What is Gaussian Process Classification?

Gaussian Process Classification is a Bayesian classification model that places a Gaussian Process prior over a latent function and maps that function to class probabilities through a non-Gaussian likelihood.

Quick Facts

SpecificationOfficial Specification

How It Works

Separate the latent function from observed classes

For binary labels, a latent function f(x) receives a GP prior and a link such as p(y=1|f)=Phi(f) or sigmoid(f). Nearby inputs share information through the kernel, while the likelihood connects an unobserved score to a discrete label. Class labels are therefore not modeled as Gaussian observations.

Gaussian Processes for Machine Learning, Chapter 3 derives binary and multiclass classification and shows why the posterior is non-Gaussian. Changing the link changes the probability model, not merely the display scale.

Approximate a nonconjugate posterior

Laplace approximation fits a Gaussian around a posterior mode; expectation propagation matches local moments; variational inference optimizes a tractable lower bound. Each method can change posterior variance, evidence estimates, convergence behavior, and predictive probabilities. Record the inference scheme as part of the model contract.

Prediction integrates the link over the approximate latent posterior. For a probit link and Gaussian latent marginal, the integral is analytic: Phi(mu / sqrt(1 + variance)). Applying the link only to the posterior mean ignores uncertainty and can produce systematically sharper probabilities.

Validate class policy, calibration, and scale

Binary GPC does not define a multiclass decomposition by itself. One-versus-rest, one-versus-one, or a joint multiclass likelihood produce different training objectives and probability semantics. The scikit-learn GPC documentation exposes implementation-specific decomposition and kernel constraints that should not be generalized to every GPC.

Evaluate Log Loss, Brier Score, reliability diagrams, discrimination metrics, threshold cost, subgroup slices, and shift behavior. Exact binary inference inherits dense GP cost, while multiclass or large-data systems may require inducing variables or other approximations whose calibration must be reassessed.

Key Characteristics

  • Places a Gaussian Process prior over latent class scores
  • Uses a non-Gaussian class likelihood and link function
  • Requires approximate posterior inference in typical settings
  • Produces probabilities by integrating over latent uncertainty
  • Needs an explicit multiclass construction for more than two classes
  • Retains kernel assumptions and dense GP scaling unless approximated

Common Use Cases

  1. Probabilistic binary classification with limited labeled data
  2. Spatial or scientific classification with structured similarity
  3. Active learning driven by validated predictive uncertainty
  4. Nonlinear classification where kernel priors encode domain structure
  5. Comparing calibrated Bayesian classifiers under small-data regimes

Example

loading...
Loading code...

Frequently Asked Questions

Why is Gaussian Process Classification not analytically conjugate?

Discrete labels use Bernoulli or categorical likelihoods, commonly parameterized by probit, logistic, or softmax links, rather than a Gaussian likelihood. Combining that likelihood with a Gaussian Process prior does not yield a Gaussian posterior, so typical GPC systems use Laplace, expectation propagation, or variational approximation.

How does GPC turn a latent score into a probability?

It integrates a link function over the posterior distribution of the latent score. For a probit link with a Gaussian latent marginal, the result is Phi(mu divided by sqrt(1 plus variance)). Applying the link only to the mean discards posterior uncertainty.

Does Gaussian Process Classification support multiple classes?

Yes, but the construction must be explicit. Joint multiclass likelihoods, one-versus-rest, and one-versus-one decompositions have different objectives, computational costs, and probability semantics. Document the choice and evaluate the resulting probabilities, not only top-one accuracy.

Are GPC probabilities automatically calibrated?

No. They are conditional on the kernel, likelihood, inference approximation, hyperparameters, and data. Approximation error, misspecification, imbalance, and shift can cause miscalibration. Check Log Loss, Brier Score, reliability, subgroup behavior, and decision thresholds.

How does Gaussian Process Classification scale?

Dense binary methods generally inherit cubic time and quadratic memory in the number of observations, often with iterative approximate inference. Multiclass structure adds cost. Inducing variables, structured kernels, or explicit features can reduce cost but change approximation error.

Related Terms

Related Articles