What is Bayesian Quadrature?
Bayesian Quadrature is a probabilistic numerical-integration method that places a prior, commonly a Gaussian Process, over an unknown integrand and conditions on function evaluations to obtain a posterior distribution over the integral.
Quick Facts
| Specification | Official Specification |
|---|
How It Works
Integrate the GP posterior analytically
For I = integral f(x) p(x) dx, a GP prior on f induces a Gaussian prior on I. Given function values y at nodes X, the posterior mean is m_I = z^T (K + noise)^-1 y, where z_i = integral k(x_i, x) p(x) dx is a kernel mean. The posterior variance is k_bar - z^T (K + noise)^-1 z, with k_bar the kernel integrated over both arguments.
These quantities are tractable only when kernel means can be computed or accurately approximated. The Bayesian quadrature convergence analysis makes explicit that rates depend on smoothness, design, and prior assumptions.
Select evaluations for integral accuracy
Nodes may be fixed, sampled from the integration measure, or chosen adaptively to reduce posterior integral variance. Unlike Bayesian optimization, the objective is to estimate an integral, not locate the maximum of the integrand. An acquisition rule must therefore be judged by integration error and uncertainty calibration rather than best observed value.
For noisy or stochastic evaluations, include the observation variance in the likelihood and repeat selected nodes when useful. Hyperparameters learned from the same small evaluation set can make posterior variance overconfident, so sensitivity analysis or hierarchical inference matters.
Check misspecification and dimensional scaling
BQ can be sample-efficient for smooth, expensive, low-dimensional integrands with kernels whose means are available. It degrades when the integrand has discontinuities, sharp local structure, heavy tails, an incorrectly specified measure, or too many active dimensions. Transformations and structured kernels may help but must preserve the target integral.
When a reference is feasible, evaluate absolute and relative error, credible-interval coverage, interval width, node budget, and wall-clock cost over a suite of integrands. Compare randomized quasi-Monte Carlo, classical quadrature, and plain Monte Carlo; a narrow BQ interval without empirical coverage is not sufficient evidence.
Key Characteristics
- Models the integrand as a random function
- Returns a posterior mean and variance for an integral
- Uses kernel means and an integrated kernel covariance
- Can adapt evaluation nodes to reduce integral uncertainty
- Makes smoothness and measure assumptions explicit
- Can become overconfident under kernel misspecification
Common Use Cases
- Integrating outputs from expensive scientific simulations
- Estimating expectations with few deterministic evaluations
- Approximating marginal likelihoods in low dimensions
- Propagating numerical-integration uncertainty into decisions
- Comparing adaptive sampling with Monte Carlo baselines
Example
Loading code...Frequently Asked Questions
How does Bayesian Quadrature differ from classical quadrature?
Classical rules return a weighted estimate and may provide deterministic error bounds under smoothness assumptions. BQ derives weights and an integral distribution from a probabilistic prior, making assumptions explicit but requiring calibration under misspecification.
How does Bayesian Quadrature differ from Bayesian Optimization?
BQ estimates an integral and selects points to reduce uncertainty about that integral. Bayesian Optimization seeks an optimum and chooses points according to potential objective improvement. They may share GP machinery but solve different decision problems.
What is a kernel mean in Bayesian Quadrature?
A kernel mean integrates one kernel argument against the target measure: z_i = integral k(x_i, x)p(x)dx. It converts observed function values into weights for the posterior integral mean and is closely related to kernel mean embeddings.
Does Bayesian Quadrature posterior variance guarantee integration error?
No. It is conditional on the GP prior, kernel hyperparameters, integration measure, likelihood, and node policy. Empirical coverage over representative integrands or theoretical conditions are needed before treating it as a reliable error bar.
When is Bayesian Quadrature a poor choice?
It is often weak for high effective dimension, discontinuities, sharp unmodeled features, heavy tails, misspecified measures, or cheap integrands where Monte Carlo offers many more evaluations. Compare accuracy and coverage per unit cost.