What is Laplace Approximation?
The Laplace Approximation is a local method that approximates a smooth probability distribution with a Gaussian centered at a mode and uses local curvature to determine its covariance.
Quick Facts
| Specification | Official Specification |
|---|
How It Works
Start from a valid MAP solution and curvature
Fit the likelihood and prior jointly, preserve their scaling, and verify that the selected point is a meaningful local optimum. Around theta_MAP, the log posterior is approximated quadratically and the Gaussian covariance is the inverse negative Hessian. If the Hessian is indefinite, singular, or dominated by arbitrary regularization, a naive covariance is invalid. Damping and generalized Gauss-Newton or Fisher approximations change the object and must be reported.
Choose parameter and curvature structure as part of the method
Laplace Redux separates four practical choices: which weights receive a posterior, which curvature structure is stored, how prior precision is selected, and how the posterior is mapped to predictions. Last-layer and diagonal variants reduce cost but discard dependencies; full or structured curvature retains more geometry at greater memory and compute cost. These variants are not interchangeable implementations of one fixed posterior.
Test local fit and posterior predictions
Compare the quadratic approximation with exact or sampling results on a reduced problem, inspect sensitivity to prior precision and damping, and verify finite positive covariance factors. Then evaluate the posterior predictive distribution using proper scores, coverage, calibration, and shift slices. A tight Gaussian can indicate strong local curvature while missing another plausible mode, so numerical stability and narrow intervals do not establish global posterior accuracy.
Key Characteristics
- Centers a Gaussian approximation at a posterior mode
- Uses second-order local curvature to define posterior precision
- Can be applied after fitting a regularized deterministic model
- Offers full, diagonal, structured, subnetwork, and last-layer variants
- Cannot faithfully represent distant modes, strong skew, or heavy tails
- Requires prior, curvature, damping, and predictive choices to be versioned
Common Use Cases
- Adding local posterior uncertainty to a pretrained neural network
- Approximating parameter covariance in Bayesian logistic regression
- Comparing prior precision or model choices with approximate evidence
- Building uncertainty-aware predictions under a constrained compute budget
- Benchmarking a local Bayesian approximation against VI and ensembles
Example
Loading code...Frequently Asked Questions
How does the Laplace Approximation work?
It finds a posterior mode, expands the log posterior to second order around that point, and interprets the inverse local negative curvature as Gaussian covariance. Predictions then integrate or sample through this approximate Gaussian rather than using only the mode.
What is the difference between Laplace Approximation and Variational Inference?
Laplace fits a local Gaussian from mode and curvature, often after ordinary MAP training. VI optimizes parameters of a chosen distribution family under a divergence objective. Both are approximations; their geometry, cost, and failure modes differ.
Why use a last-layer Laplace Approximation?
Restricting uncertainty to the final layer greatly reduces curvature storage and computation and can be applied to a fixed feature extractor. It cannot represent uncertainty in learned features, so its speed comes with a clearly narrower posterior claim.
When does the Laplace Approximation fail?
It can misrepresent multimodal, skewed, heavy-tailed, weakly identified, or non-smooth posteriors. Neural-network symmetries and flat directions also make curvature difficult. Compare with sampling or another approximation on representative reduced cases.
Does a positive covariance make a Laplace model trustworthy?
No. Positive covariance only establishes a numerically valid local Gaussian. Trust also requires a defensible prior and likelihood, stable curvature estimation, posterior predictive checks, proper scores, coverage, and evaluation under realistic distribution shift.