What is Variational Inference?

Variational Inference is a family of methods that approximates an intractable target distribution with a tractable parameterized distribution by solving an optimization problem.

Quick Facts

SpecificationOfficial Specification

How It Works

Derive the objective from the posterior identity

For joint density p(x,z) and approximation q(z), log p(x) = ELBO(q) + KL(q(z) || p(z|x)). Because the KL term is non-negative, maximizing E_q[log p(x,z) - log q(z)] tightens a lower bound on the log evidence and minimizes reverse KL within the chosen family. The Blei, Kucukelbir, and McAuliffe review develops this optimization view. The evidence is usually unavailable, so ELBO values alone do not reveal the remaining approximation gap.

Choose a variational family and gradient estimator deliberately

Mean-field VI factorizes latent variables and scales well, but cannot represent posterior correlations. Structured covariance, mixture, or normalizing-flow families add expressiveness at higher optimization and memory cost. Coordinate ascent is available for some conjugate models; black-box VI uses Monte Carlo gradients, control variates, and reparameterization where possible. Black Box Variational Inference reduces model-specific derivation, but stochastic-gradient variance and local optima remain operational concerns.

Diagnose optimization and approximation as separate failures

Run multiple initializations, inspect ELBO traces and gradient variance, and verify that samples and moments are finite. Then compare posterior and posterior predictive summaries with exact results on small cases or a trusted sampling baseline. Reverse KL often favors one high-density region and may underestimate variance when the true posterior is multimodal or strongly correlated. Validate downstream Log Loss, coverage, calibration, and decisions; optimizer convergence is not posterior correctness.

Key Characteristics

  • Approximates a target distribution with a chosen tractable family
  • Transforms probabilistic inference into numerical optimization
  • Commonly maximizes an ELBO equivalent to minimizing reverse KL
  • Supports stochastic gradients, mini-batches, and amortized inference
  • Introduces approximation bias determined by family and divergence
  • Requires separate optimization and predictive-quality diagnostics

Common Use Cases

  1. Scaling posterior inference in latent-variable and hierarchical models
  2. Training variational autoencoders with amortized latent inference
  3. Approximating weight posteriors in Bayesian Neural Networks
  4. Serving reusable approximate posteriors when repeated sampling is costly
  5. Comparing richer variational families against speed and memory budgets

Example

loading...
Loading code...

Frequently Asked Questions

What problem does Variational Inference solve?

VI approximates a posterior or other difficult conditional distribution when exact normalization and integration are infeasible. It replaces direct computation with optimization over a chosen distribution family, producing a tractable proxy for expectations and predictions.

What is the ELBO in Variational Inference?

The Evidence Lower Bound is `E_q[log p(x,z) - log q(z)]`. The log evidence equals the ELBO plus `KL(q || posterior)`, so maximizing the ELBO minimizes that reverse-KL gap within the selected variational family.

Is a higher ELBO always a better model?

Only under a controlled comparison. For the same joint model, data, variational family, estimator, and numerical convention, a higher ELBO indicates a tighter optimized bound. Across different data scaling, objectives, or estimators, raw ELBO values may not be comparable, and predictive quality still needs direct evaluation.

How does Variational Inference differ from MCMC?

VI solves an optimization problem for a tractable approximation and is often faster and easier to scale. MCMC constructs samples intended to approach the target distribution asymptotically but can be expensive and hard to diagnose. Both can fail, and their accuracy must be checked for the actual model.

Why can mean-field Variational Inference underestimate uncertainty?

Its factorization removes posterior correlations, and reverse KL strongly penalizes placing mass where the target density is low. The optimum can concentrate on one mode and be too narrow. Richer families may help, but only comparison with trusted references and predictive checks can establish adequacy.

Related Terms

Related Articles