What is Markov Chain Monte Carlo?

Markov Chain Monte Carlo is a family of algorithms that estimates expectations under a target distribution by constructing a Markov chain whose long-run stationary distribution is that target.

Quick Facts

SpecificationOfficial Specification

How It Works

Construct a chain with the intended invariant distribution

A Markov transition depends on the current state rather than the entire recorded path. MCMC designs that transition so the target pi is invariant. Detailed balance is a common sufficient construction, but it is not necessary; irreducibility, recurrence, and aperiodicity conditions determine whether the chain can approach the intended equilibrium.

The target often needs to be known only up to a constant, which is valuable for Bayesian posteriors. That mathematical property does not guarantee useful exploration: disconnected modes, constrained geometry, or a poor transition can trap finite chains.

Separate adaptation, warmup, and retained sampling

Initialization and adaptation affect early draws. Warmup can tune proposal scales or metrics and reduce dependence on initial states, but discarding a fixed prefix does not prove convergence. Adaptation that changes the transition should normally finish before retained sampling unless the method has a valid diminishing-adaptation construction.

Run multiple chains from dispersed, valid initial states and preserve seeds, warmup settings, transition parameters, rejections, and warnings. Thinning usually throws away information and is not a substitute for addressing autocorrelation or storage deliberately.

Diagnose estimands, not only chains

The Stan posterior-analysis guidance emphasizes that finite-chain convergence requires diagnostics. Inspect rank-normalized split R-hat, bulk and tail Effective Sample Size, Monte Carlo standard error, trace behavior, and sampler-specific warnings. Diagnostics near their preferred values are necessary checks, not proof that every mode was found.

Report uncertainty for each consequential transformed quantity, not only raw parameters. Compare replicated chains, alternative parameterizations, prior and posterior predictive checks, and where feasible an independent method or known simulation truth.

Key Characteristics

  • Generates dependent draws through a Markov transition kernel
  • Targets an invariant distribution known up to normalization
  • Estimates expectations with ergodic averages
  • Requires separate warmup and retained-sampling decisions
  • Uses multiple diagnostics because convergence is not observable directly
  • Can fail through poor mixing, multimodality, or invalid transitions

Common Use Cases

  1. Posterior inference for hierarchical probabilistic models
  2. Uncertainty propagation through complex latent-variable models
  3. Estimating expectations under unnormalized target densities
  4. Calibrating scientific models with correlated parameters
  5. Benchmarking approximate inference against posterior simulation

Example

loading...
Loading code...

Frequently Asked Questions

Why can MCMC use an unnormalized target density?

Many transition rules use ratios of target densities, so a common unknown normalizing constant cancels. The chain still needs a proper target and a valid transition; cancellation does not fix an improper posterior or poor exploration.

Are MCMC draws independent?

Usually not. Successive states are generated by a Markov transition and may be strongly autocorrelated. Effective Sample Size and Monte Carlo standard error quantify consequences for a particular estimand more honestly than raw draw count.

Does discarding warmup guarantee MCMC convergence?

No. Warmup can reduce initialization effects and tune a sampler, but no fixed discarded prefix proves equilibrium. Use multiple dispersed chains, rank-normalized split R-hat, bulk and tail ESS, trace checks, and method-specific diagnostics.

When should MCMC be preferred over Variational Inference?

Prefer MCMC when posterior fidelity and uncertainty are worth higher computation and the model can be diagnosed adequately. VI may be preferable for fast repeated inference or very large data, but its approximation bias should be assessed against simulation or MCMC where feasible.

Can good MCMC diagnostics prove the result is correct?

No. Diagnostics can reveal nonstationarity, autocorrelation, divergent transitions, and other failures, but they may miss isolated modes or model misspecification. Posterior predictive checks, sensitivity analysis, simulation-based calibration, and domain validation remain necessary.

Related Terms

Related Articles