What is Adaptive MCMC?

Adaptive MCMC is a class of sampling algorithms that updates parameters of a Markov transition using information collected during the run while controlling that adaptation so the target distribution remains valid.

Quick Facts

SpecificationOfficial Specification

How It Works

Treat the tuning state as part of the algorithm

At iteration n, an adaptive sampler chooses a kernel P_gamma_n using a tuning state such as a running covariance or log proposal scale. The next sample changes both the physical state and potentially gamma_n; the process is therefore not a homogeneous Markov chain in the sampled state alone.

Useful adaptation targets movement, not a remembered acceptance number in isolation. Scale, correlation, tails, constraints, dimension, and the estimand all affect efficiency, so a controller should have explicit bounds and observability.

Preserve convergence while kernels change

Roberts and Rosenthal's ergodicity analysis shows why changing among individually valid kernels is not sufficient. Diminishing Adaptation controls how much successive kernels differ, while Containment prevents the process from drifting toward kernels with arbitrarily poor convergence.

These are theoretical conditions, not log fields that can be declared true after a short run. A conservative implementation adapts during warmup, freezes the kernel before retained sampling, and restarts rather than silently retuning after a regime change.

Audit both adaptation and retained sampling

Record the complete tuning trajectory, update schedule, bounds, target statistic, accepted moves, warmup length, final kernel, random seed, and software revision. Compare the frozen kernel across dispersed chains and check whether adaptation converged to similar scales rather than one chain hiding a pathological region.

Evaluate bulk and tail Effective Sample Size, Monte Carlo error, movement, and wall time for consequential estimands. An improved acceptance statistic can coexist with biased adaptation, mode loss, or worse tail exploration, so simulation truth and nonadaptive baselines remain valuable.

Key Characteristics

  • Updates proposal or transition settings from run history
  • Makes the sampled-state sequence nonhomogeneous during adaptation
  • Requires a validity argument beyond fixed-kernel invariance
  • Often uses diminishing step sizes or a finite warmup phase
  • Can learn scale, covariance, orientation, or population schedules
  • Needs adaptation traces in addition to ordinary MCMC diagnostics

Common Use Cases

  1. Learning a random-walk proposal scale during warmup
  2. Estimating a covariance matrix for correlated posterior parameters
  3. Tuning sampler settings across heterogeneous model geometries
  4. Reducing manual calibration in repeated Bayesian workflows
  5. Building bounded adaptive components inside population samplers

Example

loading...
Loading code...

Frequently Asked Questions

Why can adaptation invalidate an MCMC chain?

The transition at one iteration depends on earlier samples through the tuning state. Even if every fixed kernel preserves the target, arbitrary switching among those kernels can destroy convergence, so the adaptation rule itself needs justification.

What are Diminishing Adaptation and Containment?

Diminishing Adaptation requires successive kernels to become increasingly similar. Containment prevents the adaptive process from visiting tuning states whose convergence times become arbitrarily poor. Together they form a widely used sufficient framework, not an automatic empirical certificate.

Is warmup tuning the same as continuing Adaptive MCMC?

No. If adaptation stops after warmup, retained draws come from a fixed transition kernel and standard MCMC reasoning applies to that phase. Continuing adaptation changes the kernel throughout sampling and requires a valid adaptive construction.

Should Adaptive MCMC target a fixed acceptance rate?

An acceptance target may guide a specific proposal family under stated assumptions, but no rate is universally optimal. Evaluate effective samples and Monte Carlo error per unit time, movement across relevant regions, and tail behavior.

What must be recorded for an adaptive chain?

Record initial state, random seed, adaptation statistic, gain schedule, bounds, warmup length, full tuning trace, final frozen kernel, rejections, warnings, and the exact retained draws used for each reported estimand.

Related Terms

Related Articles