What is Metropolis-Hastings?

Metropolis-Hastings is a Markov Chain Monte Carlo algorithm that proposes a candidate state and accepts it with a probability correcting for both the target-density ratio and proposal asymmetry.

Quick Facts

SpecificationOfficial Specification

How It Works

Correct the proposal with an acceptance ratio

For current state x, proposal x' ~ q(x' | x), and unnormalized target pi, MH accepts with alpha = min(1, pi(x') q(x | x') / (pi(x) q(x' | x))). The unknown normalizing constant of pi cancels. If q is symmetric, the proposal terms cancel; otherwise omitting them changes the stationary distribution.

Hastings' 1970 paper generalized the original symmetric Metropolis rule and discussed estimation error. Detailed balance establishes invariance for this construction, but finite-chain quality still depends on reachability and mixing.

Tune movement without targeting an acceptance number blindly

A narrow random-walk proposal accepts often but moves slowly; a wide proposal may be rejected repeatedly. The useful scale depends on dimension, posterior geometry, parameterization, and proposal family. Acceptance rate is therefore a symptom rather than a universal optimization target.

Independent, block, component-wise, and adaptive proposals have different costs. Adaptation should be confined to a valid warmup scheme, and constrained parameters should be proposed on an appropriate transformed space with the required Jacobian included in the target.

Audit support, modes, and Monte Carlo error

The proposal must be able to reach every consequential target region. A local random walk can look stable while remaining in one mode, and a high acceptance rate can coexist with negligible exploration. Run multiple chains from dispersed states and inspect trace behavior, jump distances, rejection runs, and mode occupancy.

Report Effective Sample Size and Monte Carlo standard error per estimand, not only iteration count. Compare proposal scales and parameterizations on a fixed target, and validate moments or probabilities against analytic truth or an independent method whenever possible.

Key Characteristics

  • Uses proposal, target, and reverse-proposal density ratios
  • Works with targets known only up to a constant
  • Retains the current state whenever a proposal is rejected
  • Includes random-walk and independent proposal variants
  • Can support discrete or nondifferentiable parameter spaces
  • Depends strongly on proposal support and mixing quality

Common Use Cases

  1. Sampling an unnormalized Bayesian posterior
  2. Updating discrete parameters unavailable to gradient samplers
  3. Embedding custom proposals inside larger MCMC schemes
  4. Building Metropolis-within-Gibbs samplers
  5. Providing a transparent baseline for advanced samplers

Example

loading...
Loading code...

Frequently Asked Questions

Why does Metropolis-Hastings not need the target normalizing constant?

The acceptance rule uses a ratio of target densities at the proposed and current states. Any common multiplicative normalizing constant cancels, although the remaining log density and proposal probabilities must still be evaluated correctly.

When can the proposal ratio be omitted?

Only when the proposal is symmetric for the relevant states, so q(x'|x) equals q(x|x'). Independent, log-normal, boundary-aware, and many state-dependent proposals are asymmetric and require the Hastings correction.

What acceptance rate should an MH sampler target?

There is no universal rate across dimensions, proposal families, and targets. Optimize effective samples or Monte Carlo error per unit time while checking exploration, rather than tuning solely toward a remembered acceptance percentage.

Should rejected Metropolis-Hastings states be deleted?

No. Rejection means the next chain state equals the current state. Removing repeated states changes empirical frequencies and biases estimates because dwell time is part of how the chain represents the target distribution.

How can an MH chain appear converged but still be wrong?

A local proposal may explore one mode smoothly without reaching another, or proposal support may exclude valid regions. Use dispersed chains, mode-specific checks, ESS and Monte Carlo error, sensitivity to proposals, and known-truth simulations.

Related Terms

Related Articles