What is Parallel Tempering?

Parallel Tempering is a population Markov Chain Monte Carlo method that runs replicas at different temperatures and swaps their states so the target-temperature chain can traverse separated modes.

Quick Facts

SpecificationOfficial Specification

How It Works

Run replicas across a temperature ladder

A common ladder uses inverse temperatures 1 = beta_0 > beta_1 > ... > beta_K > 0 and replica targets proportional to pi(x)^beta_k. Lower beta flattens density differences, so a hot replica can cross valleys that trap the cold chain. In Bayesian work, tempering the full posterior and tempering only the likelihood define different intermediate distributions.

Hukushima and Nemoto's Exchange Monte Carlo paper demonstrates simultaneous replicas and configuration exchanges for a slowly relaxing system. The mechanism is general; its efficiency is not guaranteed for every multimodal geometry.

Preserve the joint target with swap corrections

Each replica performs local transitions valid for its own tempered target. A proposed exchange between adjacent states x_i and x_j is accepted with the Metropolis ratio for the product distribution; for power tempering its log ratio is (beta_i - beta_j)(log pi(x_j) - log pi(x_i)).

Only samples currently associated with beta = 1 represent the original target. Pooling hot-replica draws into posterior summaries without reweighting is biased. Replica identities may move across temperatures, so implementations must distinguish a physical state from a temperature slot.

Diagnose communication, not just local acceptance

Adjacent swap rates reveal whether neighboring distributions overlap, but a uniform-looking swap rate does not prove global travel. Track temperature-index traces, complete cold-to-hot-to-cold round trips, mode occupancy at the cold target, local-kernel ESS, and wall-clock cost across the full population.

A ladder that is too sparse blocks exchanges; an unnecessarily dense ladder wastes computation. Adapt the ladder only under a justified schedule or during warmup, freeze the retained-sampling contract, and compare against dispersed independent chains and other global samplers.

Key Characteristics

  • Runs a population of replicas at different temperatures
  • Flattens barriers for hot replicas while retaining a cold target
  • Exchanges adjacent states with a Metropolis correction
  • Preserves a product distribution across temperature slots
  • Requires a temperature ladder and local kernel for every replica
  • Targets multimodal exploration at substantial compute cost

Common Use Cases

  1. Posterior distributions with well-separated modes
  2. Energy-based models with rugged probability landscapes
  3. Bayesian models affected by label or symmetry modes
  4. Scientific inverse problems with local posterior basins
  5. Stress-testing whether ordinary chains miss important regions

Example

loading...
Loading code...

Frequently Asked Questions

Why does Parallel Tempering help with multimodal targets?

Hot replicas flatten density differences and can cross barriers more often. Valid swaps transfer explored states between neighboring temperatures, allowing the cold replica to visit modes that its local transition may rarely reach alone.

Can samples from every temperature be pooled?

No. Only the state in the target-temperature slot follows the original distribution directly. Hot replicas follow tempered distributions; pooling them without a valid reweighting scheme changes the estimand.

How should a temperature ladder be chosen?

Neighboring replicas need enough overlap to exchange states while the hottest distribution must reduce relevant barriers. Monitor adjacent swaps and full round trips, then account for total compute rather than targeting one universal swap rate.

Is Parallel Tempering the same as Simulated Annealing?

No. Parallel Tempering maintains multiple equilibrium targets and uses corrected exchanges to sample the cold distribution. Simulated Annealing changes one temperature over time primarily to seek optima and does not generally produce posterior samples.

Do frequent swaps prove all modes were explored?

No. Replicas may swap frequently within the same region, and the hottest chain may still miss isolated modes. Inspect round trips, cold-chain mode occupancy, between-chain agreement, ESS, and known simulation truth where available.

Related Terms

Related Articles