What is Maximum Mean Discrepancy?
Maximum Mean Discrepancy (MMD) is an integral probability metric that measures the largest expectation difference over a chosen function class and, with an RKHS unit ball, equals the distance between two kernel mean embeddings.
Quick Facts
| Created | Developed as a kernel two-sample method by Gretton and collaborators from 2006 to 2012 |
|---|---|
| Specification | Official Specification |
How It Works
Distinguish the discrepancy from the hypothesis test
The population discrepancy is sup_{||f||_H <= 1} |E_P[f] - E_Q[f]|, which equals ||mu_P - mu_Q||_H. With a characteristic kernel, it is zero exactly when the two population distributions are equal on the stated domain.
The kernel two-sample test uses an empirical MMD statistic and calibrates a rejection threshold. A large score alone is not a p-value, and a non-rejection does not establish practical equivalence.
Choose the estimator and null calibration explicitly
The biased V-statistic includes diagonal kernel terms and is nonnegative. The unbiased U-statistic removes within-sample diagonals, is unbiased for squared population MMD, and can be negative in finite samples. Linear-time and incomplete estimators trade computation for additional variance.
Permutation calibration relies on exchangeability under the null. Time series, repeated users, paired samples, clusters, survey weights, or adaptive windows violate naive row shuffling. Use a design-preserving bootstrap or permutation scheme and report the estimator, resampling unit, repetitions, and random seed.
Validate kernels, power, and monitoring policy
Kernel bandwidth determines which scales of difference the test can detect. Very narrow kernels emphasize near-duplicates; very wide kernels approach coarse mean-like comparisons. Median-distance heuristics are starting points, not proofs of sensitivity.
Estimate power on changes that matter operationally, include no-change controls, and test multiple bandwidths with a correction or predefined aggregation rule. In drift monitoring, also account for repeated testing, reference-window reuse, delayed labels, alert cooldowns, and whether a detected distribution change affects model utility.
Key Characteristics
- Measures distance between two RKHS mean embeddings
- Avoids explicit probability-density estimation
- Identifies unequal distributions only with a suitable characteristic kernel
- Offers biased, unbiased, linear-time, and incomplete empirical estimators
- Requires calibrated null inference for a statistical test
- Is sensitive to bandwidth, preprocessing, dependence, and sample size
Common Use Cases
- Testing whether two independent samples follow the same distribution
- Monitoring multivariate feature or embedding drift
- Evaluating simulated or generated samples against a reference population
- Defining distribution-alignment losses in domain adaptation
- Comparing experimental cohorts when density estimation is impractical
Example
Loading code...Frequently Asked Questions
What does Maximum Mean Discrepancy measure?
MMD measures the largest expectation difference over the unit ball of a selected RKHS, equivalently the distance between two kernel mean embeddings. It detects only differences represented at scales and structures supported by that kernel.
Is MMD a distance or a statistical test?
Population MMD is a discrepancy and can be a metric on distributions for a characteristic kernel. An empirical MMD value becomes a test only after specifying the null, estimator, threshold calibration, significance level, and sampling assumptions.
Why can unbiased MMD squared be negative?
The unbiased U-statistic removes diagonal self-similarities so its expectation matches population MMD squared. Finite-sample fluctuations can make that estimate negative even though the population quantity is nonnegative. Do not silently clamp it before inference.
How should an RBF bandwidth be chosen for MMD?
Choose it against the scales of change that matter and evaluate sensitivity across a predefined grid. The median heuristic is a baseline, but it can lose power in high dimensions or heterogeneous data. Avoid selecting a bandwidth on the test outcome without correction.
Can MMD alone diagnose harmful model drift?
No. It can reveal a distribution difference, not its cause or business impact. Pair it with feature slices, model-performance or label analysis, reference-quality checks, repeated-testing control, and a response policy tied to operational risk.