What is Normalized Mutual Information?

Normalized Mutual Information is a symmetric partition-agreement measure that divides the mutual information between two hard cluster assignments by a declared function of their entropies.

Quick Facts

SpecificationOfficial Specification

How It Works

Name the normalization, logarithm, and degenerate convention

With arithmetic normalization, NMI=2I(U;V)/(H(U)+H(V)); other valid variants produce different finite-sample values. The logarithm base cancels only when the same base is used throughout. When both partitions contain one cluster, both entropies are zero and implementations usually define identical constant partitions as 1. Report the exact formula and library setting.

Do not confuse normalization with chance correction

Vinh, Epps, and Bailey analyze information-theoretic clustering measures and show why correction for chance matters when sample count is small relative to cluster count. NMI rescales MI but can still rise for fine random partitions. Adjusted Mutual Information subtracts expected MI under a declared random model and answers a different question.

Inspect the contingency table when NMI and ARI disagree

Current scikit-learn documentation exposes the arithmetic, geometric, minimum, and maximum entropy normalizers and confirms symmetry and permutation invariance. NMI measures entropy reduction, while ARI counts sample pairs; splitting a large class into pure subclusters can affect them differently. The contingency table reveals whether disagreement comes from splits, merges, or skew.

Key Characteristics

  • Measures shared information between two aligned hard partitions
  • Is invariant to arbitrary cluster label permutations
  • Is symmetric when the same normalization is used in both directions
  • Requires the entropy normalizer and degenerate-case convention to be declared
  • Usually ranges from 0 to 1 but is not adjusted for chance
  • Does not require equal cluster counts or one-to-one label matching

Common Use Cases

  1. Comparing document clusters with an editorial taxonomy
  2. Measuring agreement among community-detection partitions
  3. Checking clustering stability across runs or data perturbations
  4. Evaluating partitions with different numbers of clusters
  5. Reporting an information-theoretic metric beside ARI and AMI

Example

loading...
Loading code...

Frequently Asked Questions

How is Normalized Mutual Information calculated?

Build the contingency table for two aligned partitions, compute mutual information from joint and marginal probabilities, and divide by a declared entropy normalizer. With arithmetic normalization the formula is `2I(U;V)/(H(U)+H(V))`.

Is there one standard NMI formula?

No. Common implementations normalize MI by the minimum, maximum, geometric mean, arithmetic mean, or joint entropy. These variants share broad behavior but can return different numbers, so comparisons require the same formula and implementation.

What is the difference between NMI and AMI?

NMI rescales mutual information to a bounded range but does not remove agreement expected by chance. Adjusted Mutual Information subtracts an expected MI under a random-partition model, making it safer when sample counts are small or candidate cluster counts differ.

How does NMI differ from Adjusted Rand Index?

NMI measures reduction in label entropy, while ARI measures chance-adjusted agreement over sample pairs. Both ignore label names, but they weight split, merge, and cluster-size patterns differently. Report both and inspect the contingency table when they disagree.

Can high NMI prove that clusters are semantically correct?

No. It only shows that two partitions share information. A supplied reference can be noisy or encode another valid grouping, and fine partitions can inflate unadjusted NMI. Add chance-adjusted, stability, domain, and downstream evidence.

Related Terms

Related Articles