What is Adjusted Rand Index?

Adjusted Rand Index is a symmetric measure of agreement between two hard partitions that counts whether sample pairs stay together or apart and subtracts the agreement expected under a fixed-marginal random model.

Quick Facts

Created1985 by Lawrence Hubert and Phipps Arabie
SpecificationOfficial Specification

How It Works

Compute from the full contingency table

Hubert and Arabie's Comparing Partitions established the widely used chance-adjusted form. Build a contingency table between partitions, sum choose(n_ij,2) within cells, and derive row and column pair totals. Mapping each cluster to the majority reference class discards unmatched groups and is not ARI.

Interpret chance correction under its null model

The expected agreement assumes partitions are randomized while their marginal cluster sizes remain fixed. This makes the baseline useful, but not assumption-free. Highly skewed cluster sizes change pair weighting, and a large cluster contributes quadratically many pairs. Publish cluster counts, size distributions, sample alignment, uncertainty, and any omitted or noise-labeled records.

Separate partition agreement from semantic correctness

Current scikit-learn documentation confirms symmetry, label-permutation invariance, a perfect score of 1, and possible negative values. A reference taxonomy may encode only one valid grouping, so low ARI does not prove useless clusters and high ARI does not prove fairness, stability, or downstream benefit.

Key Characteristics

  • Compares co-clustering decisions over all unordered sample pairs
  • Is invariant to arbitrary permutations of cluster label values
  • Is symmetric in the two partitions being compared
  • Corrects expected pair agreement under fixed cluster-size marginals
  • Can be negative when agreement is worse than the chance baseline
  • Applies to hard exhaustive partitions, not soft memberships by default

Common Use Cases

  1. Benchmarking a clustering against a trusted reference partition
  2. Measuring stability across seeds, resamples, or model revisions
  3. Comparing two human or automated taxonomy assignments
  4. Detecting split and merge changes through pair disagreement
  5. Reporting a chance-adjusted external clustering metric

Example

loading...
Loading code...

Frequently Asked Questions

How is Adjusted Rand Index calculated?

Create the contingency table between two aligned partitions and count sample pairs that share each row, column, and cell. ARI subtracts the pair agreement expected under fixed marginal cluster sizes and divides by the remaining distance to maximum agreement.

Why is Adjusted Rand Index invariant to cluster label names?

It evaluates whether pairs are placed together or apart, not whether literal IDs match. Renaming cluster `0` to `blue` changes no pair relationship, so a perfectly relabeled partition still scores 1.

What does a negative Adjusted Rand Index mean?

It means the observed pair agreement is below the expectation of the fixed-marginal random model. The exact attainable lower value depends on sample and cluster-size structure, so negative scores should not be interpreted with a universal quality band.

How does ARI differ from classification accuracy?

Accuracy compares named predictions to named targets one sample at a time. ARI compares entire partitions through pair relationships and ignores label permutations. It can compare different numbers of clusters without first forcing a one-to-one class mapping.

Can a high ARI prove that a clustering is useful?

No. It proves agreement with one supplied partition under one sample set. The reference may be noisy or represent a different notion of similarity. Add label-quality review, stability, slice results, and downstream utility before making a deployment decision.

Related Terms

Related Articles