What is Calinski-Harabasz Index?
Calinski-Harabasz Index is an internal cluster-validity measure that divides between-cluster dispersion per cluster degree of freedom by within-cluster dispersion per residual degree of freedom.
Quick Facts
| Created | 1974 by Tadeusz Caliński and Jerzy Harabasz |
|---|---|
| Specification | Official Specification |
How It Works
Keep the degrees-of-freedom correction in the formula
Caliński and Harabasz's original paper proposed the variance ratio as an informal indicator for the number of groups within a dendrite clustering method. The factors k-1 and n-k matter: omitting them produces a different statistic and makes additional clusters look artificially favorable.
Handle zero residual dispersion explicitly
The standard score requires 2 <= k <= n-1. If every member equals its cluster centroid, W is zero and the ratio is infinite or undefined depending on the implementation. Do not silently replace that condition with an ordinary finite score. Record duplicates, singleton clusters, numeric precision, preprocessing, and the exact handling of degenerate partitions.
Avoid interpreting a high ratio as semantic truth
Current scikit-learn documentation defines the score from between- and within-cluster dispersion. Because both use squared Euclidean distances to centroids, compact convex partitions are favored. Compare compatible candidates only, inspect other geometries and stability, and validate the intended use separately.
Key Characteristics
- Uses a degrees-of-freedom-adjusted ratio of between and within dispersion
- Has no fixed upper bound, with higher values preferred under one protocol
- Requires at least two clusters and at least one residual degree of freedom
- Can be computed efficiently from cluster counts, centroids, and sums of squares
- Favors structures represented well by squared Euclidean centroid geometry
- Does not provide a p-value or prove that clusters are meaningful
Common Use Cases
- Screening candidate cluster counts for centroid-based methods
- Comparing feature pipelines on one frozen evaluation dataset
- Tracking variance-ratio changes after representation updates
- Combining several internal indices during clustering model selection
- Flagging partitions with excessive within-cluster dispersion
Example
Loading code...Frequently Asked Questions
How is Calinski-Harabasz Index calculated?
Compute between-cluster sum of squares `B` and within-cluster sum of squares `W`, then divide `B/(k-1)` by `W/(n-k)`. The degrees-of-freedom terms are part of the metric and must not be dropped.
Is a higher Calinski-Harabasz Index always better?
Higher is preferred only among compatible candidates evaluated on the same records, features, scaling, and implementation. The score has no universal upper bound, and a high value can simply reward compact convex geometry that does not match the intended semantics.
Can Calinski-Harabasz Index compare different datasets?
Not as an absolute quality score. Sample size, dimensionality, scale, outliers, and cluster count all affect the ratio. Use it for relative model selection under a frozen protocol and preserve the complete context with every result.
What happens when within-cluster dispersion is zero?
The denominator becomes zero, so the variance ratio is infinite or undefined. This can occur with duplicated points or degenerate singleton-like partitions. Detect and report the condition instead of converting it into an ordinary finite score.
How should Calinski-Harabasz Index be used in production evaluation?
Use it beside Silhouette and Davies-Bouldin scores, seed and resample stability, cluster-size and outlier diagnostics, reference-partition measures when available, and evidence that the grouping improves a real downstream workflow.