What is Davies-Bouldin Index?
Davies-Bouldin Index is an internal cluster-validity measure that averages, for every cluster, its largest ratio of combined within-cluster scatter to centroid separation from another cluster.
Quick Facts
| Created | 1979 by David L. Davies and Donald W. Bouldin |
|---|---|
| Specification | Official Specification |
How It Works
Read the score as an average worst-rival ratio
Davies and Bouldin's original paper defines a family of separation measures and a cluster-similarity construction. The common index selects the largest similarity ratio for each cluster before averaging. One ambiguous neighbor can therefore dominate a cluster's contribution even when all its other neighbors are distant.
Treat coincident centroids and scale as protocol decisions
When two distinct clusters have the same centroid, their separation denominator is zero and the ratio is undefined or infinite unless an implementation applies a special convention. Feature scaling and outliers also move centroids and scatter. Record preprocessing, the exact scatter and centroid-distance definitions, noise-label handling, and the library version before interpreting a value.
Compare only compatible candidate partitions
Current scikit-learn documentation implements a nonnegative score where lower is better. The scale is not normalized, so values from different datasets, representations, or distance conventions are not directly comparable. Use it to rank compatible candidates, then add Silhouette profiles, stability, reference-label evidence where available, and domain utility.
Key Characteristics
- Averages one worst competing-cluster ratio for every cluster
- Combines within-cluster scatter with between-centroid separation
- Has a theoretical minimum of zero and no universal upper bound
- Uses only features and assignments, so it needs no reference labels
- Is efficient but inherits centroid and distance geometry assumptions
- Cannot establish semantic validity, stability, or production value
Common Use Cases
- Ranking candidate cluster counts under one fixed feature pipeline
- Detecting clusters with a close or highly overlapping centroid rival
- Monitoring compactness and separation after embedding revisions
- Comparing repeated centroid-based clustering runs
- Triaging partitions before deeper sample-level and domain review
Example
Loading code...Frequently Asked Questions
How is Davies-Bouldin Index calculated?
Measure each cluster's average member-to-centroid scatter. For every pair, divide their combined scatter by the distance between their centroids. Keep the largest ratio for each cluster and average those worst-rival ratios across all clusters.
Is a higher or lower Davies-Bouldin Index better?
Lower is better under a fixed evaluation protocol because it represents smaller within-cluster scatter relative to centroid separation. Zero is the theoretical minimum, but there is no universal threshold that transfers across datasets or feature spaces.
Can Davies-Bouldin Index compare scores from different datasets?
Not reliably. The value depends on representation, feature scaling, outliers, cluster count, scatter definition, and distance geometry. Compare candidate partitions on the same evaluated records with an identical preprocessing and scoring protocol.
Why can Davies-Bouldin Index fail on non-convex clusters?
It compresses every cluster into a centroid and average radius. A crescent, ring, manifold, or multimodal group may be coherent under connectivity or density while looking diffuse around its centroid, so the score can favor an incorrect convex partition.
Should Davies-Bouldin Index be the only clustering release metric?
No. Combine it with sample-level diagnostics such as Silhouette Score, variation across seeds and resamples, outlier analysis, reference-label measures when appropriate, expert review, and downstream outcome checks.