What is Silhouette Score?
Silhouette Score is an internal cluster-validity measure that compares each sample's average distance to its own cluster with its average distance to the nearest alternative cluster.
Quick Facts
| Created | 1987 by Peter J. Rousseeuw |
|---|---|
| Specification | Official Specification |
How It Works
Inspect samples and clusters before trusting the mean
Rousseeuw's original paper introduced the silhouette as a graphical diagnostic, not merely one scalar. Review each cluster's coefficient distribution, width, negative tail, and nearest rival. A high global mean can coexist with a small diffuse cluster, while averaging runs can conceal assignment instability.
Declare the representation, distance, and singleton rule
The score changes when features are scaled, embeddings are normalized, or Euclidean distance is replaced with cosine or a precomputed dissimilarity. Current scikit-learn documentation requires between 2 and n-1 labels and supports explicit distance metrics. A singleton cluster has no within-cluster mean distance; implementations commonly assign its sample a neutral coefficient of 0, which must be recorded.
Use selection data separately from confirmation data
Choosing the algorithm, feature pipeline, or number of clusters that maximizes Silhouette Score on one dataset optimizes to that dataset's geometry. Recompute the frozen choice on held-out or resampled data, report variation across seeds, and add expert review or downstream outcomes. Pairwise distance computation can be quadratic in sample count, so any sampling approximation also needs a fixed design and uncertainty estimate.
Key Characteristics
- Produces a per-sample coefficient from cohesion and nearest-cluster separation
- Ranges from -1 to 1 under the standard definition
- Requires at least two clusters and fewer clusters than samples
- Depends on the feature representation, scaling, and distance function
- Can expose boundary and potentially misassigned samples through a profile plot
- Does not establish cluster stability, semantics, or downstream usefulness
Common Use Cases
- Comparing candidate cluster counts on the same representation and dataset
- Finding samples that lie near a competing cluster
- Diagnosing a small or diffuse cluster hidden by a global average
- Checking clustering stability across seeds and resamples
- Complementing reviewer judgments in document or embedding clustering
Example
Loading code...Frequently Asked Questions
How is Silhouette Score calculated?
For each sample, compute its mean distance to the rest of its own cluster, `a`, and its smallest mean distance to another cluster, `b`. The coefficient is `(b-a)/max(a,b)`. The usual overall score is the sample mean, but per-cluster distributions should also be inspected.
What does a negative Silhouette Score mean?
It means the sample is closer on average to a different cluster than to its assigned cluster under the selected distance. That may indicate a bad assignment, an overlapping boundary, an unsuitable representation, or a non-convex structure rather than a simple labeling error.
Can Silhouette Score choose the correct number of clusters?
It can rank candidate partitions on the same data and geometry, but it cannot prove that one cluster count is semantically correct. Treat the selected count as a model choice and confirm it with resampling stability, domain review, and downstream utility.
Why does feature scaling change Silhouette Score?
The coefficient is built from distances. A feature with a larger numeric scale can dominate those distances, and normalization can change nearest clusters. Freeze the representation, scaling, missing-value policy, and distance function before comparing scores.
Is a high average Silhouette Score enough to approve a clustering?
No. It measures compactness and separation under one geometry. Inspect per-sample and per-cluster values, stability across seeds and samples, outliers, protected or rare slices, expert judgments, and whether the clusters improve the intended downstream task.