What is t-SNE?
t-SNE (t-Distributed Stochastic Neighbor Embedding) is a nonlinear visualization method that places samples in two or three dimensions by matching neighborhood-probability distributions.
Quick Facts
| Full Name | t-Distributed Stochastic Neighbor Embedding |
|---|---|
| Created | 2008 by Laurens van der Maaten and Geoffrey Hinton |
| Specification | Official Specification |
How It Works
Match Gaussian neighborhoods with a heavy-tailed map
For each input point, t-SNE chooses a Gaussian bandwidth whose probability entropy corresponds to the requested Perplexity, then symmetrizes the conditional probabilities. In the low-dimensional map it uses a one-degree-of-freedom Student t distribution. The heavy tail reduces the crowding problem by allowing moderately dissimilar points to separate without forcing true neighbors apart.
The original t-SNE paper defines this probability-matching objective. Because KL divergence is asymmetric, missing a high-probability neighbor is penalized more heavily than placing unrelated points somewhat too close, which explains the method's local emphasis.
Run a parameter and optimization protocol
Perplexity is an effective neighborhood scale, not a cluster count, and must be smaller than the sample count. Initialization, Learning Rate, Early Exaggeration, stopping, distance metric, approximation, and random seed can all change the map. Current scikit-learn documentation also notes that Learning Rate conventions differ among implementations.
For very high-dimensional inputs, first fit PCA on dense data or Truncated SVD on sparse data using only the analysis training subset. Run multiple plausible Perplexities and seeds, require objective convergence, and retain every preprocessing and implementation setting with the artifact.
Validate neighborhoods outside the picture
Distill's controlled examples show that cluster size, inter-cluster distance, apparent gaps, and even noise patterns can be misleading. Evaluate trustworthiness or neighborhood recall across several k values, compare repeat runs, inspect known controls, and verify any clustering claim in the original representation with a separate algorithm and metric.
Classic t-SNE is transductive: it optimizes coordinates for the fitted sample rather than learning a simple reusable mapping. Adding records can move existing points, and the common scikit-learn estimator exposes fit_transform rather than a general transform. Use an explicit parametric or landmark method when stable out-of-sample coordinates are a requirement.
Key Characteristics
- Optimizes a non-convex KL divergence between neighborhood probabilities
- Uses per-point Gaussian bandwidths controlled through Perplexity
- Uses a heavy-tailed Student t distribution in the map
- Emphasizes local neighbor recall rather than global metric fidelity
- Depends on initialization, optimization settings, approximation, and random seed
- Usually produces a transductive visualization without reusable axes
Common Use Cases
- Exploring local neighborhoods in image, text, or biological representations
- Comparing representation models on a fixed diagnostic sample
- Inspecting suspected mixing, outliers, or batch effects before formal tests
- Displaying labeled evaluation data without using labels during fitting
- Generating hypotheses that will be validated in the original feature space
Example
Loading code...Frequently Asked Questions
Does a separate island in a t-SNE plot prove a cluster exists?
No. Optimization, Perplexity, density equalization, initialization, and sampling can create or split visible islands. Treat the map as a hypothesis generator, then evaluate clustering and separation in the original representation with stability tests and domain evidence.
What does t-SNE Perplexity control?
Perplexity sets the entropy target used to choose each point's Gaussian bandwidth, so it behaves like a smooth effective neighborhood size. It is not the number of clusters. Compare several values below the sample count rather than tuning one plot for visual appeal.
Can distances and axes in a t-SNE plot be interpreted?
Axes have no original-feature meaning, and distances between well-separated groups may not represent high-dimensional distance. Local neighborhoods are the primary signal. Cluster area, density, orientation, and empty space can all be artifacts of the embedding.
Why should PCA sometimes run before t-SNE?
Reducing extremely high-dimensional dense data to a moderate dimension can suppress weak noise and reduce pairwise-distance cost. Fit that PCA only on the intended analysis or training subset and verify that it preserves the neighborhoods relevant to the study.
Can a fitted t-SNE model place new records on the same map?
Classic t-SNE is transductive and jointly optimizes coordinates for the fitted sample. Some specialized parametric or landmark implementations support out-of-sample placement, but their mapping and validation contract is different and must be versioned explicitly.