What is UMAP?
UMAP (Uniform Manifold Approximation and Projection) is a nonlinear dimensionality-reduction method that represents input neighborhoods as a fuzzy weighted graph and optimizes a low-dimensional graph with similar memberships.
Quick Facts
| Full Name | Uniform Manifold Approximation and Projection |
|---|---|
| Created | 2018 by Leland McInnes, John Healy, Nathaniel Saul, and Lukas Grossberger |
| Specification | Official Specification |
How It Works
Construct a fuzzy neighborhood graph
The UMAP paper builds on manifold learning and fuzzy topological representations. In practical terms, each sample receives a local connectivity radius and smooth distance scale. Directed memberships decay beyond the nearest guaranteed connection, then fuzzy union combines both directions into one weighted edge.
The graph is only as meaningful as the input representation and metric. Euclidean distance on standardized measurements, cosine distance on normalized embeddings, and a domain-specific metric define different neighborhoods. Missing values, duplicates, batch effects, and insufficient sampling can change the topology before optimization begins.
Separate neighborhood scale from layout compactness
According to the current umap-learn parameter guide, n_neighbors controls how locally the manifold is estimated: smaller values emphasize fine neighborhoods, while larger values use broader context. min_dist controls how tightly points may pack in the output; it does not discover a natural cluster radius.
Run a declared grid of metrics, neighborhood sizes, output dimensions, and seeds. Use the same preprocessing and sample identity for comparison, save the implementation version, and inspect stability rather than selecting the most visually separated plot.
Version reproducibility and out-of-sample behavior
UMAP uses randomness in approximate search and optimization. The project's reproducibility guide explains that exact seeded output can require giving up some multithreaded execution. A seed makes one implementation run repeatable; it does not prove that the recovered structure is stable under resampling or parameter changes.
The reference implementation can transform new records relative to a frozen training embedding, unlike classic t-SNE, but this remains a learned approximation. Fit preprocessing and UMAP on training data only, monitor drift and neighbor recall, and do not cluster a two-dimensional map without validating the resulting partition against the original representation.
Key Characteristics
- Builds a locally scaled fuzzy nearest-neighbor graph
- Optimizes low-dimensional memberships with attractive and repulsive forces
- Uses n_neighbors for neighborhood scale and min_dist for output compactness
- Supports multiple input metrics and more than two output dimensions
- Is stochastic and can trade exact reproducibility for parallel performance
- Can transform new samples in common implementations but still requires drift validation
Common Use Cases
- Exploring local structure in large embedding or representation datasets
- Creating a nonlinear diagnostic view before formal cluster evaluation
- Reducing dimensions for a downstream experiment with held-out validation
- Comparing neighborhood stability across model or dataset versions
- Placing compatible new samples relative to a frozen reference embedding
Example
Loading code...Frequently Asked Questions
What do n_neighbors and min_dist mean in UMAP?
`n_neighbors` sets the scale used to estimate local structure: smaller values focus on finer neighborhoods and larger values use broader context. `min_dist` controls output packing. Neither parameter is a cluster count or an automatically learned truth.
Does UMAP preserve global distances and cluster sizes?
Not reliably. UMAP can retain more broad organization than some local visualization methods on some datasets, but axes, absolute distance, island area, density, and empty space remain distorted by the graph construction and optimization.
Can UMAP transform records that were not in the fit set?
The reference implementation supports transforming compatible new records relative to a fitted graph and embedding. Freeze preprocessing, metric, model version, parameters, and reference sample, then evaluate neighborhood recall and downstream performance under drift.
Is setting random_state enough to make UMAP conclusions stable?
It can reproduce one supported execution path, sometimes with reduced parallelism, but it does not establish scientific stability. Repeat across seeds, resamples, parameter ranges, and implementation versions, and compare neighborhoods rather than only plot appearance.
Should clustering be performed directly on a 2D UMAP map?
Usually not as the only evidence. Two dimensions deliberately distort structure and density. If UMAP is part of a clustering pipeline, use a justified higher output dimension, tune only inside training data, and validate the final partition in the original representation and on held-out data.