t-SNE / UMAP Explorer
The t-SNE / UMAP Explorer provides nonlinear dimensionality reduction, visualization, clustering, manual selection, and spatial mapping of hyperspectral image data.
IDCubePro 2026 - t-SNE / UMAP Explorer
Overview
The t-SNE / UMAP Explorer provides nonlinear dimensionality reduction, visualization, clustering, manual selection, and spatial mapping of hyperspectral image data.
Each valid image pixel is represented by its spectrum and projected into a low-dimensional embedding space.
Pixels with similar spectral characteristics tend to appear near one another in the embedding.
t-distributed Stochastic Neighbor Embedding (t-SNE) is a nonlinear dimensionality-reduction method designed to preserve local relationships between observations.
t-SNE is particularly useful for visualizing groups or populations of pixels with similar spectral characteristics.
Uniform Manifold Approximation and Projection (UMAP) is a nonlinear dimensionality-reduction method that preserves local structure while often retaining more global organization than t-SNE.
UMAP can be useful for identifying spectral populations, gradients, and relationships within hyperspectral datasets.
Before t-SNE or UMAP is calculated, the hyperspectral spectra are reduced using principal component analysis (PCA).
The PCA Components setting determines how many principal components are passed to the nonlinear embedding algorithm.
Using PCA preprocessing reduces computational complexity and removes highly redundant spectral information.
A value of approximately 10 components is a useful general starting point for many hyperspectral datasets.
Creates a two-dimensional embedding using Dimension 1 and Dimension 2.
The 2D mode is required for manual Lasso selection.
Creates a three-dimensional embedding using Dimensions 1, 2, and 3.
The 3D display can be interactively rotated.
Changing a previously calculated 2D embedding to 3D requires recalculating the embedding.
Perplexity controls the approximate neighborhood size considered by t-SNE.
Smaller values emphasize very local relationships, while larger values consider broader neighborhoods.
A value of 30 is a useful general starting point.
Exaggeration influences the separation of groups during optimization.
The default value provided by the Explorer is generally appropriate as a starting point.
The number of neighbors controls the balance between local and broader structure in the UMAP embedding.
Smaller values emphasize local spectral populations. Larger values emphasize broader relationships within the dataset.
A value of approximately 15 is a useful starting point.
Minimum Distance controls how tightly UMAP allows neighboring observations to be packed in the embedding.
Smaller values can produce more compact groups, while larger values create more diffuse embeddings.
A value of 0.1 is a useful starting point.
Click Run to calculate the selected t-SNE or UMAP embedding.
Each valid hyperspectral image pixel is represented by one point in the embedding.
Calculation time depends on the number of image pixels, number of PCA components, selected algorithm, and computer performance.
t-SNE may require substantially more computation time for large hyperspectral images.
The Embedding panel displays the calculated low-dimensional representation of the hyperspectral pixels.
Points that appear close together have similar representations according to the selected dimensionality-reduction method.
The embedding coordinates are analytical coordinates and should not be interpreted as physical spatial coordinates.
K-medoids clustering groups pixels according to their positions in the calculated t-SNE or UMAP embedding space.
Clustering does not rerun t-SNE or UMAP.
Select the desired number of K-medoids classes.
The selected number determines how many groups are identified within the embedding.
Select the distance metric used for K-medoids clustering.
Squared Euclidean is a useful general starting point.
Runs K-medoids clustering on the current embedding coordinates.
The resulting classes are displayed both in embedding space and mapped back to their original spatial locations in the hyperspectral image.
The Cluster Image displays the spatial locations of the embedding-derived classes in the original image coordinates.
Each cluster is assigned a different color.
After clustering, use the class menu to inspect an individual cluster.
Selecting Class 1, Class 2, and so forth displays only the spatial pixels belonging to that class.
Selecting All Classes restores the complete cluster image.
The cluster view control changes how a calculated embedding is displayed.
For a 3D embedding, the clustered points can be displayed in either 2D or 3D.
A 2D embedding contains only two embedding coordinates and therefore cannot provide a true 3D cluster display.
Lasso provides manual selection directly in embedding space.
1. Calculate a t-SNE or UMAP embedding.
3. Display the embedding in 2D.
5. Draw a polygon around the desired population of embedding points.
6. Complete the polygon to create the selection.
The corresponding pixels are mapped automatically back to their original spatial locations in the hyperspectral image.
Lasso operates in 2D embedding space.
Clear removes the current class or Lasso selection.
The calculated embedding and K-medoids clustering are preserved.
The complete cluster image and clustered embedding are restored.
Exports hyperspectral image data corresponding to the currently selected class or Lasso region.
Pixels outside the selected region are excluded from the exported selection.
Exports the complementary image region.
Pixels inside the current class or Lasso selection are excluded, while pixels outside the selection are retained.
Exports the current class or Lasso selection as a binary spatial mask.
Selected pixels have a value of 1 and unselected pixels have a value of 0.
Exports the calculated t-SNE or UMAP embedding coordinates.
Each row corresponds to a valid hyperspectral image pixel and each column corresponds to an embedding dimension.
t-SNE and UMAP are nonlinear exploratory methods.
The absolute positions, orientation, and scale of the embedding axes do not have direct physical or spectral meaning.
The primary information is contained in the relative organization and grouping of observations in embedding space.
Different parameter settings may produce different embeddings.
For reproducible quantitative analysis, record the embedding method and parameter settings used.
2. Select 2D or 3D embedding.
3. Choose the number of PCA preprocessing components.
4. Set the embedding parameters.
6. Inspect the embedding structure.
7. Select the number of clusters and distance metric.
8. Run K-medoids clustering.
9. Inspect the spatial cluster image.
10. Select individual classes as needed.
11. Use Lasso for manual population selection when appropriate.
12. Export the selected region, inverse region, mask, or embedding as needed.