Documentation  ›  Documentation Library  ›  K-means Clustering
AI & Machine Learning

K-means Clustering

K-means clustering groups hyperspectral data into a user-selected number of clusters based on similarity in the input data.

Overview

K-means clustering groups hyperspectral data into a user-selected number of clusters based on similarity in the input data.

IDCubePro provides two K-means modes: Global 2D and 3D Volume.

Global 2D uses spectral bands as feature channels and produces one global 2D cluster-label map.

Each pixel is assigned to one cluster according to its spectral feature vector across the hyperspectral bands.

The true result is a 2D label map. IDCubePro also replicates this map across spectral bands so it can be displayed through the standard hyperspectral viewer.

3D Volume applies K-means directly to the full three-dimensional data volume and produces a 3D cluster-label volume.

Number of clusters specifies K, the number of groups K-means will attempt to identify.

K must be 2 or greater. Too few clusters can merge distinct populations, while too many clusters can divide similar populations into smaller groups.

Max iterations specifies the maximum number of K-means optimization iterations.

Increase this value if clustering has not converged adequately. More iterations can require additional computation.

Cluster labels are categorical values 1, 2, 3, and so forth.

They identify cluster membership and are not physical spectral intensities.

The numerical label assigned to a cluster does not imply ranking, magnitude, concentration, or biological meaning.

The original source hyperspectral cube is preserved in kmeansSourceCube.

Cluster labels are stored in kmeansLabels and segmentationWorkflowLabels.

The active myData.Images dataset is replaced with a display-compatible label result so that the clustering can be visualized in IDCubePro.

In Global 2D mode, the true result is one 2D cluster-label map.

For compatibility with the hyperspectral viewer, that same 2D map is replicated across the spectral dimension.

The replicated bands should not be interpreted as independent spectral measurements.

In 3D Volume mode, the active result is the full 3D cluster-label volume.

Each voxel is assigned to one of the requested K clusters.

K-means is an unsupervised clustering method.

It groups data according to similarity but does not automatically identify the physical, chemical, biological, or material meaning of each cluster.

Cluster interpretation should be supported by the original spectra, spatial context, known references, controls, or other independent information.

1. Load the hyperspectral dataset.

3. Choose Global 2D or 3D Volume.

4. Select the desired number of clusters.

5. Set the maximum number of iterations.

7. Inspect the resulting cluster-label display.

8. Compare the clusters with the original image and spectra.

9. Adjust K and repeat if the segmentation is too coarse or too fragmented.

Important

K-means results depend on the scaling, preprocessing, spectral quality, number of clusters, and structure of the input dataset.

Different values of K can produce substantially different segmentations.

The cluster numbers themselves are arbitrary categorical labels and may not correspond between separate K-means runs.

Close K-means clustering help.