Clustering
Group spectrally similar pixels using unsupervised K-Means or DBSCAN workflows.
1. Overview
Clustering groups pixels with similar spectral signatures into classes - without requiring any labeled training data. It is the primary unsupervised classification tool in IDCubeCloud and is ideal for land-cover mapping, anomaly detection, and exploratory spectral analysis.
Key points
- No training data required - the algorithm discovers natural groupings in the spectral data.
- Two algorithms: K-Means (compact clusters) and DBSCAN (density-based, finds irregular shapes).
- Results are displayed as a color-coded classified map in the viewer and as a 3D PCA scatter plot.
- Works on any loaded scene - satellite or lab hyperspectral.
2. What Clustering Does
1. Loaded hyperspectral cube [H x W x B]
2. Each pixel to a point in B-dimensional spectral feature space (Optional) PCA dimensionality reduction applied internally to reduce noise before clustering
3. Clustering algorithm groups nearby points into classes
4. K-Means: partitions into exactly K clusters by minimising within-cluster sum of squares
5. DBSCAN: finds dense regions separated by low-density space (number of clusters is automatic - determined by the data)
6. Cluster label assigned to every pixel
7. Output
8. ① Classified map [H x W] - colour-coded by cluster ID
9. ② 3D PCA scatter plot - each point coloured by cluster
10. ③ Mean spectrum per cluster for identification
3. Where to Find It
1. Load any dataset from My Files or the Satellite page.
2. Click Clustering in the feature toolbar.
3. The Clustering Parameters panel appears in the left sidebar; the classified map and 3D scatter plot appear in the main area after running.
4. Step-by-Step Workflow
Step 1 Load Data
Any loaded scene works. For best results on large or high-dimensional data, consider running PCA first and using the PCA-reduced representation.
Step 2 Open Clustering
Click Clustering in the feature toolbar.
Step 3 Select Algorithm
Choose K-Means or DBSCAN depending on the expected cluster structure (see Section 5).
Step 4 Set Parameters
- For K-Means: set Number of Clusters (default: 5).
- Adjust 3D Scatter Points to control how many sample points appear in the PCA scatter plot (default: 5,000).
Step 5 Run Clustering
Click Run Clustering. Processing time depends on scene size and number of bands.
Step 6 Inspect Results
- Classified map - colour-coded raster shown in the viewer.
- 3D PCA scatter - review cluster separation in spectral space.
- Click individual clusters in the map to inspect their mean spectrum.
Step 7 Interpret and Label
Cross-reference cluster spectra against the Spectral Library to assign material labels (e.g., vegetation, water, bare soil, urban).
5. Clustering Algorithms
K-Means
1. Initialise K random centroids
2. Assign each pixel to nearest centroid (Euclidean distance in feature space)
3. Recalculate centroids as mean of assigned pixels
4. Repeat until centroids stabilise (convergence)
DBSCAN
6. Parameters Explained
7. Output and Interpretation
Reading the classified map
- Spatially coherent regions with the same colour = pixels with similar spectral signatures = likely same material.
- Scattered, speckled colours = high spectral variability or too many clusters (reduce K).
- One cluster dominates the map = data is spectrally uniform (try increasing K or applying PCA first).
8. Choosing the Right Number of Clusters
9. Known Limitations
10. Common Use Cases
Part of the IDCubeCloud User Documentation - see also: PCA Guide, Spectral Library Guide, Vegetation Indices Guide.
Manual revision: September 2026
↑ Back to top