IDCubeCloud Documentation

Clustering

Group spectrally similar pixels using unsupervised K-Means or DBSCAN workflows.

Technical Reference Revised September 2026

1. Overview

Clustering groups pixels with similar spectral signatures into classes - without requiring any labeled training data. It is the primary unsupervised classification tool in IDCubeCloud and is ideal for land-cover mapping, anomaly detection, and exploratory spectral analysis.

Key points

  • No training data required - the algorithm discovers natural groupings in the spectral data.
  • Two algorithms: K-Means (compact clusters) and DBSCAN (density-based, finds irregular shapes).
  • Results are displayed as a color-coded classified map in the viewer and as a 3D PCA scatter plot.
  • Works on any loaded scene - satellite or lab hyperspectral.

2. What Clustering Does

1. Loaded hyperspectral cube [H x W x B]

2. Each pixel to a point in B-dimensional spectral feature space (Optional) PCA dimensionality reduction applied internally to reduce noise before clustering

3. Clustering algorithm groups nearby points into classes

4. K-Means: partitions into exactly K clusters by minimising within-cluster sum of squares

5. DBSCAN: finds dense regions separated by low-density space (number of clusters is automatic - determined by the data)

6. Cluster label assigned to every pixel

7. Output

8. ① Classified map [H x W] - colour-coded by cluster ID

9. ② 3D PCA scatter plot - each point coloured by cluster

10. ③ Mean spectrum per cluster for identification

3. Where to Find It

1. Load any dataset from My Files or the Satellite page.

2. Click Clustering in the feature toolbar.

3. The Clustering Parameters panel appears in the left sidebar; the classified map and 3D scatter plot appear in the main area after running.

4. Step-by-Step Workflow

Step 1 Load Data

Any loaded scene works. For best results on large or high-dimensional data, consider running PCA first and using the PCA-reduced representation.

Step 2 Open Clustering

Click Clustering in the feature toolbar.

Step 3 Select Algorithm

Choose K-Means or DBSCAN depending on the expected cluster structure (see Section 5).

Step 4 Set Parameters

  • For K-Means: set Number of Clusters (default: 5).
  • Adjust 3D Scatter Points to control how many sample points appear in the PCA scatter plot (default: 5,000).

Step 5 Run Clustering

Click Run Clustering. Processing time depends on scene size and number of bands.

Step 6 Inspect Results

  • Classified map - colour-coded raster shown in the viewer.
  • 3D PCA scatter - review cluster separation in spectral space.
  • Click individual clusters in the map to inspect their mean spectrum.

Step 7 Interpret and Label

Cross-reference cluster spectra against the Spectral Library to assign material labels (e.g., vegetation, water, bare soil, urban).

5. Clustering Algorithms

K-Means

Property
Detail
Type
Centroid-based partitioning
Parameters
Number of clusters K
Output clusters
Always exactly K clusters
Best for
Roughly spherical, similarly sized groups; general land-cover mapping
Limitation
Must specify K in advance; sensitive to initial centroid placement

1. Initialise K random centroids

2. Assign each pixel to nearest centroid (Euclidean distance in feature space)

3. Recalculate centroids as mean of assigned pixels

4. Repeat until centroids stabilise (convergence)

DBSCAN

Property
Detail
Type
Density-based spatial clustering
Parameters
Epsilon (neighbourhood radius), min_samples
Output clusters
Variable - number determined automatically
Best for
Irregular cluster shapes, detecting anomalous/outlier pixels
Limitation
Sensitive to epsilon choice; poor on high-dimensional data without PCA pre-processing

6. Parameters Explained

Parameter
Description
Notes
Algorithm
K-Means or DBSCAN
K-Means recommended for general use
Number of Clusters (K)
How many classes to find (K-Means only)
Start with 5; adjust after inspecting results
3D Scatter Points
Number of pixels sampled for the 3D PCA scatter plot
Lower value = faster rendering; higher value = more representative plot

7. Output and Interpretation

Output
What It Shows
Classified map
Each pixel coloured by its cluster ID
3D PCA scatter
Pixels plotted in PC1/PC2/PC3 space, coloured by cluster
Cluster mean spectra
Average spectrum of all pixels in each cluster (click cluster in map)

Reading the classified map

  • Spatially coherent regions with the same colour = pixels with similar spectral signatures = likely same material.
  • Scattered, speckled colours = high spectral variability or too many clusters (reduce K).
  • One cluster dominates the map = data is spectrally uniform (try increasing K or applying PCA first).

8. Choosing the Right Number of Clusters

Guidance
Detail
Start with K = 5
A reasonable default for most satellite scenes (water, urban, vegetation, bare soil, mixed)
Increase K if clusters look spectrally mixed
Check mean spectra - if two spectra look similar, merge them
Decrease K if the map looks noisy
Too many clusters creates fragmented maps with no clear spatial pattern
Use PCA first
For data with > 20 bands, run PCA and cluster on the first 5-10 components
Elbow method
Run clustering at K = 3, 5, 7, 10 and compare within-cluster variance (available in export)

9. Known Limitations

Limitation
Detail
K-Means initialisation
Results can vary between runs due to random centroid initialisation. Run multiple times to verify stability.
High-dimensional data
K-Means degrades with many bands due to the "curse of dimensionality". Apply PCA first.
DBSCAN on hyperspectral data
Works best after PCA reduction to 3-5 components. Raw hyperspectral data is too high-dimensional.
Cluster labels have no inherent meaning
Cluster 1 is not always "vegetation" - interpret by inspecting mean spectra and comparing to spectral library.
Large scenes
Scenes > 2000 x 2000 pixels may take significant processing time. Crop if needed.

10. Common Use Cases

Use Case
Workflow
Land-cover mapping
Clustering (K-Means, K = 5-8) to classify map to label clusters via Spectral Library
Anomaly detection
Clustering to identify small, isolated clusters to inspect spectra
Pre-classification exploration
PCA (3-5 components) to Clustering to review 3D scatter for natural groupings
Vegetation health segmentation
Clustering to compare cluster spectra to NDVI range
Mineral / geology mapping
Clustering (K = 8-12) to match cluster spectra to spectral library

Part of the IDCubeCloud User Documentation - see also: PCA Guide, Spectral Library Guide, Vegetation Indices Guide.

Manual revision: September 2026

↑ Back to top