Machine Learning Studio
Machine Learning Studio trains supervised pixel classifiers from hyperspectral spectra and a spatial label image.
Supervised Hyperspectral Machine Learning
IDCubePro 2026 - Supervised Hyperspectral Machine Learning
Overview
Machine Learning Studio trains supervised pixel classifiers from hyperspectral spectra and a spatial label image.
Each labeled pixel contributes its full spectral feature vector from myData.Images and its corresponding class value from the loaded label or mask image.
The trained model can then classify the complete hyperspectral cube, generate a prediction map, provide confidence information when available, create a prediction overlay, and be applied to another compatible cube.
Use supervised machine learning when you already have meaningful class labels or masks and want to learn a predictive relationship between hyperspectral spectra and those classes.
This tool is appropriate when class identity is expected to be encoded primarily in the spectral feature vector of each pixel.
Unlike the Deep Learning Studio, this ML Studio does not use spatial patches around each pixel. It is primarily a per-pixel spectral classifier.
When To Consider Deep Learning Instead
Consider Deep Learning Studio when local spatial morphology, texture, or neighborhood structure is expected to contribute strongly to class identity.
For smaller labeled datasets, simpler spectral classifiers can be easier to train, faster to evaluate, and easier to interpret.
The hyperspectral cube comes from myData.Images.
The cube must be three-dimensional with rows x columns x spectral bands.
Each pixel spectrum becomes one feature vector for supervised learning.
Labels can come from current IDCubePro workflow memory or from a saved file.
Current workflow sources include:
- segmentationWorkflowLabels
Use Current Labels retrieves labels or masks already present in IDCubePro memory.
Use this option when the labels were created with Label Creation, Mask Creation, Binary Threshold, or another compatible workflow.
MAT is the preferred format because it can preserve exact class values and metadata.
TIFF and PNG are also accepted as secondary formats.
Color image files may need to be converted to sequential class IDs because display colors are not necessarily the original scientific class values.
The label or mask image must match the spatial rows and columns of the hyperspectral cube.
If the cube was cropped, resized, rotated, registered, or otherwise transformed after label creation, verify that the labels still align pixel-for-pixel.
The training-data preparation treats class value 0 differently depending on the number of foreground classes.
When more than one foreground class exists, only positive class values are used for training.
When there is only one positive foreground class, background 0 is also included so that a two-class foreground-versus-background model can be trained.
IDCubePro balances the number of sampled pixels across classes before training.
The smallest available class determines the maximum balanced count that can be used from every class.
This prevents a very large class from dominating training simply because it contains more labeled pixels.
Training pixels / class is the maximum number of pixels that may be used from each class after balancing.
For example, if the field is 5000 but the smallest class contains only 1200 valid pixels, the tool will use at most 1200 pixels from every class.
Increasing this value does not help if the smallest class contains fewer usable pixels.
HOW MANY TRAINING PIXELS SHOULD I USE?
There is no universal ideal count.
More pixels can improve statistical stability, but many neighboring pixels from one compact ROI are not equivalent to many independent examples.
Representative diversity across multiple ROIs, samples, subjects, or specimens is usually more valuable than simply maximizing the pixel count.
The tool reports the original pixel counts by class before balanced sampling.
If the largest class is more than approximately 5 times the smallest class, a warning is produced.
If the imbalance exceeds approximately 10 times, the warning is stronger.
Balanced sampling corrects the numerical imbalance used for training, but it cannot create biological or spectral diversity that is absent from a small class.
The current implementation uses a 25% holdout test set.
Approximately 75% of the balanced data are used for training and 25% are reserved for evaluation.
The split is performed with MATLAB cvpartition using HoldOut = 0.25.
Important Validation Limitation
The holdout is pixel-based rather than specimen-based or explicitly spatially separated.
Neighboring hyperspectral pixels can be highly correlated.
Therefore, evaluation from a random pixel holdout may be optimistic when training and test pixels come from the same spatial regions.
For strong scientific validation, test the final model on independent images, specimens, subjects, batches, or acquisitions whenever possible.
Feature normalization is learned using the training subset only.
For each spectral feature, the training mean and standard deviation are calculated.
Training and test spectra are then z-score normalized using those training-derived values.
The same stored FeatureMean and FeatureStd are reused for full-image prediction and prediction on new datasets.
Why Training-only Normalization Matters
Calculating normalization from all data before evaluation would allow information from the held-out test set to influence preprocessing.
Using training-only statistics reduces this form of information leakage.
Classifier - Auto Recommended
Auto recommended currently selects Linear SVM implemented through ECOC with linear learners and one-versus-all coding.
At present, Auto recommended and the explicit Linear SVM option use the same underlying method.
Use Auto recommended as a strong default starting point for most hyperspectral classification problems.
Linear SVM is a strong baseline for high-dimensional hyperspectral data.
It searches for linear decision boundaries between classes in spectral-feature space.
Advantages include good performance in high-dimensional settings, relatively low overfitting risk compared with very flexible nonlinear models, and efficient prediction.
Choose Linear SVM first when you do not have a strong reason to prefer another classifier.
kNN classifies a spectrum according to nearby training spectra in feature space.
The current implementation uses 5 nearest neighbors.
kNN can work well when classes form compact local groups and decision boundaries are irregular.
It can become sensitive to irrelevant features, noise, very large training sets, and differences in local sampling density.
Because IDCubePro performs z-score normalization first, distance comparisons are less dominated by features with large numerical scale.
Decision Tree recursively separates the feature space using spectral-feature thresholds.
Trees are easy to conceptualize and can model nonlinear class boundaries.
However, a single decision tree can overfit noisy or high-dimensional hyperspectral data.
Use it mainly as a comparison model or when interpretable threshold-like decisions are of interest.
LDA stands for Linear Discriminant Analysis.
It models classes using linear discriminant boundaries under distributional assumptions about class covariance.
LDA can perform well when class distributions are approximately compatible with those assumptions and the number of training samples is adequate.
It may be less robust when spectral features are strongly collinear, distributions are very non-Gaussian, or covariance estimation is unstable.
Naive Bayes estimates class probabilities while making a simplifying conditional-independence assumption between features.
Hyperspectral bands are often strongly correlated, so that assumption may not be realistic.
Nevertheless, Naive Bayes can provide a useful simple baseline and can perform surprisingly well in some datasets.
WHICH CLASSIFIER SHOULD I USE?
Recommended comparison sequence:
1. Start with Auto recommended / Linear SVM.
2. Record Accuracy, Macro-F1, and confusion patterns.
3. Compare kNN if nonlinear local structure may be important.
4. Compare LDA when the class distributions appear approximately linear and well sampled.
5. Use Decision Tree or Naive Bayes as additional baselines when appropriate.
Do not choose the classifier based only on training speed or one aggregate metric.
Prefer the method that performs consistently across scientifically important classes and remains stable on independent data.
Train Model performs balanced sampling, the 75/25 holdout split, training-only z-score normalization, classifier fitting, prediction on the held-out test subset, and metric calculation.
At least two classes are required.
Accuracy is the fraction of held-out test samples classified correctly.
It is easy to interpret but can hide poor performance in individual classes.
Balanced training reduces some class-imbalance effects, but accuracy should still be interpreted together with Macro-F1 and the confusion matrix.
Macro-F1 calculates an F1 score for each class and then averages across classes so that each class contributes equally.
This makes Macro-F1 particularly useful when all classes are scientifically important, including smaller classes.
A high overall accuracy with a substantially lower Macro-F1 can indicate that some classes are performing poorly even when the total number of correct predictions looks good.
How To Interpret Accuracy And Macro-f1 Together
If both Accuracy and Macro-F1 are high, performance is more likely to be broadly distributed across classes.
If Accuracy is high but Macro-F1 is much lower, inspect the confusion matrix for weak minority or difficult classes.
If both are low, the model may be underpowered, labels may be noisy, classes may overlap spectrally, or preprocessing may be insufficient.
The software provides a qualitative model-quality grade based on Accuracy and Macro-F1.
Use the grade as a convenient summary only.
It should not replace inspection of the underlying metrics, confusion matrix, class representation, and independent validation.
Rows represent true classes and columns represent predicted classes.
Values on the main diagonal are correct classifications.
Off-diagonal values are misclassifications.
Large off-diagonal values identify specific class pairs that the model has difficulty separating.
How To Read The Confusion Matrix
If Class 2 is frequently predicted as Class 3, inspect the spectra and annotations for those two classes.
Possible causes include genuine spectral similarity, mislabeled pixels, mixed boundary pixels, insufficient examples, inconsistent preprocessing, or class definitions that overlap scientifically.
After Predict Full Image has been run, clicking a confusion-matrix cell highlights pixels whose true and predicted labels correspond to that cell.
This can help identify where specific errors occur spatially.
Predict Full Image classifies every pixel in the currently loaded hyperspectral cube using the trained classifier.
The stored training-set normalization is applied before prediction.
The result is stored as a predicted class image.
Full-image Prediction Is Not Validation
A visually plausible full-image prediction is useful, but it is not an independent measure of model accuracy.
Prediction maps should be interpreted together with held-out metrics and preferably independent validation data.
When the selected MATLAB classifier returns prediction scores, IDCubePro stores the maximum score at each pixel as the confidence image.
The exact numerical meaning of the score depends on the classifier implementation.
Therefore, confidence should be interpreted primarily as a relative indicator of how strongly the model favors its selected class, not as a universal calibrated probability.
How To Use The Confidence Map
Low-confidence regions may indicate mixed pixels, class boundaries, unusual spectra, noise, artifacts, or data that differ from the training distribution.
High confidence does not guarantee correctness.
Inspect confidence together with the predicted class map, known sample structure, and independent ground truth.
Overlay Prediction blends the predicted class colors with the original display image.
This is useful for checking spatial plausibility and locating obvious artifacts.
The current overlay uses a fixed blend with approximately 45% prediction color.
What A Good Prediction Map Looks Like
A useful map should generally be consistent with expected sample anatomy, materials, morphology, or known spatial organization.
Uniform regions should not show unexplained salt-and-pepper class switching unless the underlying sample is genuinely heterogeneous.
Class boundaries should be scientifically plausible.
Warning Signs In The Prediction
Potential warning signs include:
- One class dominating almost the entire image unexpectedly.
- Salt-and-pepper predictions in visually uniform regions.
- Predictions concentrated in glare, shadow, saturation, or detector artifacts.
- Strong disagreement with known anatomy or material layout.
- Low confidence throughout large portions of the image.
These patterns may reflect poor labels, domain mismatch, spectral overlap, preprocessing problems, or an unsuitable classifier.
Apply Model to New Data loads another MAT-file hyperspectral cube and applies the trained classifier.
The new cube may have different spatial rows and columns.
However, it must contain the same number of spectral features as the training cube.
Matching the number of bands is necessary but not sufficient.
The spectral bands should also correspond to the same wavelengths and ordering used during training.
If the band order changes, the model interprets the wrong physical wavelengths as its learned features.
When wavelength vectors are available in both datasets and have matching length, IDCubePro calculates the maximum wavelength difference.
A warning is shown if the maximum difference exceeds approximately 5 nm.
This warning is useful but does not guarantee full acquisition compatibility.
A model trained on one dataset may perform worse on data collected with another instrument, illumination condition, acquisition day, sample preparation, subject population, or preprocessing pipeline.
This phenomenon is often called domain shift.
Always validate transferred models on representative independent data before relying on them scientifically.
Export Model saves the trained classifier and the information needed to reproduce its prediction preprocessing.
The exported structure includes:
- PredictedImage when available
- ConfidenceImage when available
Why Featuremean And Featurestd Are Saved
The model expects new spectra to be normalized exactly as the training data were normalized.
Recomputing mean and standard deviation independently on a new image would change the feature space and can invalidate the trained decision boundaries.
Clear Results removes the trained model, metrics, predictions, confidence image, normalization parameters, classifier choice, and class-balance report from the current ML Studio session.
Loaded labels remain available unless changed separately.
Recommended Baseline Workflow
1. Create scientifically defensible labels.
2. Use multiple representative ROIs per class.
3. Load the labels into ML Studio.
4. Start with Auto recommended / Linear SVM.
5. Leave Training pixels / class at 5000 unless the dataset is very small or computational constraints require less.
7. Review class-balance information.
8. Review Accuracy and Macro-F1.
9. Inspect the confusion matrix for specific class errors.
10. Compare another classifier only when there is a reason to do so.
12. Inspect prediction and confidence maps.
13. Review the overlay for spatial plausibility.
14. Validate on independent data whenever possible.
15. Export the final model only after evaluation is satisfactory.
1. Recheck label quality before changing the classifier.
2. Add more representative ROIs to weak or undersampled classes.
3. Remove ambiguous boundary pixels or mislabeled regions.
4. Inspect whether the confused classes are spectrally separable.
5. Confirm consistent reference correction and preprocessing.
6. Compare Linear SVM with kNN or LDA.
7. Consider feature selection or dimensionality reduction if many bands are noisy or redundant.
8. Use independent validation to determine whether apparent improvements generalize.
This ML Studio currently uses the full spectral feature vector supplied by the hyperspectral cube.
If many bands are noisy, redundant, or irrelevant, external feature-selection tools in IDCubePro may improve efficiency or robustness before supervised training.
Any feature selection used for model development must be applied consistently to future prediction data.
The classifier can learn differences caused by sample chemistry, biology, illumination, instrument response, background, or artifacts.
Therefore, consistent preprocessing is critical.
Reference correction, masking, denoising, band removal, and other preprocessing steps should be selected according to the scientific acquisition workflow and then applied consistently.
For a reproducible model, preserve:
- The hyperspectral acquisition conditions.
- Wavelength calibration and band order.
- Training / validation strategy.
- Independent validation results.
Common Mistake - Too Many Pixels From One Roi
A huge number of neighboring pixels can create the appearance of a large training dataset while containing little independent biological or material variation.
Use multiple independent regions and samples whenever possible.
Common Mistake - Optimizing On The Same Holdout Repeatedly
If many classifier choices or preprocessing variants are selected based on the same holdout results, the holdout gradually becomes part of model selection.
The final performance should therefore be confirmed on independent data not used for those choices.
Common Mistake - Treating Confidence As Certainty
Prediction scores are not automatically calibrated probabilities.
A high score can still be wrong, especially under domain shift.
Common Mistake - Applying A Model To Mismatched Spectra
The same band count does not guarantee the same feature meaning.
Verify wavelength order, spectral calibration, preprocessing, and acquisition compatibility.
Common Mistake - Using Only Overall Accuracy
Overall accuracy can conceal failure in a scientifically important class.
Always inspect Macro-F1 and the confusion matrix.
Choosing The Best Final Model
Prefer the model that provides:
- Strong and reasonably balanced class performance.
- High Macro-F1 as well as good overall Accuracy.
- Limited confusion between scientifically important classes.
- Stable performance across repeated or independent datasets.
- Spatially plausible prediction maps.
- Sensible confidence behavior.
- Compatibility with the intended future acquisition workflow.
The classifier with the highest single holdout score is not automatically the best scientific model.
Important Scientific Limitation
Supervised machine learning identifies statistical relationships between spectra and supplied class labels.
It does not by itself establish causation, molecular identity, biological mechanism, or clinical validity.
Interpret predictions within the experimental context and validate them against appropriate independent ground truth.
Supervised Hyperspectral Machine Learning Close Machine Learning Studio help.