Documentation  ›  Documentation Library  ›  Deep Learning Studio
AI & Machine Learning

Deep Learning Studio

Deep Learning Studio trains and applies 3D convolutional neural networks (CNNs) for supervised hyperspectral image classification.

Overview

Deep Learning Studio trains and applies 3D convolutional neural networks (CNNs) for supervised hyperspectral image classification.

The network learns from labeled pixels and uses both spectral information and local spatial context to predict a class for every pixel in the image.

The current workflow uses a 75% training / 25% validation split, balanced class sampling, PCA learned from training pixels only, and normalization learned from training pixels only.

The exact PCA and normalization learned during training are stored with the model and reused during prediction.

Use Deep Learning Studio when you have labeled examples and expect the classes to differ not only spectrally but also in their local spatial structure.

Typical applications include tissue classification, lesion segmentation, material classification, vegetation mapping, microscopy, and other supervised hyperspectral segmentation problems.

Deep learning is especially useful when spectral and spatial context together provide better discrimination than either one alone.

If only a very small number of labeled examples are available, simpler machine-learning methods may be easier to train and interpret.

Recommended Starting Workflow

1. Confirm that the correct hyperspectral image is loaded.

2. Load reliable class labels.

5. Use Epochs = 5-10 for an initial run.

9. Inspect validation accuracy, training behavior, and the confusion matrix.

10. Adjust settings only if the baseline model is inadequate.

12. Inspect the prediction overlay and confidence map.

13. Export the model only after the results are satisfactory.

The studio automatically uses the current IDCubePro hyperspectral cube when it opens.

Load HSI can be used to replace the current studio dataset with another MAT-file.

For model transfer, the new hyperspectral image should have compatible wavelength sampling, calibration, preprocessing, and acquisition conditions.

Supervised deep learning requires class labels for representative pixels.

Labels can be loaded from IDCubePro annotation memory or from a file.

MAT is preferred because exact class IDs are preserved.

Label value 0 is treated as unlabeled or background.

Positive integer values represent classes.

The quality of the labels strongly controls model quality.

A more complex CNN cannot compensate for incorrect or ambiguous training labels.

Avoid uncertain boundaries, mixed pixels, obvious artifacts, and mislabeled regions whenever possible.

Use multiple spatially separated examples of each class instead of labeling only one small compact region.

Very small classes provide limited training information even though the studio balances class sampling.

Approximately 75% of labeled samples are used for training and 25% are reserved for validation.

When possible, IDCubePro uses a spatial holdout so that training and validation pixels come from different image regions.

A guard band is placed between them based on the spatial patch radius.

This reduces the chance that nearly identical overlapping patches appear in both training and validation.

If a spatial split cannot include every class in both partitions, the software falls back to a class-wise random 75/25 split.

Neighboring hyperspectral pixels are often highly correlated.

If nearby pixels are randomly divided between training and validation, validation accuracy can become artificially optimistic.

Spatial holdout is therefore preferred when the label geometry allows it.

PCA reduces the original hyperspectral bands to a smaller number of principal components before CNN training.

This reduces spectral redundancy, memory use, and computation.

PCA is fitted using TRAINING pixels only.

Validation and prediction data are transformed using the stored training-derived PCA model.

Choosing The Number Of Pca Components

PCA = 5 is a good starting point for many hyperspectral datasets.

Use fewer components when the bands are highly redundant or when the training set is small.

Use more components when important class differences are subtle and may require additional spectral detail.

Too few components can remove discriminative information.

Too many components can retain noise and redundant information and increase computation.

A practical comparison is to test 5 components first, then compare with approximately 8-15 components if needed.

Window defines the spatial patch width and height centered on each labeled pixel.

For example, Window = 11 means that the network analyzes an 11 x 11 neighborhood around the center pixel.

The window must be odd so that there is a well-defined center pixel.

Window = 7 or 9 emphasizes very local structure.

Window = 11 is a good general starting point.

Window = 15 or larger can be useful when class identity depends strongly on larger-scale morphology or texture.

If the window is too large, it may contain multiple classes and reduce the specificity of the center-pixel label.

If the prediction is very noisy or fragmented, a moderately larger window may improve spatial consistency.

If small structures disappear or boundaries become overly smooth, reduce the window.

An epoch is one complete pass through the training dataset.

More epochs allow additional optimization.

Too few epochs can produce underfitting.

Too many epochs can produce overfitting.

The default of 5 epochs is useful for rapid testing.

For an initial analysis, 5-10 epochs are reasonable.

Increase epochs only if training has not converged and validation performance is still improving.

Learning Rate controls the size of optimization steps used by the Adam optimizer.

LR = 0.001 is a strong default starting value.

If loss oscillates strongly or training appears unstable, try a lower value such as 0.0005 or 0.0001.

If training improves extremely slowly, a somewhat larger value may help, but large changes can destabilize optimization.

Simple CNN is the recommended starting architecture.

It has the lowest memory requirement, trains fastest, and has the lowest risk of overfitting on small datasets.

Use Simple CNN first to establish whether the classification problem is learnable.

Deeper CNN has greater capacity and can learn more complex spectral-spatial features.

Use it when Simple CNN appears to underfit and enough labeled data are available.

It requires more computation and has a greater risk of overfitting.

Hybrid CNN uses several 3D convolutional kernels with broader spectral extent and greater model capacity.

It is intended for problems in which subtle spectral relationships and spatial morphology are both important.

Use it after establishing a baseline with Simple CNN.

Hybrid CNN generally benefits from a larger and more diverse labeled training set.

2. Verify that labels and preprocessing are correct.

3. If validation performance is inadequate, try Deeper CNN.

4. Try Hybrid CNN when spectral-spatial interactions appear especially important.

Do not automatically select the most complex model.

A simpler model with stable validation performance is often preferable to a more complex model that overfits.

After PCA, component means and standard deviations are calculated from TRAINING pixels only.

Those stored values are reused for validation and prediction.

This prevents validation or prediction data from influencing preprocessing learned during training.

Click Train after loading valid labels and selecting the desired settings.

The studio creates the training and validation partitions, fits PCA, normalizes the components, extracts patches, builds the selected CNN, trains the network, and evaluates the validation data.

The Training quality plot shows loss during optimization.

Loss measures disagreement between predicted and true labels.

Lower training loss is generally better, but training loss alone should never be used to judge model quality.

Training loss should generally decrease while validation behavior remains stable or improves.

The goal is not simply to minimize training loss, but to obtain a model that performs well on held-out data.

If both training and validation performance remain poor, the model may be underfitting.

  • Increase epochs if training has not converged.
  • Verify that the classes are actually separable.

If training performance becomes excellent but validation performance remains much worse, the model may be overfitting.

  • Collect more labeled examples.
  • Increase spatial diversity of training labels.
  • Reduce unnecessary model complexity.

Validation accuracy is the fraction of held-out validation samples that are classified correctly.

Higher accuracy is better, but overall accuracy can hide poor performance in an important minority class.

Always inspect the confusion matrix in addition to overall validation accuracy.

The confusion matrix shows which true classes are predicted as which classes.

Values concentrated along the main diagonal indicate correct classification.

Large off-diagonal values identify specific class pairs that the model confuses.

How To Use The Confusion Matrix

If two classes are repeatedly confused, inspect their spectra, class definitions, spatial appearance, labeling quality, and number of training examples.

  • Spectra are genuinely similar.
  • Classes overlap biologically or chemically.
  • Labels include boundary or mixed pixels.
  • Too few examples are available.
  • PCA removed useful information.
  • Window size is inappropriate.

Predict applies the trained or imported model to the complete current hyperspectral image.

The exact PCA and normalization learned during training are reused automatically.

Do not independently recalculate PCA or normalization when applying a stored model.

Prediction Mode - Safe Batch

Safe batch is the recommended default.

It processes a bounded number of patches at a time and is generally more memory efficient.

Use Safe batch when GPU or CPU memory is limited or when reliability is more important than maximum speed.

Prediction Mode - Fast Block

Fast block predicts spatial blocks for higher throughput.

Use Fast block when speed is important and sufficient memory is available.

If memory errors occur, switch to Safe batch.

Batch controls how many prediction patches are processed together.

Larger batches can improve throughput but use more memory.

The default value of 3000 is a practical starting point.

Reduce Batch if GPU or system-memory errors occur.

Batch size mainly affects computational efficiency; it does not inherently improve model accuracy.

Overlay controls the opacity of the predicted-class map over the base hyperspectral image.

Lower values show more of the original image.

Higher values emphasize the predicted classes.

Overlay affects visualization only and does not change the prediction.

When Confidence is On, the studio displays the highest class score associated with each predicted pixel when scores are available.

High confidence means that one class receives a substantially stronger model score than the alternatives.

Low confidence commonly occurs at class boundaries, mixed pixels, noisy regions, unusual samples, or regions that differ from the training data.

A high-confidence prediction is not automatically correct.

Interpret confidence together with validation performance and spatial plausibility.

How To Evaluate The Prediction Map

Inspect whether the predicted classes form spatial patterns that make sense for the application.

  • Salt-and-pepper predictions in uniform regions.
  • One class unexpectedly dominating the entire image.
  • Classes appearing in physically impossible locations.
  • Predictions concentrated in known noisy or saturated regions.

These patterns can indicate poor labels, inappropriate preprocessing, insufficient spectral information, an unsuitable Window size, or poor generalization.

Import Model loads a previously exported Deep Learning Studio model.

The imported file contains the trained network and its stored preprocessing.

The current HSI must contain the number of spectral bands expected by the model.

Wavelength compatibility also matters.

Two cubes with the same number of bands are not necessarily equivalent if wavelength calibration differs.

Export Model saves the trained network together with PCA, normalization, wavelength information, Window size, model architecture, and training settings.

Export the model only after reviewing validation accuracy, confusion behavior, and prediction quality.

A model is most reliable when new data were acquired with compatible instrumentation, wavelength calibration, preprocessing, and sample conditions.

Domain differences can reduce prediction accuracy even when the model performs very well on its original validation data.

Reset clears the trained or imported network, labels, stored preprocessing, validation results, and prediction displays.

The current hyperspectral dataset is then reloaded into the studio.

Recommended Starting Settings

Change one major setting at a time so that you can understand why model behavior changes.

2. Confirm that every class has enough representative examples.

3. Use the confusion matrix to identify specific problem classes.

4. Increase PCA components if useful spectral information may be missing.

5. Adjust Window size if more or less spatial context is needed.

6. Increase Epochs only if training has not converged.

7. Try Deeper CNN or Hybrid CNN only after evaluating Simple CNN.

8. Reinspect prediction maps and confidence for spatial plausibility.

How To Choose The Best Model

Do not choose a model based only on training accuracy or training loss.

Prefer a model that provides:

  • Strong validation accuracy.
  • Acceptable performance across all important classes.
  • Limited off-diagonal confusion.
  • Spatially plausible prediction maps.
  • Sensible confidence patterns.
  • Reproducible behavior on independent data when available.

Deep-learning predictions are data-driven classifications, not direct chemical or biological measurements.

Performance depends on label quality, acquisition conditions, preprocessing, spectral calibration, training-sample diversity, class definitions, and model complexity.

Validation performance from one image does not guarantee equivalent performance on different instruments, batches, subjects, tissues, or acquisition conditions.

For strong scientific conclusions, evaluate the final model on independent data that were not used for training or model selection.

Close Deep Learning Studio help.