Documentation  ›  Documentation Library  ›  Label Creation
AI & Machine Learning

Label Creation

Manual Label Creation lets you create multiclass spatial ground-truth labels by drawing regions of interest (ROIs) on the current IDCubePro display image.

Idcubepro 2026 - Manual Label Creation

Overview

Manual Label Creation lets you create multiclass spatial ground-truth labels by drawing regions of interest (ROIs) on the current IDCubePro display image.

The resulting label image can be used for supervised machine learning, deep learning, segmentation validation, region-based analysis, and scientific annotation.

Each labeled pixel receives an integer class ID. Unlabeled pixels remain background.

For supervised learning, the model learns from the labels you provide.

Incorrect, inconsistent, ambiguous, or poorly placed labels can directly reduce model accuracy even when the machine-learning algorithm itself is well designed.

High-quality labels should represent the biological, chemical, material, or morphological classes that you actually want the model to distinguish.

Background is stored as class value 0.

User-defined classes are stored as integer values 1, 2, 3, and so forth.

For example, a three-class problem may use:

The numeric class IDs do not themselves carry biological meaning. Their meaning should be documented separately in your experimental workflow.

The # classes control defines how many class IDs are available for annotation.

Choose the number of scientifically meaningful categories you intend to distinguish.

Do not create extra classes merely because an image contains visual variation. A class should normally correspond to a meaningful analysis target.

The Class dropdown selects the integer class value assigned to newly drawn ROIs.

Always verify the selected class before drawing.

Accidentally labeling a region with the wrong class introduces training-label error.

Four ROI shapes are supported: Rectangle, Ellipse, Polygon, and Freehand.

Rectangle is fast and useful for compact regions with approximately rectangular boundaries.

It is often appropriate for homogeneous calibration regions, background areas, or simple sample geometry.

It is less suitable for irregular biological structures because it may include unwanted surrounding pixels.

Ellipse is useful for approximately round or oval structures.

It can provide cleaner annotation than a rectangle when the target has smooth curved boundaries.

Polygon is useful for irregular regions with reasonably well-defined boundaries.

It provides more geometric control than Rectangle or Ellipse while remaining easier to reproduce than detailed Freehand drawing.

Polygon is often a good general-purpose choice for tissue regions and irregular objects.

Freehand provides the greatest flexibility for complex boundaries.

Use it when the target has an irregular shape that cannot be represented adequately by simpler ROIs.

Because freehand boundaries depend strongly on the operator, use consistent annotation criteria across samples.

WHICH ROI TYPE SHOULD I USE?

Prefer the simplest ROI shape that captures the intended region without including substantial unwanted tissue or background.

For machine-learning training, a smaller clean ROI is usually preferable to a larger ROI contaminated by pixels from another class.

You may draw multiple separate ROIs for the same class.

This is useful because a class may appear in several disconnected locations or may have natural variability across the image.

Sampling several representative regions is generally better than labeling only one visually convenient area.

Click Draw ROI(s) to begin continuous drawing for the selected class.

After one ROI is completed, the tool allows another ROI for the same class.

Continue until you have sampled the desired regions, then click Stop Drawing.

Stop Drawing sets the continuous-drawing state to stop.

Complete the current ROI and then stop before changing classes or moving to another annotation task.

If ROIs from different classes overlap, the most recently drawn ROI determines the class value in the overlapping pixels.

In other words, newer labels overwrite older labels in the overlap.

Use this behavior intentionally. Unnoticed overlaps can silently change previously assigned class pixels.

ROIs are drawn on the current IDCubePro display image.

The display may be an RGB representation, a selected band, or another current image representation depending on the active IDCubePro view.

The hyperspectral cube remains the underlying scientific dataset; the display image provides the spatial reference used for annotation.

The Label Image panel shows the current integer label map.

Different class IDs are displayed with different colors for visual inspection.

The colors are visualization aids only. The scientific label information is the integer class value stored at each pixel.

Choosing A Display For Annotation

Use a display image in which the structures you intend to label can be identified reliably.

If a class boundary is unclear in the current RGB or band image, consider selecting a more informative wavelength, spectral composite, processed image, or other validated visualization before labeling.

Do not label a boundary simply because you expect it to be present if it cannot be supported by the image or independent ground truth.

Labels used for supervised learning should ideally reflect reliable ground truth.

Depending on the application, ground truth may come from morphology, histology, pathology, known sample composition, independent measurements, expert annotation, or another validated reference.

Visual appearance alone may be insufficient when different classes look similar in the displayed image.

Try to label representative examples of each class rather than only the easiest or most obvious examples.

A model trained only on ideal examples may perform poorly on realistic variation.

When appropriate, include variation in brightness, morphology, spectral appearance, position, and sample condition.

Large differences in the number of labeled pixels between classes can influence model training.

You do not always need exactly equal pixel counts, but avoid unintentionally providing enormous regions for one class and only a few pixels for another.

Inspect the amount and diversity of annotation for every class before training.

Avoiding Boundary Contamination

Pixels near boundaries may contain mixed signal from adjacent structures because of optical resolution, scattering, registration error, or partial-volume effects.

If the scientific class boundary is uncertain, consider labeling a conservative interior region rather than forcing labels directly along an ambiguous edge.

Clean interior training pixels can be more useful than a larger but contaminated training region.

Neighboring hyperspectral pixels are often highly correlated.

Thousands of labeled pixels from one small ROI do not necessarily provide the same diversity as pixels sampled from multiple independent regions or specimens.

For robust model development, distribute labels spatially and across independent samples whenever the study design allows.

For predictive modeling, avoid allowing nearly identical neighboring pixels from the same annotated region to appear independently in both training and validation sets unless the validation method explicitly accounts for spatial correlation.

Spatially separated or specimen-level validation is generally more informative when assessing generalization.

In this tool, value 0 represents unlabeled/background pixels.

Do not assume that every pixel with value 0 is automatically a scientifically defined background training class.

Whether value 0 is ignored or treated as a class depends on the downstream analysis or training workflow.

Clear Class removes label pixels and ROI objects associated with the currently selected class.

Use it when one class needs to be redrawn without deleting annotations from the other classes.

Clear All removes every ROI and resets the complete label image to 0.

Use this only when you intend to restart the annotation.

The current label image is stored in IDCubePro application data as:

  • segmentationWorkflowLabels

These entries allow downstream IDCubePro segmentation and machine-learning tools to access the labels.

Save exports the label dataset in three forms: MAT, PNG, and TIFF.

The three formats serve different purposes.

Mat File - Recommended Scientific Format

The MAT file contains a LabelData structure with the label image and associated metadata.

The current implementation stores:

Use the MAT file when preserving the scientific label data and metadata is important.

The TIFF stores the label image as uint16 integer values.

This preserves the numerical class IDs and is useful for interoperability with image-analysis software.

For multiclass scientific labels, the integer TIFF is more informative than a color preview.

The PNG is a color visualization of the label image.

It is useful for inspection, figures, presentations, and visual documentation.

Do not use the RGB colors in the PNG as a substitute for the original integer class IDs when exact scientific labels are required.

Recommended Annotation Workflow

1. Decide what each class means before drawing.

2. Set the total number of classes.

3. Select an informative display image.

4. Select Class 1 and an appropriate ROI type.

6. Draw several clean representative regions for Class 1.

8. Select the next class and repeat.

9. Inspect the Label Image for incorrect labels and overlaps.

10. Check that each class has adequate and representative coverage.

11. Clear and redraw questionable regions if necessary.

12. Save the completed labels.

13. Preserve the class definitions with the experiment or model documentation.

Labels define the target output that the supervised model will attempt to learn.

Good performance requires both useful hyperspectral features and reliable labels.

If the classes cannot be distinguished spectrally, increasing the number of training pixels alone will not solve the problem.

Patch-based deep-learning models may use spatial neighborhoods around labeled pixels.

Labels placed very close to class boundaries can therefore create patches containing mixed classes.

When appropriate, use clean regions and allow the downstream training workflow to manage spatial separation and guard bands.

Do not judge a model only by how well it reproduces the pixels used for training.

Validation should test performance on data not used to fit the model.

For hyperspectral images, spatially separated validation is often preferable to a purely random pixel split because nearby pixels may be highly similar.

There is no single correct number of ROIs or pixels.

The required amount depends on class complexity, spectral variability, image quality, model type, number of classes, and the independence of the samples.

Prioritize representative diversity and label correctness rather than simply maximizing pixel count.

Common Mistake - Labeling Only The Most Obvious Regions

This can create an artificially easy training set.

Include realistic within-class variation if the model is expected to recognize that variation later.

Common Mistake - Large Impure Rois

A very large ROI may cross tissue boundaries or include several spectral populations.

Smaller high-confidence ROIs are often safer for supervised training.

Common Mistake - Too Few Independent Regions

Many adjacent pixels from one ROI can create a large pixel count without providing much independent information.

Whenever possible, annotate multiple spatial regions and multiple independent specimens.

Common Mistake - Class Definitions That Overlap

If two classes are not conceptually distinct, different annotators or different samples may receive inconsistent labels.

Define classes using explicit scientific criteria before building a large training dataset.

Common Mistake - Changing The Image Geometry

Label dimensions must remain spatially compatible with the hyperspectral cube used for training.

If the cube is cropped, resized, rotated, registered, or otherwise geometrically transformed after labeling, the labels must undergo the corresponding transformation or be regenerated.

Quality Control Before Saving

Before saving, inspect the complete Label Image and ask:

  • Are the intended classes represented?
  • Are any ROIs assigned to the wrong class?
  • Are there accidental overlaps?
  • Are boundaries reasonable?
  • Are there enough representative regions?
  • Are any classes dominated by artifacts?
  • Does the label image still align with the hyperspectral dataset?

For scientific studies, document what each class represents, who or what defined the ground truth, which image representation was used for annotation, and any exclusion criteria.

Consistent annotation rules are especially important when several users label data or when labels are created over a long study period.

Important

Manual labels are not merely graphical annotations; they may become the ground truth used to train and evaluate predictive models.

Label conservatively when class identity is uncertain.

Preserve exact integer class IDs and metadata whenever possible.

Always verify spatial compatibility between the label image and the hyperspectral cube before model training.