Volumetric AI Studio
Volumetric AI Studio trains and applies a 3D convolutional neural network to segment structures in true X-Y-Z image volumes.
Volumetric AI Studio — 3D Segmentation
Idcubepro 2026 — Volumetric Ai Studio
Overview
Volumetric AI Studio trains and applies a 3D convolutional neural network to segment structures in true X-Y-Z image volumes.
The validated model in the current release is a MATLAB-native 3D U-Net trained with trainnet.
The current release performs binary volumetric segmentation: background = 0 and the structure of interest = 1.
Examples of appropriate volumetric data include MRI, micro-CT, confocal or multiphoton Z-stacks, reconstructed microscopy volumes, and other true 3D image datasets.
Important — True Volume Vs Hyperspectral Cube
IDCubePro also stores hyperspectral data as 3D arrays, but a hyperspectral cube is usually X × Y × wavelength, not X × Y × Z.
Do not treat the spectral-band dimension of a conventional hyperspectral cube as spatial depth unless the dataset is genuinely a reconstructed volumetric image.
If the current dataset contains wavelength metadata matching dimension 3, Volumetric AI Studio asks you to confirm that the array should be interpreted as a true X-Y-Z volume.
Volume input — MAT, NIfTI (.nii/.nii.gz), or multi-page TIFF.
Mask input — MAT, NIfTI, multi-page TIFF, or a current IDCubePro mask.
The volume and mask must have identical X-Y-Z dimensions.
Use Current Dataset loads the current IDCubePro 3D array as the active volume.
Use Current Mask loads the mask generated by the IDCubePro Mask Creation workflow when available.
The preferred mask source is MaskData.MaskVolume because it preserves the full 3D segmentation together with slice-label metadata.
For binary segmentation, voxels with value 0 are background and voxels with values greater than 0 are treated as foreground.
If your mask contains several numeric classes, the current binary Volumetric AI workflow collapses all nonzero values into one foreground class.
The current mask workflow distinguishes labeled slices from slices that have not been reviewed.
Manual anchor slices and interpolated slices are marked as labeled.
Slices outside the annotated/interpolated region can remain unlabeled and are excluded from training rather than being silently treated as confirmed background.
This is important when the biological structure may be present on slices that have not yet been annotated.
Recommended Mask Creation Workflow
1. Open Mask Creation for the 3D volume.
2. Draw and Store ROIs on representative anchor slices.
3. Click Interpolate to label the slices between manual anchors.
5. Adjust any interpolated ROI that needs correction.
6. Click Store on a corrected slice to promote it to a manual anchor.
7. Continue until the labeled volume is acceptable.
Single-volume training is useful for software validation, troubleshooting, or a small proof-of-concept.
1. Click Use Current Dataset or Load Volume.
2. Click Use Current Mask or Load Mask.
3. Confirm that the volume and mask dimensions match.
4. Select training parameters.
6. Click Predict Full Volume.
7. Inspect the prediction, Dice/IoU metrics, and probability map.
8. Save the model or export the predicted mask if desired.
A model trained and validated from only one specimen should not be interpreted as evidence of generalization to independent biological specimens.
Recommended Multi-volume Workflow
For meaningful model development, use multiple independently acquired and labeled specimens or scans.
2. Add Current Pair or Add Pair From Files.
3. Add multiple independent volume + mask pairs.
4. Confirm that each volume has its matching mask.
7. Review the held-out validation result.
8. Apply the trained model to additional unseen volumes for independent testing.
Whole-volume Train/validation Split
With multiple datasets, Volumetric AI Studio splits by whole volume rather than mixing patches from the same specimen into both training and validation.
This reduces information leakage and gives a more realistic validation estimate.
For example, with 10 datasets and Train fraction = 0.8, the intended split is approximately 8 training volumes and 2 validation volumes.
Add Current Pair — adds the currently loaded volume and current mask as one training pair.
Add Pair From Files — selects a volume file and its matching mask file.
Remove Selected — removes the selected pair from the current training list.
Clear Dataset — removes all pairs from the current training list.
The dataset table reports the pair name, size, foreground voxel count, and source.
Patch size is the cubic X-Y-Z region used by the 3D U-Net during training.
A patch size of 32 means 32 × 32 × 32 voxels.
Larger patches provide more spatial context but require substantially more GPU memory and computation.
A practical starting point is 32. Reduce the patch size if GPU memory is insufficient or if your contiguous labeled Z-range is small.
Because the network downsamples internally, the implementation normalizes patch size to values compatible with the network architecture.
Important For Sparse Labeling
A training patch must lie within a contiguous labeled Z-range.
If Patch size = 32, at least 32 contiguous labeled slices must be available for a valid 32-slice-deep training patch.
If your manually labeled/interpolated span is smaller, use a smaller compatible patch size such as 16 or 24 when appropriate.
Epochs controls how many training passes are performed over the generated patch set.
More epochs can improve convergence, but unnecessarily large values increase training time and may promote overfitting.
For initial pipeline testing, use a small number such as 5–10 epochs. Increase later after confirming that the workflow is functioning correctly.
Batch size controls how many 3D patches are processed together during one training update.
3D networks are memory intensive. Batch size 1 or 2 is a reasonable starting point for many GPUs.
If MATLAB reports GPU out-of-memory errors, reduce batch size first, then reduce patch size if needed.
Learning rate controls the magnitude of optimizer updates.
The default 1e-3 is a practical starting value for the validated training workflow.
If training is unstable, a smaller value such as 1e-4 may help. If learning is extremely slow, the learning rate may be too small.
Train fraction controls the proportion of independent volumes assigned to the training group when multiple datasets are available.
The remaining volumes are reserved for validation.
Typical values are approximately 0.7–0.8 when enough independent volumes are available.
Patches / epoch controls how many training patches are sampled during each epoch.
Increasing this value exposes the network to more spatial locations but increases training time.
Val patches controls how many patches are sampled from the held-out validation volume or volumes for validation calculations.
The validation patches are not used to update the network.
Threshold converts foreground probability into the final binary predicted mask.
The default threshold is 0.50.
Probability greater than or equal to the threshold is classified as foreground.
Changing the threshold does not retrain the model; it changes how the probability output is converted into the binary segmentation.
Use GPU enables GPU acceleration when a compatible MATLAB-supported GPU is available.
3D U-Net training and full-volume inference can be substantially faster on a GPU than on a CPU.
The device indicator reports GPU when the requested GPU is available, or CPU if GPU support is unavailable.
GPU memory is often the limiting resource. If training fails because of memory, reduce Batch size or Patch size.
Train 3D U-Net launches the validated MATLAB trainnet-based 3D U-Net workflow.
For multiple datasets, whole volumes are divided into training and validation groups before patches are sampled.
For a single current volume, the tool trains using the current volume, current mask, and labeled-slice information.
The trained network, model information, training information, and validation statistics are retained in the Studio.
Dice measures spatial overlap between predicted foreground and ground-truth foreground.
Dice = 1 indicates complete overlap. Dice = 0 indicates no overlap.
Validation Dice should be interpreted together with visual inspection and independent testing.
Intersection-over-Union (IoU) is the intersection of prediction and ground truth divided by their union.
Like Dice, higher values indicate greater overlap.
P(fg|GT) summarizes foreground probability inside ground-truth foreground voxels.
P(fg|BG) summarizes foreground probability in ground-truth background voxels.
A useful model generally produces high foreground probability inside the target and low foreground probability in background.
Predict Full Volume applies the trained or loaded network to the currently displayed 3D volume.
Prediction uses overlapping 3D tiles with smooth blending to reduce tile-boundary artifacts.
The prediction panel shows the binary segmentation and the overlay panel shows the predicted foreground on the original image.
Show Probability Map displays the voxel-wise foreground probability for the current slice.
The probability map is useful for identifying uncertain boundaries and for understanding how strongly the model supports the foreground class.
Save Model stores the trained network together with model information and training history in a MAT file.
A saved model can later be loaded into Volumetric AI Studio without retraining.
Load Model restores a previously saved Volumetric AI model.
When available, model metadata such as patch size and prediction threshold are restored to the Studio controls.
Export Mask saves the current predicted 3D segmentation.
MAT export stores the predicted mask and foreground probability volume.
NIfTI export writes the binary predicted mask as a NIfTI volume.
Reset clears the current volume, mask, model, prediction, probability map, dataset list, metrics, and log from Volumetric AI Studio.
Reset does not delete source files from disk.
Original — current slice of the input volume.
Ground Truth — current slice of the loaded reference mask.
Prediction — current slice of the model-generated binary segmentation.
Overlay — prediction superimposed on the original image.
Use the slice slider to inspect the segmentation throughout the Z dimension.
Recommended Development Strategy
Start with approximately 5–10 well-labeled independent volumes to validate the complete workflow before investing time in a large annotation campaign.
Confirm that loading, masks, labeled-slice handling, training, GPU execution, prediction, and export all work as expected.
After the pipeline is stable, expand the dataset with additional representative specimens, experimental conditions, imaging sessions, and biological variability.
Keep a final independent test group that is not used during training or model-development decisions whenever possible.
Use consistent orientation, voxel spacing, intensity preprocessing, and biological labeling rules across datasets whenever possible.
Large differences in acquisition protocol or voxel dimensions can make model training more difficult.
Document modality, resolution, specimen, condition, preprocessing, and annotation conventions for reproducibility.
A high validation Dice from a small number of similar volumes does not by itself demonstrate robust performance on new specimens.
Always inspect predicted masks visually, especially around small structures, weak-contrast boundaries, and volume edges.
Prediction quality depends strongly on annotation quality. Systematic annotation errors can be learned by the network.
Interpolated masks accelerate annotation, but they should be reviewed and manually corrected where anatomy changes rapidly.
The current release is binary. Named mask classes are useful for annotation organization, but all nonzero mask values are currently merged into one foreground class for Volumetric AI training.
A future multiclass Volumetric AI implementation can preserve separate mask values such as Bone = 1, Marrow = 2, and Lesion = 3.
Until multiclass training is enabled, train separate binary models when distinct biological structures must be segmented independently.
Recommended End-to-end Workflow
1. Acquire or reconstruct a true 3D volume.
2. Load the volume into IDCubePro.
3. Create a 3D mask using manual anchor slices.
4. Interpolate between anchors.
5. Review and correct the interpolated ROIs.
6. Save the final labeled volume/mask pair.
7. Repeat for multiple independent specimens.
8. Add the pairs in Training Dataset...
9. Select GPU when available.
10. Set patch size, epochs, batch size, learning rate, train fraction, and patch counts.
12. Review held-out validation Dice and visual predictions.
13. Predict on new independent volumes.
14. Inspect probability maps and segmentation boundaries.
15. Export masks and perform downstream quantitative analysis.
Important
Volumetric AI Studio is intended for true 3D segmentation. Confirm that the third dimension represents spatial depth, not ordinary hyperspectral wavelength channels.
Treat validation as model-development evidence, not as a substitute for independent biological testing.