IMERSE Lab

Margin Model

Identifies tumor margins in real-time endoscopic video, distinguishing tissue still requiring resection from tissue already removed or healthy anatomy.

Two-stage architecture:

  • Stage 1, Spatial Constraint Model: full tumor segmentation, narrows the search region
  • Stage 2, Margin Segmentation Model (margin58-SWA): precise margin segmentation within that narrowed region

Configuration selected through both quantitative and qualitative review, favoring consistent real-world performance over the single highest validation score. Trained on several hundred hand-annotated frames spanning approximately 65 tumor cases.

Stage 1: Spatial Constraint Model

  • Performs full tumor segmentation to localize the tumor within the frame
  • Narrows the search area before margin segmentation runs
  • Retrained on approximately 600 selected frames, about half the data used by the model it replaced
  • Outperformed that larger-data version
  • Reliable across most cases, with reduced accuracy on a small number of atypical tumors
Original
Model Output

Stage 2: Margin Segmentation Model (margin58-SWA)

  • Performs fine-grained segmentation within the region identified by Stage 1
  • U-Net with ResNet-34 encoder, trained with a combined Focal, Tversky, and boundary-aware loss
  • Refined through stochastic weight averaging
  • Configuration selected over a numerically higher-scoring checkpoint after video review found that model produced undetected holes in otherwise-correct masks
Original
Model Output

Architecture and Loss Function

ComponentConfiguration
ArchitectureU-Net, ResNet-34 encoder, ImageNet-pretrained
TaskBinary segmentation (background / tumor margin), 512x512
Focal Lossgamma = 2.75, focuses training on hard-to-classify pixels
Tversky Lossalpha = 0.10, beta = 0.9, penalizes missed tumor more heavily than false positives
Boundary Losssigma = 3.0, emphasizes accuracy near the mask boundary
Class Weighting5.0x on tumor pixels (approximately 2% of all pixels)

Training Configuration

ComponentConfiguration
AugmentationRotation (+/-20 degrees), color jitter, CutMix
OptimizerAdamW, learning rate 3e-4, weight decay 5e-5
SchedulerReduceLROnPlateau, 12-epoch warmup
StabilizationGradient clipping, early stopping (patience 35)
Final WeightsStochastic Weight Averaging from epoch 35, with BatchNorm recalibration

Hyperparameters were not tuned solely for the highest Dice score. The deployed configuration (0.842 Dice) was selected over a higher-scoring alternative (0.851 Dice) after video review found undetected holes in otherwise-correct masks.

Validating Design Decisions Through Direct Comparison

Training With vs. Without Unlabeled Pre-Resection Frames

  • Fifty unlabeled frames captured before tumor resection began, added to training
  • Teaches the model to predict no margin when no tumor boundary is present
  • Reduces false positives in pre-resection frames
With Pre-Resection Frames
Without Pre-Resection Frames

With vs. Without EMA (Exponential Moving Average)

  • Running weighted average of mask probability over time
  • Reduces frame-to-frame flicker at the tumor edge from lighting, glare, and smoke
  • Does not alter the boundary's underlying sharpness
With EMA
Without EMA

With vs. Without the Spatial-Constraint Stage

  • Stage 1 narrows the search region before fine-grained segmentation
  • Suppresses false positives the second stage would otherwise register elsewhere in frame
With Spatial Constraint
Without Spatial Constraint

Upcoming: A formal ablation study evaluating which individual components contribute most to performance is underway, using a newly constructed test set, in preparation for a conference submission.