SurgDepth: Test-Time Adaptation of Depth Foundation Models for Surgical Scene Understanding
Synopsis
The work proposes SurgDepth, a test-time adaptation framework that adapts natural-image-pretrained depth foundation models to surgical endoscopy without any labeled surgical data, combining two self-supervised signals (stereo photometric consistency and flip equivariance, the latter requiring only single images) with selective decoder adaptation and two tuning-free reliability mechanisms (progress-aware anchor regularization and a structural-trust safeguard for poorly-illuminated sequences); the authors report AbsRel reductions of up to 63% with stereo pairs and 42% without, improvement in four of five cross-domain evaluation settings, generalization across seven foundation models, and the lowest AbsRel (0.
Interpretation
Proposes SurgDepth, a test-time adaptation framework requiring no labeled surgical data, using stereo photometric consistency and flip equivariance (single images only) as self-supervised signals together with selective decoder adaptation. Existing adaptation methods require surgical training data for each new camera and clinical environment, whereas this framework adapts at test time without labeled surgical data. The abstract presents the framework design and the two self-supervised signals, and reports AbsRel reductions in both stereo and non-stereo settings (up to 63% and 42%).
Introduces two tuning-free reliability mechanisms: progress-aware anchor regularization and a structural-trust safeguard for poorly-illuminated sequences. These mechanisms target the stability of test-time adaptation in surgical scenes and require no additional hyperparameter tuning. The abstract states the mechanisms' target setting (poorly-illuminated sequences) and their tuning-free property; detailed ablations are not expanded in the abstract.
Conducts the first systematic study of test-time adaptation for dense depth regression, comparing eight baselines, and finds that existing TTA methods consistently fail in this setting. Systematic examination of TTA for dense depth regression was previously lacking; the study reports that entropy-based methods degrade regardless of the number of adapted parameters and that temporal-photometric adaptation diverges on non-rigid tissue. The conclusion comes from a systematic comparison across eight baselines, a methodological observation within this setting.
Improves four of five cross-domain evaluation settings (three of four unseen datasets; Hamlyn evaluated under two protocols), generalizes across seven foundation models, and attains the lowest AbsRel (0.128) on the complete standard Hamlyn benchmark among all reported methods. Extends adaptation results from a single model and dataset to multiple models and datasets, and provides a comparable number on a standard benchmark. Results rest on cross-domain evaluation and standard-benchmark comparison, explicitly limited to relative (median-scaled) depth; 20-step convergence (0.4 s) supports its time efficiency.
Perspective
The result targets surgical endoscopy applications that need relative (median-scaled) depth, such as augmented-reality overlay and reconstruction initialization; adaptation happens at test time without labeled surgical data, can use stereo pairs or single images only, and has been validated across seven foundation models and multiple cross-domain settings. For tasks that depend on absolute-scale depth, applicability depends on how later work combines relative depth with scale information.
The currently available text is the abstract, without figures, per-dataset results, or ablations, so the individual contribution of each reliability mechanism, the specific conditions under which the eight baselines fail, and the differences between the two Hamlyn evaluation protocols remain to be confirmed in the full text; moreover, all results are median-scaled relative depth, and behavior under absolute-scale depth is an open question.
