SlideRuler calibrates pathology foundation models using within-slide controls, cutting embedding distance by 16.3–38.5% across five scanners
Synopsis
The authors introduce SlideRuler, which treats other regions of the same slide as internal controls and learns a transfer map from paired rescans to estimate and correct acquisition-induced shifts, enabling calibration from a single scan at inference while keeping the foundation model frozen; across two encoders and five SCORPION scanners, learned transfer reduces mean target-to-source embedding distance by 16.3–38.5% relative to raw embeddings, same-slide controls outperform unrelated same-scanner controls, and a source-anchored variant reduces source-feature displacement by 47.7–83.6% while retaining most of its alignment gain.
Figure 1: One-slide calibration with learned transfer. Top: paired rescans define the acquisition basis U U from residuals centered across scanners within each registered region and train the transfer map. Bottom: a single scan supplies control displacements relative to source prototypes; learned transfer corrects the query while preserving Q z j Qz_{j} . Controls exclude the query; training prototype matches exclude its entire slide. The nine-control SCORPION example is schematic.
arXivInterpretation
Same-slide controls carry information beyond the shared same-device correction: replacing a slide's own controls with those of another slide on the same scanner leaves the paired distance improvement positive in all four evaluation settings, at 0.535 and 0.303 for Prost40M and 0.330 and 0.249 for DistillPath, with all four pooled intervals excluding zero and the pooled effect positive under each of the 32 permutations. Prior paired-rescan methods learn a global transformation shared across slides; by holding the scanner and query fixed and changing only the identity of the control tissue, this work separates the current slide's own contribution from the shared acquisition correction. Thirty-two seeded derangements with pointwise 95% intervals from 5,000 physical-slide bootstrap resamples, plus 16 scanner-specific intervals that all exclude zero; the same-slide increment accounts for 6.3–8.9% of learned transfer's total distance reduction.
The source-anchored variant makes the tradeoff between target alignment and source-feature displacement explicit: relative to learned transfer it reduces source-feature displacement by 47.7–83.6%, with decreases on all 48 slides in every setting, at the cost of a 0.291–0.749 increase in mean target distance, while retaining 79.8–93.2% of learned transfer's distance reduction over raw embeddings. Learned transfer optimizes target alignment alone; this work adds an anchored residual estimator whose model selection is guided by source-feature displacement, giving deployers a choice of operating point between alignment and source-side fidelity. All four paired intervals exclude zero; training-only model selection retains a residual component in 42 of 60 unique fits and selects the control-modulated global correction in the remaining 18, with all Prost40M fits retaining a residual and all six calibrated DistillPath fits selecting the global correction.
Availability of tissue-defined controls is limited by sampling density: an external three-class selector achieves patient-weighted precision of 89.37% for benign glands, 99.98% for residual stroma, and 96.32% for tumor at 96.75% coverage, and on PAR raising sampling from 512 to 2,048 fields per scan increases scans meeting both control-count requirements from 29/105 to 58/105, with benign-gland sampling the remaining constraint. Prior work largely assumes control tissue is available; this work treats control availability itself as a measurable endpoint and quantifies the relationship between selector precision and sampling policy. 8,325 fields from 37 slides and 32 patients with each patient weighted to one and 2,000 patient bootstrap resamples; PAR uses a fixed 20 patients, 35 slides, and 105 scans with all assigned scans kept in the denominators.
Controlled synthetic experiments attribute the benefit of learned transfer to tissue-response differences and the downstream head's sensitivity to acquisition shifts: in the acquisition-sensitive regime the two-compartment method reduces shift RMSE from 4.014 to 1.894, raises worst-target QWK from 0.828 to 0.949, and cuts grade flips from 56.7% to 15.7%, while the QWK gain is only 0.009 for an already robust head and learned ungated transfer improves RMSE from 3.766 for direct subtraction to 1.662 when lesion and control responses differ strongly. Real-scanner data cannot supply known shifts; this work uses 64 synthetic runs across eight regimes and eight seeds to provide a controlled comparison with known acquisition shifts, showing under which tissue and head conditions the method helps. All eight seeds contribute to the reported means, with paired inference using 10,000 seed-level percentile bootstrap replicates and exact two-sided sign-flip tests with Holm adjustment over three comparisons; RMSE reduction and QWK gain in the acquisition-sensitive regime are each significant after adjustment.
Perspective
The work addresses two questions, embedding calibration and tissue-control availability, and applies when inference sees only a single scan, the foundation model stays frozen, and identifiable control tissue exists within the slide, as benign glands and stroma do in prostate biopsies. It lets acquisition shifts be estimated and corrected across hospitals and devices without paired rescans, scanner identity, or target-cohort statistics, and it offers an optional anchored operating point for downstream tasks sensitive to source-side fidelity. The selector and sampling results give a quantitative basis for designing control acquisition: on PAR, raising sampling from 512 to 2,048 fields per scan increases scans meeting both control-count requirements from 29/105 to 58/105. The method is positioned for research use.
SCORPION provides aligned regions and physical-slide groups but no patient identifiers, compartment annotations, or clinical grades; its bootstrap intervals condition on the fitted models and development choices, and public metadata do not permit a complete audit of pretraining overlap. Acquisition coordinates may contain biological variation, and reduced source-feature displacement does not establish preservation of downstream predictions. The optional strict gate accepts no SCORPION corrections at this sample size, so the reported real-scanner results concern the ungated estimators. PAR control counts are an intermediate endpoint: registration and matched-tumor requirements leave at most six eligible training patients, below the fixed ten-patient minimum, so no end-to-end PAR calibration or grading comparison was run; control purity on PAR requires expert validation, and denser sampling yields spatially clustered fields. Synthetic stress tests show that agreement between control compartments does not protect against shared control drift or altered lesion responses. These point to next validation priorities: expert-verified controls, patient-separated downstream evaluation, and assessment of correction decisions under broader acquisition changes.
