Skip to main content
Back to timeline
arXivSource publication:

Retention-Constrained Post-Training Quantization of Cellpose–SAM: An Auditable Compression Protocol for Stem Cell Microscopy

Synopsis

The work proposes a pre-specified retention protocol and applies it to several post-training quantization schemes for Cellpose–SAM on a stratified 176-field public panel spanning BBBC038 nuclei, BBBC039 U2OS fluorescence, and NIST iPSC images: weight-only W8A16 preserves instance F1 across all modalities, a sensitivity-guided mixed W4/W8 scheme with four INT8 exception operators reduces weight storage from 1162.07 MiB to 171.86 MiB with no observed catastrophic failures, while ternary W2A16-G64 compresses to 96.21 MiB but fails catastrophically on 169 of 176 fields, showing that compression should be judged by modality-stratified downstream retention rather than a single accuracy number.

AI-generated editorial illustration: Retention-Constrained Post-Training Quantization of Cellpose-SAM for Stem Cell Microscopy

Interpretation

It defines and executes a pre-specified retention protocol: paired instance F1 at an IoU threshold together with AP endpoints as the downstream metric, a 1% mean-change margin fixed before any hold-out data are examined, and per-(scheme, modality) verdicts computed with a cluster-bootstrap 95% interval over experimental units. Relative to reporting compression by a single accuracy number or an output-tensor proxy, the verdict is pushed down to the mask-level metric and forced to be modality-specific, so a passing verdict on one modality cannot silently obscure a failing verdict on another. The protocol elements are stated item by item in the methods; verdicts rest on 176 hold-out fields with the experimental unit as the bootstrap resampling atom, and released audit artifacts allow independent re-derivation.

Across the three imaging modalities, weight-only W8A16 and the sensitivity-guided mixed W4/W8 scheme both meet the retention criterion, and neither shows an observed catastrophic failure on the 176 fields (rule-of-three upper bound 0.028). The mixed scheme identifies the four highest-sensitivity operators (blocks.0.mlp.lin2, blocks.23.attn.qkv, neck.2, and out) via one-operator-at-a-time W4 perturbation on the development split and keeps them at INT8, deepening compression from 1162.07 MiB to 171.86 MiB (6.76x) while tying W8A16 for the tightest failure count. Conclusions come from the modality-stratified hold-out panel and cluster-bootstrap intervals; the authors state explicitly that no observed failure is not a zero failure rate and quote a rule-of-three upper bound.

Calibrated W8A16-QDQ-obs (weight-only packing) and W4A16-G64 each have 1 catastrophic field out of 176 (interval [0.000, 0.001]), with storage of 295.78 MiB and 168.47 MiB respectively. Reporting mean retention and catastrophic-failure rate side by side lets readers see different facets of the same compression tier: average preservation versus worst-field behavior. Failure counts and bootstrap intervals are given in the results tables; W4A16-G64's widest arm on NIST iPSC remains above the threshold but is labeled marginal retention.

Ternary W2A16-G64 fails catastrophically on the segmentation-dominant modalities: it fails severely on BBBC038 and BBBC039, and although the NIST iPSC arm shows only a small paired F1 drop because FP32 is already floor-limited there, the compressed model produces essentially no correct instances on that modality either; the panel-level failure rate is 169/176 with interval [0.949, 1.000]. The result places the intuition that higher compression is better under the retention criterion: a 12.08x storage advantage does not by itself constitute deployability. Failure rates and intervals are given in the results tables and paired with the storage table, giving a direct contrast between compression gain and retention cost.

Perspective

The protocol addresses cell-imaging deployers who need auditable compression decisions, especially iPSC monitoring pipelines on lab-bench and edge hardware; it requires a retention target appropriate to the downstream use, a stratified panel covering the distributional axes that matter, an experimental unit over which resampling makes physical sense, and a release path for audit artifacts. Passing per-modality retention on the NIST iPSC arm is a necessary but not sufficient condition for iPSC deployment: necessary, because monitoring protocols sample the full density trajectory and a scheme that silently degrades in the high-density regime would delay passaging decisions or contaminate downstream release-testing counts; not sufficient, because FP32 itself does not segment reliably on high-density iPSC in this panel.

The retention verdicts are scoped to the modalities evaluated on this panel; passing on BBBC038, BBBC039, or NIST iPSC does not certify the same scheme on modalities with different statistics (3D volumetric microscopy, time-lapse phase contrast at extended intervals, or non-nuclear stains). The 1% margin suits a monitoring pipeline that can absorb a bounded per-modality F1 drop, whereas regulatory release-testing pipelines require a tighter margin; the protocol scales but the verdicts do not transfer. In addition, the FP32 reference is itself floor-limited on the NIST iPSC arm, so that arm's verdict primarily shows the protocol correctly flagging a floor-limited modality rather than showing that the compressed model preserves detection quality on iPSC; a sufficient condition would need a NIST-analogue arm where the FP32 reference actually detects instances, assembled from a workflow-specific dataset before deployment.

Sources