CA-WCP gives per-class 90% volume coverage on BraTS 2020 and synthetic CT with intervals 8–14% narrower than symmetric weighted conformal prediction
Synopsis
The work proposes Class-Aware Asymmetric Weighted Conformal Prediction (CA-WCP), which combines latent-space density-ratio weighting with directional quantiles that separately calibrate lower and upper volume bounds and scales each bound by a class-specific asymmetry factor derived from validation-set false-positive and false-negative rates; the authors prove it retains the weighted-exchangeability marginal coverage guarantee per class and show on BraTS 2020 and a covariate-shifted synthetic multi-organ CT benchmark that 95% Clopper–Pearson intervals for observed coverage contain the nominal 90% level for every class while interval width shrinks 8–14% versus symmetric weighted conformal prediction.
Figure 1: Class-specific error patterns revealing asymmetric biases in brain tumor segmentation, measured on the validation split. (Left) False positive versus false negative rates. Rates are normalized by ground-truth volume v ∗ v^{*} (Eq. ( 5 )), so values above 100% are well defined and indicate severe over-segmentation. Necrotic core shows a false-negative bias, edema a large false-positive bias, and enhancing tumor a moderate false-positive bias. (Right) Distribution of volume errors across classes, showing the skew that motivates asymmetric interval design.
arXivInterpretation
CA-WCP composes density-ratio weighted conformal prediction, directional quantiles, and class-conditional (Mondrian) calibration in the 3D multi-class volumetric setting, using validation-set false-positive and false-negative rates to couple interval geometry to the per-class error structure. Prior volumetric conformal formulations largely assume symmetric error distributions and handle class effects only implicitly through adaptive score functions; here class effects are calibrated explicitly and tied to the directional structure of the volumetric error. Evaluated on BraTS 2020 and a synthetic multi-organ CT benchmark against three baselines sharing the backbone, density-ratio weights, and calibration split, so the Directional WCP versus CA-WCP comparison is a direct ablation of the asymmetry mechanism.
Directional quantiles calibrate the lower and upper bounds separately, so the dominant error direction no longer sets both margins, which supplies the efficiency gain. Relative to symmetric weighted conformal prediction computed on the maximum directional score, directional quantiles cut mean width from 9.7 to 6.7 mL. Per-class exact coverage with 95% Clopper–Pearson intervals and mean widths are reported on BraTS 2020; backbone Dice is identical across all conformal variants (necrosis 0.63, edema 0.81, enhancing 0.88).
The asymmetry factors add a conservative, class-specific margin on the side where validation errors concentrate, lifting point coverage to the nominal level. CA-WCP raises point coverage to 90.4% for all three BraTS classes at a mean width of 8.5 mL, 12% narrower than symmetric weighted conformal prediction and 27% wider than directional conformal prediction. Every branch yields an asymmetry factor of at least 1, so the margin can only widen the interval; the authors give a proof of per-class marginal coverage under weighted exchangeability and note the approximation introduced by estimated weights.
Encoding the calibrated intervals into structured prompts lets a multimodal LLM produce reports whose hedging language tracks interval width. Across 10 test cases, adding calibrated intervals raised uncertainty concordance from 2.1 to 4.6, factual accuracy from 3.8 to 4.8, and BLEU-4 from 0.28 to 0.42. Rated by an independent radiologist blinded to the prompt configuration, with reference reports written by a board-certified radiologist mapping ground-truth masks to a standardized template; the authors frame this as a preliminary assessment rather than clinical validation.
Perspective
The framework targets 3D multi-class segmentation settings that need per-class volume coverage guarantees under distribution shift, such as brain tumor subregions and abdominal organs with directional error structure. As a post-hoc wrapper it does not modify segmentation masks and adds no inference cost beyond standard split conformal prediction, so it can be layered onto existing backbones. The calibrated intervals can also be encoded into structured prompts so a multimodal LLM produces radiology reports whose hedging tracks interval width, offering a route from shift-aware uncertainty quantification to interpretable clinical communication.
Results come from a single calibration/test split, and the coverage gap between directional conformal prediction and CA-WCP is not statistically resolved at this sample size, so repeated random splits would be needed to establish whether weight-estimation error causes systematic under-coverage. The synthetic CT data are produced by a learned generator and inherit its anatomical priors; while they isolate the structure-scale effect under a controlled shift, validation on acquired multi-center abdominal CT remains necessary. Test cohort sizes limit the precision of coverage estimation, which is why exact Clopper–Pearson intervals are reported. The report assessment covers 10 cases and a single rater and is best read as an indication of how uncertainty conditioning changes reporting language rather than as clinical validation.
