In four-way mood and psychosis classification across 1,520 subjects, class-conditional (Mondrian) calibration cut the between-diagnosis coverage gap from 12.3 points to 0.4 points at a cost of 0.07 labels in mean set size
Synopsis
On a four-way mood and psychosis classification task over 1,520 subjects from three studies and 14 acquisition sites, the work shows that marginal split-conformal calibration reached 0.9000 empirical coverage against a nominal 0.90 while healthy controls were covered at 0.941 and schizoaffective disorder at 0.818, and that class-conditional (Mondrian) calibration reduced this 12.3-point disparity to 0.4 points at a cost of 0.07 labels in mean set size (under 3%), making set size at matched coverage interpretable as a property of the subject and separating subjects into confident, boundary, ambiguous and unresolved strata, with the proportion independently flagged as label-ambiguous by a structural-MRI model rising monotonically across these strata (34.1%, 57.3%, 68.1%, 81.8%; p = 8.8e-18).
Interpretation
In multi-class diagnostic classification, marginal coverage can be met exactly while individual diagnoses are covered very unevenly: healthy controls at 0.941 and schizoaffective disorder at 0.818, with the surplus and deficit cancelling so the aggregate figure is uninformative about either. Puts the marginal-coverage guarantee, usually treated as sufficient, into a concrete four-way mood and psychosis classification setting and shows it does not describe per-diagnosis reliability when classes are unbalanced. Based on a four-way task over 1,520 subjects from three studies and 14 acquisition sites, reporting 0.9000 empirical coverage for marginal split-conformal calibration against a nominal 0.90, together with per-class coverage values.
Class-conditional (Mondrian) calibration reduced the between-diagnosis coverage disparity from 12.3 points to 0.4 points, at a cost of 0.07 labels in mean set size, under 3%. Provides a calibration approach that equalizes coverage across diagnoses while barely sacrificing prediction-set compactness, and quantifies that cost. Reports the coverage disparity before and after calibration (12.3 points versus 0.4 points) and the mean set-size increase (0.07 labels, under 3%).
Once coverage is held fixed across diagnoses, residual variation in set size can no longer be attributed to class prior or per-class accuracy and becomes interpretable as a property of the subject; the resulting prediction sets separate subjects into confident, boundary, ambiguous and unresolved strata, and the proportion independently flagged as label-ambiguous by a structural-MRI model rises monotonically across these strata (34.1%, 57.3%, 68.1%, 81.8%; p = 8.8e-18). Repositions set size from a by-product of model performance to an interpretable subject-level signal, supported by external agreement from a structural-MRI model trained separately on the same cohort. The stratum proportions rise monotonically with p = 8.8e-18; the text also reports that no schizoaffective subject and 0.9% of bipolar subjects reach the confident stratum, against 18.0% of controls and 13.4% of schizophrenia subjects.
At matched coverage, set size provides a comparison between representations that accuracy cannot: across structural MRI, functional MRI and their fusion there is a sign reversal, with fusion reducing set size for bipolar (-0.25) and schizoaffective (-0.23) subjects and increasing it for controls (+0.16), while the aggregate fusion benefit itself changes sign below alpha = 0.10 and top-1 accuracy rises monotonically from 0.495 to 0.578 to 0.611 across the three models. Shows that set size can expose class-heterogeneous effects concealed by accuracy, and notes that the aggregate fusion benefit is sensitive to the significance level. Reports top-1 accuracy across the three models (0.495, 0.578, 0.611) and the per-class set-size changes under fusion (-0.25, -0.23, +0.16), and states that the aggregate fusion benefit changes sign below alpha = 0.10.
Perspective
The results are meant for multi-class, class-imbalanced diagnostic classification, in particular the four-way mood and psychosis classification in psychiatric neuroimaging. For researchers and developers who want to use conformal prediction in clinical decision support, it indicates that class-conditional coverage should be reported alongside marginal coverage, and that set size at matched coverage can compare representations such as structural MRI, functional MRI and their fusion. The agreement between the subject strata (confident, boundary, ambiguous, unresolved) and the independent structural-MRI label-ambiguity flag offers a direction for interpreting prediction sets at the individual level.
The visible text is an abstract and does not include figures, per-site or per-study breakdowns, calibration details, or the full specification of the statistical tests; therefore the stability of class-conditional calibration across sites, the mechanism behind the relationship between set-size strata and the structural-MRI label-ambiguity flag, and the behavior of the fusion benefit sign as alpha varies remain open questions to confirm against the full text.
