CARing combines compositional medical semantic IDs with a coverage reward to reach the highest recall in next-visit diagnosis prediction on MIMIC-III/IV (R@30 46.04%/46.52%)
Synopsis
The work proposes CARing, which compresses disease names and ICD ontology paths into compositional medical Semantic IDs (SIDs) via residual quantization, grounds them through six alignment tasks plus teacher-generated semantic-enrichment corpora and reasoning activation so a single autoregressive model reasons before emitting a diagnosis SID, and optimizes multi-label coverage with a coverage reward and multi-positive supervision; on MIMIC-III and MIMIC-IV it exceeds all EHR-trained baselines in weighted score and attains the highest recall at every reported cutoff, with R@30 of 46.04% and 46.52% in reasoning mode.
Figure 1: Motivation for multi-label reinforcement learning in next-visit diagnosis prediction. Standard RL can collapse onto one correct diagnosis, leaving others uncovered. Our coverage reward discounts repeated predictions and encourages coverage of distinct correct diagnoses.
arXivInterpretation
Representing diseases as compositional medical SIDs lets one autoregressive LLM read longitudinal histories, reason in natural language, and generate diagnoses within a single interface. Earlier medical tokenizers such as MedTok and MedRQ mainly feed learned discrete representations to downstream predictors, and MERA assigns one atomic token per disease so its vocabulary grows with the diagnosis catalog and natural-language reasoning stays outside its prediction path; CARing composes each disease from shared codebook tokens and places reasoning and SID generation in the same model. The method specifies residual quantization, six alignment tasks (T1–T6), and a reasoning-activation stage; Appendix C.4 reports higher first-token semantic purity for the name-plus-ICD-path variant (top-level ICD weighted purity 69.36% to 90.51%, CCS 47.63% to 69.78%), and Table 2 reports R@30 rising from 38.97 to 39.73.
A coverage reward divides each correct prediction's credit by its frequency in the sampled group, so the group reward equals the number of distinct correct diagnoses, discouraging repeated hits and encouraging coverage of complementary diagnoses. Common RLVR rewards each trajectory only for the correctness of its own final answer and cannot distinguish repeating one correct diagnosis from covering different correct diagnoses; this reward gives the frequency adjustment a direct interpretation as diagnosis coverage and comes with a proof that the group reward counts distinct correct predictions. Appendix B.4 derives this property; in the ablation, the EM reward (S5) gives 18.88 w- and 34.67 R@30, while replacing only the reward with the coverage reward (S6) raises these to 25.41 and 40.27, gains of 6.53 and 5.60 points, also exceeding the reasoning-activation checkpoint S4 (23.57 w-, 39.47 R@30).
Multi-positive supervision supplies learning signal at the answer level for correct diagnoses that sampling does not cover, complementing the coverage reward. The coverage reward can only distinguish correct SIDs once they are sampled, giving no policy-gradient signal for correct diagnoses absent from the rollout group; the auxiliary loss supervises only the SID answer tokens and assigns target positives round-robin over a fixed enumeration. Adding the multi-positive loss to S6 (S7) raises w- and R@30 to 27.58 and 46.04, gains of 2.17 and 5.77 points over S6; training monitoring shows unique SIDs per group at step 60 rising from 7.02 to 12.40, batch-level SID entropy from 4.58 to 5.58, and mean inverse frequency among correct predictions from 0.32 to 0.60.
One framework supports both direct constrained decoding and multi-chain reasoning with reciprocal rank fusion, with both settings restricting predictions to valid medical SIDs. The direct mode generates no rationale and conditions only on the longitudinal SID history under catalog-constrained beam search; the reasoning mode samples several chains, decodes a SID ranking after each chain, and fuses the rankings with RRF. In Table 1 the thinking setting leads at seven of eight recall cutoffs, while direct inference is 0.09 points higher on MIMIC-III R@40; the thinking model scores 2.49 and 2.04 points higher in w- than the direct setting on MIMIC-III and MIMIC-IV.
Perspective
The results apply to retrospective next-visit diagnosis prediction from structured diagnosis sequences: on MIMIC-III and MIMIC-IV, split by patient at 7:1:2 into training, validation, and test sets, with a 10,000-patient sampled MIMIC-IV cohort, evaluated by weighted score and top-k recall. The authors frame the intended use as clinical review, where candidate diagnoses and the thinking mode's readable traces could help physicians revisit earlier visits and consider conditions meriting further assessment, while clinical decisions still rest on examination, testing, and professional judgment and prospective studies would be needed to examine how such assistance changes case review and diagnostic decisions. Methodologically, because SID construction uses concept descriptions and ontology paths, the pipeline could be adapted to other coding systems; adding medication, procedure, and laboratory events to the history could let prediction use a fuller record and support medication recommendation and broader EHR event recommendation.
The clinical meaning of a generated trace requires separate study: the authors note that plausible chains of thought can omit factors that influence a prediction and that their faithfulness varies across models and tasks, and that the teacher receives a reference diagnosis when constructing training traces while the student sees only the history at inference, so the clinical accuracy of both kinds of trace and any artifacts linked to target guidance merit closer study. Proposed follow-up analyses include holding the checkpoint, SID history, and answer format fixed and comparing SID log scores and ranks with and without a trace for each valid SID, measuring how many additional target diagnoses enter top-k candidate lists across sampled chains, and comparing direct and thinking inference with matched sample counts and computation to quantify the contributions of trace conditioning, repeated sampling, and rank fusion. In addition, the coverage-reward analysis notes that finite sampling and advantage normalization do not guarantee a globally uniform policy over all ground-truth diagnoses, which is worth keeping in mind when interpreting coverage gains.
