Skip to main content
Back to timeline
arXivSource publication:

EviGen: Predictive Evidence Scaffolding for Verifiable Clinical Rationale Generation

Synopsis

EviGen proposes a three-layer framework that first retrieves and ranks evidence from a patient's longitudinal EHR by predictive contribution rather than textual relevance using learnable queries trained on outcome labels, then has an LLM generate a citation-grounded clinical rationale within that evidence scaffold, and finally applies a process-supervised verifier to check each reasoning step and flag unreliable claims; across MIMIC-IV one-year mortality, autism, and ADHD tasks it improves prediction performance and rationale faithfulness over full-context and RAG baselines and is preferred by most reviewers in a clinical pilot.

Source-provided article image: EviGen: Predictive Evidence Scaffolding for Verifiable Clinical Rationale Generation
Figure 2 ·

Figure 2: EviGen three-layer pipeline. (a) Evidence Selection Layer (§ 3.2 ); (b) Rationale Generation Layer (§ 3.3 ); (c) Process Verification Layer (§ 3.4 ).

arXiv

Interpretation

Introduces a three-layer framework for verifiable clinical rationale generation over full longitudinal EHRs, connecting prediction, rationale generation, and verification through the same patient-specific evidence. The authors describe it as the first framework for clinical rationale generation over full longitudinal EHR histories spanning years of patient care, whereas prior work largely targets single encounters or document classification. Supported by the framework design and end-to-end experiments on three tasks, including comparisons against full-context, RAG, and several supervised baselines.

Uses learnable queries as evidence detectors with patient-conditioned activation, enabling efficient retrieval of heterogeneous evidence (clinical notes and ICD codes) over arbitrarily long input. Compared with IRIS's fixed query set, it adds patient-conditioned gating so only queries matching a patient's profile fire; compared with RAG, retrieval is driven by predictive utility rather than textual similarity. The IRIS comparison is the exact ablation of gating, and EviGen outperforms IRIS on all three datasets, with the gap widening to 2.7-3.5 percentage points on the longer autism and ADHD cohorts.

Constrains generation with an evidence scaffold so each reasoning step cites a passage traceable to the original record and carries a signed attribution score. Extends predictive evidence retrieval beyond document classification to guide generation, making the rationale mechanically checkable via citation IDs and verbatim quotes. A matched 2x2 experiment shows adding the evidence scaffold raises the joint pass rate by 9.7-21.2 percentage points, and switching to IG-ranked evidence adds a further 3.4-11.6 percentage points.

A process-supervised verifier flags unreliable claims at the reasoning-step level and attaches an error type. Adapts the process-supervision paradigm from clinical note verification to auditing generated clinical rationales, producing type-aware error flags. Step-level binary error-detection AUROC of 0.978, calibrated F1 of 0.947 and ECE of 0.028, with 88.2% correct error typing; in a ranking evaluation on naturally generated rationales, the lowest-scored group had a faithfulness pass rate 11.81 percentage points below the average of the other nine ranks.

Perspective

The work targets longitudinal clinical prediction tasks that require integrating sparse evidence across time, in settings where outcome labels are available from structured fields and records may exceed model context length; the authors note future extension to additional modalities such as lab values and clinical imaging, and integrating the verifier into the generation loop rather than only as a post-hoc check. Clinical utility findings currently come from a pilot with 7 medical students, and the authors state that a larger study with licensed clinicians is needed to further confirm it, with such a study underway.

The verifier's evaluation on naturally generated rationales uses a judge-based ranking rather than human-annotated step-level labels, and the authors state they have not conducted a controlled reader study testing whether displaying verifier flags improves human review quality or efficiency, so the verifier should be regarded as a preliminary screening tool. The clinical reviewer pilot is small and uses pre-licensure trainees, making its findings preliminary. The autism and ADHD cohorts come from protected institutional data and cannot be shared, while the MIMIC-IV setting is fully reproducible; the authors note they did not exhaustively audit target-related free-text mentions before the cutoff. In addition, a negative label denotes no recorded target diagnosis in the available EHR rather than confirmed lifetime absence of the condition.

Sources