An Explainable Multimodal Deep Learning Framework for Alzheimer's Disease Classification Using MRI, PET, and Clinical Data
Synopsis
Using ADNI data, this study builds a patient-level multimodal fusion framework based on ResNet50 transfer learning that concatenates MRI and PET imaging features with clinical variables (ADAS11, ADAS13, APOE4, age, sex, education, MMSE total) in a fully connected network to produce three-way AD/MCI/CN predictions, with Grad-CAM heatmaps for explanation; the ablation shows accuracy rising as modalities are added, from 39.61% for MRI alone and 51.72% for PET alone to 66.67% for clinical-only, 54.17% for MRI+PET, and 79.17% for the full three-modality model on a 24-patient test cohort (95% Wilson interval roughly 59.5%-90.8%), with the addition of clinical data producing the single largest gain.
Interpretation
It proposes and evaluates a patient-level feature-level fusion framework that incorporates MRI, PET, and clinical variables together for three-way AD/MCI/CN classification rather than binary classification. In the qualitative comparison of Table I, the authors note that prior MRI+PET fusion studies (Odusami, Zhang, Song, Goel, Castellano) all omit clinical data, while the one study using clinical data, El-Sappagh et al., omits imaging entirely, so the combination of three modalities, three-way labeling, and Grad-CAM has no precedent in that literature. Evidence comes from a five-configuration ablation on ADNI data (Tables IV and V), but the patient-level fusion results rest on a 24-patient held-out test set, for which the authors also report 95% Wilson confidence intervals and note the intervals are wide.
The ablation quantifies each modality's marginal contribution and finds that clinical variables alone outperform either imaging modality and their combination in this cohort. The authors add a Clinical-only baseline (66.67%) to Table V, note it exceeds MRI (39.61%), PET (51.72%), and MRI+PET (54.17%), and state that prior work generally does not isolate the individual contribution of clinical data within a single ablation. Based on the same 24-patient test cohort; the authors report Fisher's exact tests: MRI+PET (13/24) versus the three-modality model (19/24) gave p = 0.125, and Clinical-only (16/24) versus the three-modality model (19/24) gave p = 0.517, noting that per-patient joint correctness records were not retained, so the unpaired tests are conservative.
Grad-CAM is integrated into the imaging branches alongside an interactive diagnostic interface that reports the predicted class, a confidence score, per-class probabilities, a list of affected brain regions from a class-conditioned atlas lookup, and a plain-language explanation. A class-conditioned diagnostic check indicates the PET branch's Grad-CAM response varies with the class being explained: computing heatmaps for one fixed PET slice under the AD, CN, and MCI hypotheses in turn produced non-identical maps (activation toward the right hemisphere under the AD hypothesis, a lower-central hotspot under CN, and near-absent activation under MCI). This is a qualitative inspection of three representative correctly classified cases; the authors explicitly describe it as illustrative single examples rather than a systematic evaluation, and note the affected-region labels come from a fixed lookup table rather than heatmap coordinates.
It uses a lightweight 2D slice-level ResNet50 backbone with mean pooling to aggregate to the patient level, avoiding the compute cost of 3D convolutional or transformer architectures. In the Table I comparison, the authors note that most fusion frameworks rely on computationally intensive end-to-end 3D convolutional architectures, whereas this framework substitutes feature-level concatenation fusion, and they explicitly flag the slice aggregation scheme (averaging features from the central 50 slices) as a design choice rather than an incidental detail. The method is described in full (Sections III-D to III-G), but the authors note that alternatives such as max-pooling, attention-weighted pooling, or aggregating softmax probabilities may yield different downstream fusion performance, leaving validation to future work.
Perspective
The framework targets the research setting of three-way AD/MCI/CN classification within the ADNI cohort and applies to subjects for whom MRI, FDG-PET, and structured clinical assessments (ADAS11, ADAS13, APOE4, age, sex, education, MMSE total) are all available; the imaging branches are evaluated at the slice level (MRI n=1608, PET n=696) and the fusion branches at the patient level (n=24). The authors identify enlarging the cohort, adopting 3D convolutional networks or Vision Transformers to exploit volumetric spatial context, adding biomarkers such as CSF measures or genetic risk factors, extending to longitudinal multi-scan analysis, and using federated learning across institutions as directions the framework can be extended toward.
A careful reader might watch several things: the stability of the 79.17% patient-level point estimate on a 24-patient test set awaits confirmation on larger, multicenter cohorts; the 25-point gain from clinical data and the 12.5-point residual gain from adding imaging on top of clinical-only can only be tested with unpaired Fisher tests because per-patient joint correctness records were not retained, limiting statistical power; the Grad-CAM heatmaps share a similar shape across the three classes on MRI and extend into peripheral skull structures, while PET overlays retain isolated activation patches near the interior brain-mask boundary, which the authors attribute to upsampling the low-resolution final convolutional feature map without a tighter brain-tissue mask, so they should not yet be used for precise anatomical claims; the affected-region labels come from a fixed class-conditioned lookup table rather than heatmap coordinates; and the explainability analysis covers only three representative cases, with a quantitative overlap analysis across the full cohort left for future work. In addition, the experimental environment in the loaded text omits exact hardware and library version numbers, which the authors state will be reported in the camera-ready version.
