M²PFN turns a frozen TabPFN into a multimodal Alzheimer's predictor, reaching 65.55% macro-F1 on ADNI and transferring to two external cohorts without retraining
Synopsis
The work proposes M²PFN, an end-to-end framework that back-propagates task gradients into 3D-MRI and tabular encoders through differentiable inference, aligns the two modalities into a shared subspace matched to the in-context learning prior via disentanglement and a contrastive objective, and folds in a frozen tabular-only prediction through a learnable gated shortcut; on ADNI (n=2240, three-class CN/MCI/AD) it attains 65.55% macro-F1 and 82.21% macro-AUC, surpassing the compared unimodal and multimodal baselines, regresses baseline MMSE by swapping only the head (1250-subject sub-cohort, test MAE 1.743), and achieves the best AUC and lowest MMSE MAE across all baselines on two external cohorts (OASIS-3 and SCAN) with no retraining, transferring even when the cognitive instrument changes.
Figure 1: Overview of the proposed M 2 PFN framework (§ 3 ). (a) Overall architecture. (b) Disentangler module for separating modality-specific and shared representations. (c) Class prototype alignment for encouraging cross-modal consistency. (d) Differentiable TabPFN backbone enabling end-to-end gradient propagation through the frozen PFN engine.
arXivInterpretation
M²PFN extends the in-context learning capability of the tabular foundation model TabPFN to multimodal Alzheimer's disease analysis, keeping the frozen ICL engine's in-context mechanism for test-time generalization while end-to-end training shapes the 3D-MRI and tabular encoders into features that engine can exploit. Prior multimodal AD methods combining imaging and tabular data are limited in cross-cohort generalization, and TabPFN's meta-trained synthetic tabular priors do not naturally match the statistical structure of image-derived features; this work addresses that mismatch directly by performing differentiable inference through TabPFN's transformer and back-propagating task gradients. The abstract reports 65.55% macro-F1 and 82.21% macro-AUC on ADNI (n=2240, three-class CN/MCI/AD) and states it surpasses a comprehensive set of unimodal and multimodal baselines; the specific baseline list and statistical tests are not given in the loaded text.
The framework aligns the two modalities into a shared subspace matched to the ICL engine's prior via disentanglement and a contrastive objective, and folds in a frozen tabular-only prediction through a learnable gated shortcut. The alignment objective is not generic multimodal fusion but is explicitly designed to match the in-context learning engine's prior, while retaining a tabular-only prediction path as a gated shortcut. The abstract presents this as a method description without ablation numbers or a quantitative decomposition of each component's contribution.
The same architecture regresses baseline MMSE by swapping only the head for a TabPFN regressor, reaching test MAE 1.743 on a 1250-subject sub-cohort and outperforming every multimodal baseline. This shows the architecture is not limited to classification and can move to continuous cognitive-score prediction by changing the head, exceeding the compared multimodal baselines. The abstract reports test MAE 1.743 and the 1250-subject sub-cohort size and states it outperforms every multimodal baseline; confidence intervals or per-baseline values are not given.
On two external cohorts, OASIS-3 and SCAN, M²PFN achieves the best AUC and the lowest MMSE MAE across all baselines with no retraining, and transfers even when the cognitive instrument changes. It extends cross-cohort generalization from same-instrument settings to an instrument-change setting without relying on external-cohort retraining. The abstract states the external-cohort results and instrument-change transfer but does not give specific AUC or MAE values or external cohort sample sizes.
Perspective
The result targets Alzheimer's disease research settings that take 3D-MRI and tabular data as input and perform CN/MCI/AD classification or MMSE regression, in multi-cohort setups such as ADNI training with OASIS-3 and SCAN external validation; its value lies in letting the frozen ICL engine retain its in-context mechanism at test time while training encoders into features that engine can exploit, making it most directly relevant to researchers and model developers who want to reuse a tabular foundation model without retraining on external cohorts.
The loaded text is abstract-level and contains no figures, per-baseline numbers, ablations, or statistical tests, so the relative contribution of each component (differentiable inference, disentangled alignment, gated shortcut), the specific AUC and MAE on external cohorts, and the magnitude of instrument-change transfer remain open questions to verify; moreover, how to interpret the absolute level of 65.55% macro-F1 in a three-class task requires further judgment in light of class distribution and clinically usable thresholds.
