MPFlow guides a rectified-flow prior with auxiliary MRI at inference, matching diffusion baselines at 20% of sampling steps and cutting tumor-hallucination Dice by 15%
Synopsis
The work proposes MPFlow, a zero-shot multi-modal MRI reconstruction framework built on rectified flow that uses a self-supervised pretraining strategy, PAMRI, to learn shared cross-modal representations and jointly guides the unconditional prior with data consistency and cross-modal feature alignment at inference; on HCP T2 4x super-resolution and BraTS FLAIR 8x k-space reconstruction it matches diffusion baselines in image quality using only 20% of the sampling steps while improving tumor segmentation Dice by 15% and reducing the SHAFE hallucination score by 26%.
Interpretation
It formulates multi-modal zero-shot MRI reconstruction, where an unconditional generative prior leverages auxiliary modalities (e.g. fully sampled T1) at inference without updating or retraining that prior. Prior zero-shot reconstruction methods are unimodal and lack a principled mechanism to incorporate auxiliary modalities when the generative prior is unconditional; this work is, to the authors' knowledge, the first zero-shot MRI reconstruction framework that incorporates auxiliary imaging modalities at inference time. The paper gives an information-theoretic argument: H(x|y,xaux)=H(x|y)-I(x;xaux|y), and since registered MR modalities share overlapping but non-identical anatomical information, I(x;xaux|y)>0, so uncertainty decreases; this is then tested empirically on HCP and BraTS.
It proposes PAMRI, a patch-level cross-modal contrastive pretraining strategy using an adaptive InfoNCE loss whose temperature is set by the normalized mutual information of paired patches, plus a patch reconstruction regularizer. Conventional contrastive learning discards spatial detail while MRI reconstruction is a dense task; PAMRI operates on 32x32 patches with an auxiliary reconstruction task to preserve structure, and uses mutual information to weight the latent contrastive space rather than as a direct pixel-level alignment loss, which struggles with multi-modal intensity variation (e.g. a tumor bright in FLAIR but nearly isointense in T1). On HCP at T=100, PAMRI achieves the highest SSIM and lowest SHAFE compared with simpler auxiliary guidance losses (NMI, Canny edge, pixel-MSE); Canny edge and pixel-MSE hallucinate more than the data-consistency-only baseline.
It proposes the MPFlow inference framework, in which the rectified-flow prior is jointly guided by data consistency (suppressing intrinsic hallucinations) and PAMRI latent alignment (suppressing extrinsic hallucinations), with initial noise optimization to mitigate poor trajectory initialization. Previous methods optimize via data consistency only, whereas here the composite objective Phi(x)=||F(x)-y||^2 + lambda_P LP(x, xaux) is also used for seed selection, targeting intrinsic and extrinsic hallucinations together. Ablations show the two modules address complementary failure modes: adding PAMRI yields a 27% SHAFE reduction on HCP and 5% Dice improvement on BraTS with comparatively smaller measurement-loss change, while noise optimization alone reduces measurement-space loss by 82% on HCP and 89% on BraTS with smaller SHAFE/Dice gains.
On HCP and BraTS, MPFlow outperforms zero-shot baselines on both image quality and hallucination metrics while being substantially more efficient. MPFlow matches diffusion baselines' reconstruction quality at T=500 using only T=100 steps; at T=500 it exceeds the second-best baseline DynamicDPS by 2-4% in SSIM and 6-22% in LPIPS, and relative to DPS it reduces measurement-space loss by over 80% on HCP and 88% on BraTS while improving SHAFE by 38% and Dice by 16%. All improvements are statistically significant (p<0.05); HCP testing uses N=200 and BraTS uses the official validation set N=250; in the time-constrained setting (T=100) diffusion methods drop 15-83% in SSIM and LPIPS with increased variance, whereas MPFlow degrades only 2.5-14%.
Perspective
The framework targets brain MRI settings with a registered auxiliary modality: T2 4x super-resolution on HCP and FLAIR 8x k-space reconstruction on BraTS, both using fully sampled T1 as the auxiliary modality, with experiments on axial slices. It lets routinely acquired multi-sequence protocols serve directly as an extra information source at inference, without retraining or updating the unconditional prior, yielding more reliable anatomical fidelity and fewer sampling steps under severe ill-posedness. It is meant for imaging workflows that have multi-contrast or multi-parametric protocols and can perform inter-modality registration; the paper lists non-imaging modalities and adaptive guidance schemes as future directions.
Several open questions remain for a careful reader: how the registration quality between auxiliary and target modalities affects cross-modal guidance is not developed in the main text; the benefit of PAMRI grows with task severity (Delta SSIM rising from 3.81% at 4x to 6.02% at 8x), and whether this trend holds for other degradation types and anatomies remains to be verified; hallucination metrics such as SHAFE and Dice depend on pretrained models (DINO, Swin-UNet), whose own biases are not discussed in the main text; and this parse is based on the main text and tables, so the visual details of Fig. 2 and Fig. 3 are known only from the written descriptions.
