MPAD synthesizes 3D multi-contrast MRI with 2D diffusion and multi-plane autoregression, cutting training FLOPs 7x and inference FLOPs 3x
Synopsis
The authors propose MPAD: a 3D autoencoder compresses MRI volumes into an isotropic latent representation, a 2D diffusion model reconstructs masked target-contrast latent slices conditioned on source-contrast slices, unmasked target slices, and a text modality prompt, and inference generates slices autoregressively per plane while propagating completed volumes as priors to orthogonal planes, achieving full-volume 3D synthesis with 2D operations, reducing training and inference FLOPs by 7x and 3x versus 3D latent diffusion baselines, and improving PSNR, SSIM, and NMSE on ADNI and IXI.
Figure 1 : Overview of the proposed 3D multi-contrast MR image translation framework. Given a source-contrast (e.g., T1), the model synthesizes a target contrast (e.g., T2 or PD) in latent space using a Multi-modal Conditioning Encoder (MCE) and a 2D diffusion model with plane-wise autoregressive synthesis. Bottom: Training FLOPs per iteration (GFLOPs), inference FLOPs per volume (TFLOPs), inference time, peak memory, and PSNR averaged across all translation tasks on the ADNI dataset.
arXivInterpretation
Reformulates 3D multi-contrast MRI synthesis as a 2D latent masked-slice prediction task, with a 3D autoencoder providing an isotropic latent representation sliceable along any anatomical plane, avoiding the cubic computational scaling of volumetric 3D diffusion. Prior multi-contrast synthesis was largely 2D slice-to-slice translation lacking volumetric coherence, while full 3D generative models preserve structure but scale cubically in compute and memory; this work carries 3D conditioning through a 2D denoising network. The method section details a two-stage training procedure and loss composition; ablations show the 3D multi-modal conditioning encoder outperforms its 2D counterpart (ADNI PSNR +0.72, IXI +1.19), indicating volumetric context contributes to conditioning.
Introduces multi-plane autoregressive inference: slices are generated autoregressively in random order within a plane to maintain intra-plane continuity, subsequent planes are initialized from an intermediate timestep using completed volumes as priors rather than pure noise, and four candidates are aggregated by voxel-wise averaging. Compared with generating each plane independently from noise and averaging, inter-plane priors let later planes inherit structural information from completed planes and reduce sampling steps. Ablation shows enabling priors improves ADNI PSNR by +1.12 and SSIM by +0.037, and IXI PSNR by +1.88 and SSIM by +0.013; performance rises monotonically from one plane to three planes plus a refinement pass; different plane orderings give nearly identical results (e.g., 23.48/23.49/23.54 PSNR), suggesting order insensitivity.
On ADNI and IXI, a single unified model supports one-to-many contrast synthesis and matches or exceeds five baselines (two GANs, three diffusion methods) on most contrast pairs and metrics, with only ALDM also offering one-to-many capability. Most baselines are trained separately per contrast pair, whereas MPAD handles multiple target contrasts in one model while reporting lower training and inference FLOPs, inference time, and peak memory. On ADNI it achieves the highest SSIM in five of six tasks and the lowest NMSE throughout, with PD-to-T1 and PD-to-T2 PSNR more than 2 dB above the second-best method; on IXI it ranks first across all metrics. Evaluation uses PSNR, SSIM, and NMSE, with all competing methods trained for the same number of epochs.
For training, filling masked target latent slices with Gaussian noise rather than learnable mask tokens works better. Learnable mask tokens are a common choice, but here constant values are easily detected and potentially ignored by the conditioning encoder, weakening target-side context; Gaussian noise stays continuous and within the natural range of latent activations. Ablation shows Gaussian noise masking improves ADNI PSNR by +0.32 and SSIM by +0.029, and IXI PSNR by +2.00 and SSIM by +0.017.
Perspective
The result targets 3D brain MRI scenarios where missing contrasts are synthesized from acquired ones, covering T1w, T2w, and PDw conversions, validated on ADNI (737 volumes, 1.5T GE scanners, patients with Alzheimer's disease) and IXI (577 volumes, 1.5T and 3T Philips scanners, healthy participants), with 100 randomly selected subjects per dataset for evaluation. Its value lies in obtaining full-volume output from a 2D denoising network, supporting one-to-many synthesis at lower training and inference cost on standard GPUs; the authors note a controllable trade-off between cross-plane refinement iterations and inference time, where reducing refinement iterations lowers inference time to approximately 1.8 s per volume, allowing adaptation to deployment constraints.
Evaluation metrics are image-similarity measures such as PSNR, SSIM, and NMSE; the text does not report downstream diagnostic task performance, so the relationship between synthesis quality and clinical usability remains an open question. Experiments are limited to three brain MRI contrasts and two datasets, leaving generalization across anatomical regions, additional contrasts, and different field strengths or vendors to be tested. The authors mention distillation, fewer-step sampling, and parallel multi-plane generation as future directions, indicating remaining room to reduce sampling cost. In addition, several equations and specific values such as latent channel count and resolution, noise schedule endpoints, and learning rate are not fully rendered in the provided text, so reproduction would require the supplementary material.
