UniField merges 64mT-to-3T and 3T-to-7T brain MRI enhancement into one framework, gaining about 1.81 dB PSNR and 9.47% SSIM on average
Synopsis
The work proposes UniField, a unified multi-modality, multi-task MRI field-strength enhancement framework that consolidates T1, T2, and FLAIR modalities and cross-field tasks such as 64mT-to-3T and 3T-to-7T into a single model, embeds 3D structural priors by operating in the latent space of the pretrained video super-resolution model FlashVSR with LoRA fine-tuning, and introduces a Field-Aware Spectral Rectification Mechanism (FASRM) that adjusts low-, mid-, and high-frequency loss weights according to the physical properties of the source and target fields; it also organizes and publicly releases a paired multi-field MRI dataset from five institutions that is an order of magnitude larger than existing benchmarks, reporting average improvements of about 1.81 dB in PSNR and 9.
Interpretation
Unified modeling merges different modalities and field transitions into one model and yields consistent gains. Prior methods mostly train isolated networks for a specific modality or transition such as 64mT-to-3T; this work conditions on modality and transition to consolidate tasks, letting shared degradation patterns act as implicit data augmentation. The ablation table shows unified modalities outperform single-modality models (for 64mT-to-3T T1, PSNR rises from 18.41 to 18.71), adding unified tasks raises it further to 19.06, and adding FASRM reaches 19.75; 3T-to-7T shows the same trend.
FASRM mitigates the spectral bias and over-smoothing of flow-matching models through field-conditioned frequency-band weights. Unlike uniform or task-agnostic spectral penalties, FASRM adjusts low-, mid-, and high-frequency weights according to the physical properties of the source and target fields, for example relaxing high-frequency constraints in 64mT-to-3T to avoid blind hallucination and suppressing low-frequency weights in 3T-to-7T to prevent memorizing artifacts such as B1 inhomogeneity. Ablation shows that removing FASRM leads to poor high-frequency reconstruction, blurred images, and anatomically inconsistent structures, while adding it improves metrics; training settings report λfreq=0.1, α=1.0, and band weights of [0.2,0.5,0.3] for 64mT-to-3T and [0.1,0.3,0.6] for 3T-to-7T.
A pretrained video super-resolution model serves as a 3D structural prior, replacing slice-by-slice 2D processing and training from scratch. Conventional methods treat 3D volumes as independent 2D slices, discarding inter-slice continuous anatomy; this work performs flow matching in the FlashVSR latent space, freezing the encoder and decoder and fine-tuning the main body with LoRA and sparse attention. The method section states that the encoder and decoder use frozen FlashVSR weights and the main body is fine-tuned with LoRA, and it gives the ODE formulation used at inference; in the comparison tables the FlashVSR baseline is below UniField on most metrics.
A multi-center, multi-task paired multi-field MRI dataset is organized and publicly released, an order of magnitude larger than existing benchmarks. Prior studies often rely on only a few dozen strictly paired cases; this work aggregates data from five institutions and registers previously unaligned pairs, covering both 64mT-to-3T and 3T-to-7T tasks. The dataset table lists ULF-EnC (50 cases), Leiden (10), KCL (23), UNC (10), and BNU (20), with a uniform 8:2 train-test split and preprocessing including MONAI intensity normalization, 1mm z-resampling, and 256×256×160 resizing.
Perspective
The result targets brain MRI field-strength enhancement and applies to research and clinical workflows that have paired low/high-field or clinical/ultra-high-field data and want a single model covering multiple modalities and transitions; the authors propose extending the paradigm to more anatomical organs, magnetic field strengths, and disease types. The dataset and code are released, which supports reproduction and extension under the same task settings.
Test sets are small, with only 2 to 4 test cases per center in 3T-to-7T, so readers will wonder whether the gains hold on larger, more multi-center data; FASRM band weights are set per task by hand, and how to choose them for other field combinations remains an open question; moreover, evaluation relies mainly on image metrics such as PSNR, SSIM, NRMSE, and LPIPS, and the relationship between these metrics and clinical diagnostic usability still needs further observation.
