VRPTR predicts individual language activation maps from resting-state fMRI and provides calibrated uncertainty estimates
Synopsis
The study introduces VRPTR, a three-dimensional encoder-decoder combining a compressed Transformer bottleneck, variational latent sampling, and multiscale skip connections, trained on 360 healthy adults from the WU-Minn Human Connectome Project and evaluated on 40 held-out participants for the story-versus-math language contrast, achieving mean voxel-map Pearson r=0.642 and Dice AUC=0.519, exceeding compact volumetric BrainSurfCNN-like and SWIFUN-like comparators by Δr=0.0376 and 0.0335 respectively, while raw 95% intervals covered only 4.7% of observed values and five-fold calibration within the held-out cohort raised coverage to 94.8%.
Interpretation
VRPTR predicts individual activation maps for the story-versus-math language contrast in held-out participants, with mean voxel-map Pearson r=0.642 and Dice AUC=0.519. It exceeds compact volumetric BrainSurfCNN-like and SWIFUN-like comparators by Δr=0.0376 and 0.0335 respectively, indicating higher map correlation under a comparably compact setting. Training used 360 healthy adults from the WU-Minn Human Connectome Project and evaluation used 40 held-out participants from the same dataset, with voxel-map Pearson correlation and Dice AUC as metrics.
Against a template fixed from the 360 training maps, VRPTR improved correlation but not mean absolute or root mean squared error; the template itself yielded residual correlation 0.364. This contrast separates individualized prediction from a shared population pattern, indicating that improved correlation does not equal across-the-board voxelwise error improvement. A template fixed from the 360 training maps served as the baseline, with residual correlation 0.364 reported alongside comparisons of correlation, mean absolute error, and root mean squared error.
A matched factorial ablation associated Transformer processing with improved map correlation (Δr=0.0168), while the variational main effect did not survive multiple-comparison correction; component swaps indicated that subject-specific activation patterns depended more on multiscale skip features than on the bottleneck representation. It separates the role of Transformer processing from subject-specific spatial conditioning and from the bottleneck representation, indicating that different components serve different functions. Two experimental designs were used, a matched factorial ablation and component swaps, with multiple-comparison correction reported.
Latent sampling produced only a modest association between posterior standard deviation and voxelwise absolute error (mean Spearman ρ=0.091); raw 95% intervals covered only 4.7% of observed values, five-fold calibration within the held-out cohort raised coverage to 94.8%, spatially varying uncertainty modestly improved probabilistic scores over a constant-width model, and uncertainty increased under input degradation. It shows that relative uncertainty estimation requires calibration to reach nominal coverage and demonstrates that uncertainty responds to input degradation. Mean Spearman ρ, raw and calibrated interval coverage, probabilistic score comparisons, and uncertainty changes under input degradation are reported.
Perspective
The work targets individualized prediction of story-versus-math language contrast activation maps from resting-state fMRI in healthy adults, with training and evaluation within the same dataset; it applies to research settings that need individualized language-map estimates and a relative uncertainty scale, and its calibrated intervals and spatially varying uncertainty scores offer a starting point for testing in independent samples and clinical cohorts.
The abstract does not report generalization across datasets or contrasts, nor clinical cohort results; the calibrated 94.8% coverage comes from five-fold calibration within the held-out cohort, so independent validation remains an open question; moreover, the finding that raw 95% intervals covered only 4.7% of observed values suggests the uncertainty scale is sensitive to the calibration procedure, and the mechanism awaits the full text.
