Skip to main content
Back to timeline
arXivSource publication:

ICHOR self-supervised pretraining on 11,405 ASL CBF scans outperforms structural-MRI pretrained baselines across four downstream tasks

Synopsis

The study introduces ICHOR, a self-supervised pretraining framework for ASL cerebral blood flow (CBF) maps based on 3D masked autoencoders (a ViT-Base encoder with a light decoder), pretrained on 11,405 ASL CBF scans from 14 studies spanning multiple sites and protocols, and evaluated on three diagnostic classification tasks plus one ASL CBF map quality prediction regression task, where it outperformed structural-MRI pretrained baselines BrainIAC, BrainSegFounder, and MedicalNet overall across all four tasks.

Source-provided article image: ICHOR: A Robust Representation Learning Approach for ASL CBF Maps with Self-supervised Masked Autoencoders
Fig. 1

Fig. 1. Overview of the presented study.

· Page 3

Interpretation

ICHOR provides a dedicated pretrained encoder for ASL-derived CBF maps, addressing the prior gap in which neuroimaging pretrained models focused mainly on structural or multi-contrast anatomical MRI and did not incorporate physiological modalities such as ASL. Prior large-scale brain MRI pretrained models such as BrainIAC and BrainSegFounder focused on structural imaging, leaving ASL CBF without a dedicated pretrained backbone; ICHOR performs masked image modeling on ASL CBF maps with a 3D masked autoencoder to learn transferable representations. The paper implements a ViT-Base encoder (12 transformer blocks, 12 attention heads, MLP dimension 3072) with a 4-block light decoder; inputs are CBF volumes registered to MNI space and resampled to 96×96×96, split into 512 non-overlapping 12×12×12 3D patches with masking ratio ρ=0.5, and the MSE loss is computed only over masked patches.

The study curated one of the largest ASL datasets to date for pretraining across heterogeneous acquisition settings. The dataset comprises 11,405 ASL CBF scans from 14 studies, spanning multiple sites, protocols (PASL/PCASL, 2D/3D, single/multi post-labeling delay), and heterogeneous populations, providing a large multi-site resource that was previously lacking for ASL pretraining. Table A1 lists per-dataset sample sizes (e.g., UKB 5,383; PNC 1,397; ADNI 966), totaling 11,405 scans with ages 0–100 and 54.7% female; all scans underwent automated quality control and standardized preprocessing.

Across four downstream tasks, ICHOR outperformed structural-MRI pretrained neuroimaging self-supervised baselines overall. Compared with BrainIAC, BrainSegFounder, and MedicalNet, ICHOR achieved the best overall performance on CU Aβ− vs CI Aβ+ classification (AUC 78.93%), HOA vs SVD classification (AUC 73.33%), AD vs bvFTD classification (AUC 100.00%), and ASL quality prediction (R2 86.29%, PC 93.11%). Evaluation used nested 5-fold cross-validation with one fold held out as an independent test set per iteration and an inner 80%/20% stratified train/validation split; downstream adaptation used LoRA (rank r=8, αLoRA=16, dropout p=0.2), with results reported as the mean across the five outer folds.

The masking ratio has a non-monotonic effect on downstream performance, with ρ=0.5 achieving the best downstream results on most metrics. The study compared ρ∈{0.25, 0.5, 0.75} and found that increasing ρ yields visually sharper reconstructions, but downstream performance does not improve monotonically with reconstruction quality; on CU Aβ− vs CI Aβ+, ρ=0.5 reached AUC 78.93% versus 76.74% for ρ=0.25 and 71.90% for ρ=0.75. This comparison used the same downstream training protocol, with results in Table 2; the paper interprets it as a tradeoff in which low masking makes the reconstruction objective less informative while high masking provides insufficient visible context.

Perspective

The work targets analysis of ASL-derived CBF maps and is intended for research settings and the growing use of ASL in clinical MRI protocols; its pretraining corpus spans multiple sites, protocols (PASL/PCASL, 2D/3D, single/multi post-labeling delay), and heterogeneous populations, while downstream evaluation uses institutional cohorts independent of the pretraining data. ICHOR is designed as a general-purpose encoder adaptable via LoRA to classification and regression tasks, and the authors propose extending it to quality enhancement tasks such as denoising and artifact correction, longitudinal disease progression and treatment response modeling, and expanding the pretraining corpus to strengthen generalization across populations, protocols, and scanner platforms.

Downstream cohorts are small (e.g., HOA n=20, SVD n=15, AD n=27, bvFTD n=36), so some task metrics may vary across folds; the randomly initialized baseline achieved higher recall and F1 in some settings, which the paper partly attributes to overfitting in small cohorts and to LoRA limiting adaptation under strong modality mismatch, an explanation that remains speculative. The relationship between masking ratio and downstream performance was quantified mainly on a single classification task, and whether it generalizes to other tasks remains to be seen. In addition, pretrained weights and code are described as to be made publicly available, so their actual availability and reproduction details depend on the release.

Sources