Review of 86 MRI-based autism AI studies: single-site accuracy reaches 99%, but leave-one-site-out validation lands near 67%
Synopsis
Following PRISMA, this review searched Scopus, Web of Science, and PubMed for 2015–2026 studies and included 86 Q1 journal articles, systematically mapping MRI-based machine learning and deep learning studies for autism spectrum disorder (ASD) classification across datasets, preprocessing pipelines, brain atlases, model architectures, and validation strategies; it finds fMRI is the most used modality with graph neural networks and transformers as dominant trends, and shows reported accuracy depends heavily on evaluation setup—single-site studies reach 87.4%–99.39%, the full ABIDE cohort sits around 70%–75%, and the strictest leave-one-site-out testing lands near 67%.
Interpretation
The review synthesizes 86 MRI-based AI studies for ASD classification (2015–2026) within a unified structure covering datasets, preprocessing pipelines, brain atlas selection, model architectures, performance trends, explainability, and multimodal fusion. Prior reviews often span questionnaires, EEG, eye-tracking, facial features, speech, and wearable sensors, offering comparatively limited technical analysis of MRI-specific datasets, preprocessing pipelines, parcellation strategies, modelling architectures, and validation methods; this review restricts scope strictly to MRI and provides finer technical detail. It follows PRISMA, searching Scopus, Web of Science, and PubMed with a supplementary Google Scholar search: 716 records initially, 712 after deduplication, 562 excluded at title/abstract screening, 150 full texts reviewed, and 86 studies included; only Q1 journals were kept, filtered by five criteria scored out of 10 with a threshold of 7.
The review identifies the field's dominant technical trajectory: fMRI is the most frequently used modality, while graph neural networks, transformers, attention mechanisms, and multimodal fusion frameworks are the leading methodological trends. Compared with earlier broad multimodal overviews, this review characterizes the methodological shift within MRI itself and explains why graph models map naturally onto the brain-parcellation graph, where nodes are regions of interest and edges encode functional connectivity. Based on categorical counts of included studies and a keyword co-occurrence network visualization (VOSviewer), where 'deep learning' clusters with 'graph neural networks', 'transfer learning', 'convolution network', 'autoencoder', and related terms.
The review finds reported performance is driven mainly by evaluation setup rather than the model itself: single-site studies (mostly NYU) report accuracy from 87.4% on sMRI to 99.39% on fMRI, the full ABIDE cohort drops to 70%–75%, and the strictest leave-one-site-out testing lands near 67%. This contrast attributes a gap of almost thirty percentage points to experimental design rather than model superiority, and notes that accuracies of 97%–100% often come from slice-level splitting (where slices from the same person enter both training and test sets), small balanced subsets, or validation that is not genuinely site-independent. The review aggregates reported values across differing validation protocols and notes most studies report accuracy alone, rarely providing threshold-independent measures; it also notes that the actual number of scans used under 'ABIDE I' ranges from about 870 to over 1100 across papers, making numbers not directly comparable.
The review reports that multimodal fusion does not reliably outperform single-modality models, and that the brain regions highlighted across studies disagree, making them better treated as hypotheses than confirmed biomarkers. Contrary to the common expectation that fusion should improve performance, the collected fusion accuracies vary widely (64.9% up to 87%–97%) with no reliable advantage over the best fMRI-only models; larger gains tend to appear when phenotypic variables such as age, sex, site, and IQ are added, which may reflect site or demographic cues rather than structural–functional complementarity. The review compares reported results across sMRI, fMRI, and fusion studies and lists the discriminative regions singled out by different studies (left amygdala and right hippocampus; paracingulate, supramarginal, and middle temporal gyri; calcarine sulcus and cuneus; thalamic and rectus regions), noting limited agreement.
Perspective
This review is aimed at researchers in MRI neuroimaging and computational ASD diagnosis, and at methodologically oriented readers concerned with medical AI evaluation, and it applies to mapping the state of sMRI-, fMRI-, and fusion-based ASD classification studies published in Q1 journals between 2015 and 2026. Its conclusions apply to understanding under which evaluation settings the current evidence holds: single-site results reflect performance under specific acquisition and population conditions, whereas full-ABIDE and leave-one-site-out results are closer to cross-site deployment scenarios. The priorities it proposes—multimodal integration, explainable AI, large harmonized datasets, site-held-out validation, and uncertainty estimation—offer a framework that follow-up study designs can adopt directly.
Readers should note that the review's conclusions rest on synthesis of published reports, and included studies differ substantially in subject subsets, sites, data-split levels, and validation strategies, so accuracy values are not directly comparable across studies. The brain regions reported differ across studies, and whether they remain stable across sites and preprocessing pipelines is still an open question. Whether multimodal fusion delivers genuine structural–functional complementarity or partly reflects site and demographic cues is flagged by the review as an important unresolved question. In addition, the loaded text does not expand the specific values in appendix Tables A1–A3 or the raw data behind some figures (such as Figures 2, 4, 8, and 9); verifying each study's accuracy, sensitivity, and specificity item by item would require consulting the publicly available underlying data files.
