Using OCD as its example, a review argues that big data in psychiatry must separate etiology from pathophysiology and stay hypothesis-driven to answer a specific disorder's three core questions
Synopsis
This review, using obsessive-compulsive disorder (OCD) as its example, surveys the advances and constraints of big data approaches to three questions—what causes the disorder, which treatments work best for whom, and how to widen access to evidence-based care—reporting that large-scale pooled data such as ENIGMA-OCD show some previously reported effects to be small but reproducible while others do not replicate at scale or appear to be medication effects, that the Global OCD Study harmonized methods across five international sites and recruited 250 unmedicated adults with OCD and 250 matched healthy volunteers and found relatively few case-control brain differences but more informative associations between specific brain and cognitive measures and OC clinical profiles, and that big data ar
Interpretation
The review proposes organizing big data research on OCD within a clear causal framework that distinguishes etiology from pathophysiology. Rather than listing genetic, environmental, imaging and cognitive findings side by side, it explicitly frames brain alterations as "intermediate phenotypes"—mechanisms shaped by upstream etiologic processes and likely altered by disease duration and treatment—rather than causes in their own right. A conceptual framing argument, drawing on the text's review of genome-wide association studies, rare variant analyses, transcriptomic, proteomic and metabolomic processes, and environmental factors (perinatal complications, infection and inflammation, trauma, toxins, cultural context); it presents no new empirical data.
Large-scale collaborative data are recalibrating the field's expectations about effect sizes: pooling neuroimaging data worldwide, ENIGMA-OCD found some previously reported effects to be small but reproducible, while others do not replicate at scale or appear to be medication effects. Earlier studies often recruited small samples with heterogeneous methods and proved hard to replicate; cross-national meta- and mega-analyses are the first to separate "small but reproducible" effects from non-replicating ones at sufficient scale. Based on ENIGMA-OCD meta- and mega-analyses of structural and functional neuroimaging; the text also notes its data are retrospectively pooled with clinical and imaging protocols not harmonized across sites, introducing substantial between-site variance that limits finer-grained brain-behavior analyses and subgroup identification.
The Global OCD Study harmonized clinical, neurocognitive and neuroimaging methods across five international sites, recruiting 250 unmedicated adults with OCD and 250 matched healthy volunteers; preliminary findings suggest relatively few case-control brain differences but more informative associations between specific brain and cognitive measures and OC clinical profiles, supporting a brain-behavior rather than brain-diagnosis framework. Relative to the retrospectively pooled ENIGMA-OCD, it achieves deeper phenotyping with harmonized protocols and unmedicated participants, and shifts the target from diagnosis to OC clinical profiles. 250 matched participants, multimodal neuroimaging and cognitive tasks probing cortico-striato-thalamo-cortical circuits; the text states its sample size limits more flexible data-driven approaches, motivating the new MEGA-OCD initiative (R01MH138569) to combine large-scale samples with rich phenotyping via multimodal fusion, normative modeling and symptom-anchored clustering.
On treatment and access, big data are being used to identify predictors, moderators and mediators of treatment response and to document and address gaps in care. First-line interventions (SRIs and CBT with exposure and response prevention) enable up to half of patients to achieve minimal symptoms, with the remainder selected largely by trial-and-error; an ongoing trial in China (NCT04539951) enrolling 1,600 treatment-naive individuals with OCD to determine optimal pharmacotherapy after first-line SRI is the largest OCD trial to date and is powered to detect clinically meaningful moderators and mediators. Grounded in trial design and prior literature; the text also cautions that larger trials are not inherently better, since with very large samples small differences can reach statistical significance without clinical relevance, and interventions benefiting only specific subgroups may fail to replicate if heterogeneity is not explicitly anticipated in the design—so treatment research should be hypothesis-driven and clinically anchored.
Perspective
This is a methodological and agenda-setting review using OCD as its example, aimed at psychiatric researchers, trial designers and service-system decision-makers, for planning how to integrate genetic, environmental, imaging, cognitive and clinical data within a clear causal framework and how to use digital platforms and AI tools to widen access to evidence-based care. It offers directions and principles rather than a ready-to-apply analytic pipeline or efficacy conclusions.
The loaded text is incomplete in scope and lacks figures and supplementary material, so the specific effect sizes, statistics and subgroup results of ENIGMA-OCD and the Global OCD Study cannot be verified here and are reported only as the text states them. The Global OCD Study findings are described as "preliminary," and MEGA-OCD and the Chinese trial are ongoing, so their conclusions await publication; the real-world effects of AI digital platforms and clinician-training tools, and whether they might exacerbate existing inequities, are also flagged in the text as open questions requiring careful evaluation.
