Skip to main content
Back to timeline
medRxivSource publication:

AI-Based Synthetic Data in Biomedicine: A Decade of Growth and a Persistent Translation Gap

Synopsis

This study conducted a systematic mapping and bibliometric analysis of 4,143 publications from 2015 to 2025 on AI-generated synthetic data in biomedicine, combining expert annotation with LLM-assisted classification across data modality, medical domain, paper type, deployment status, and research stance, finding continuous growth in publication volume, 77.8% of papers strongly supportive with critical work below 1%, medical imaging dominating the corpus, highly cited primary research concentrated in molecular and pharmaceutical applications, and only 27 publications reporting operational use, thereby revealing a gap between methodological growth and deployment.

AI-generated editorial illustration: AI-Based Synthetic Data in Biomedicine: A Decade of Growth and a Persistent Translation Gap

Interpretation

The study constructed a systematic mapping and bibliometric landscape of 4,143 publications spanning 2015 to 2025, classified across data modality, medical domain, paper type, deployment status, and research stance. Relative to prior reviews focused on a single method or domain, this work quantitatively characterizes the overall landscape of biomedical AI synthetic data research across a decade and multiple annotation dimensions. Based on a systematic mapping of 4,143 publications, combining expert annotation with LLM-assisted classification, with sample size and classification dimensions explicitly stated in the text.

Publication volume grew continuously, and research stance was heavily skewed toward support: 77.8% of papers were strongly supportive while critical work remained below 1%. This result quantifies the distribution of research stances in the field, showing that supportive literature is overwhelmingly dominant while critical voices are extremely rare. Derived from annotation of 4,143 publications by research stance, with the specific proportions of 77.8% and below 1% given in the text.

Medical imaging dominated the corpus, consistent with well-characterized transformation-group invariances supporting data augmentation and generative modeling; highly cited primary research concentrated disproportionately in molecular and pharmaceutical applications, where SE(3)-equivariant architectures and structure-prediction models accelerated generative approaches. This finding links the distribution of research hotspots to the degree to which invariance structures can be characterized, pointing to differences in the suitability of different modalities for generative modeling. Based on domain distribution statistics of the corpus and the concentration of highly cited primary research, with transformation-group invariances and SE(3)-equivariant architectures offered as explanatory cues in the text.

Only 27 publications reported operational use, and omics and tabular clinical data remained underrepresented due to lacking well-characterized invariance structures, revealing an overall gap between methodological growth and deployment. This result incorporates deployment status into the bibliometric analysis, explicitly identifying a clear gap between methodological research and practical application, and pointing to omics and tabular clinical data as underrepresented directions. Based on 27 records of operational use obtained from deployment status annotation, along with representation statistics for omics and tabular clinical data in the corpus.

Perspective

This study is applicable to understanding the overall distribution and deployment status of biomedical AI synthetic data research from 2015 to 2025, and its conclusions are aimed at researchers and funders concerned with the field's research landscape, evaluation standards, and deployment reporting; it characterizes patterns at the literature level rather than validating the effectiveness of any specific method in a particular clinical or experimental setting.

This paper is presented in abstract form and does not elaborate on the specific annotation criteria for each classification dimension, the consistency between expert annotation and LLM-assisted classification, or the specific composition of the 27 publications reporting operational use; readers seeking to judge the deployment maturity of a particular modality or domain would still need to consult the figures and classification details in the full paper.

Sources