Skip to main content
Back to timeline
Frontiers in OncologySource publication:

Bibliometric analysis of 608 AI breast cancer imaging studies finds diagnosis at 71.22% and explainable AI with a 1.00 burst ratio as the hottest frontier

Synopsis

Using Scopus and Web of Science and a PRISMA workflow that narrowed 4,831 records to 608 peer-reviewed journal articles and reviews from 2020 to 2026, this study applied Bibliometrix and VOSviewer for a task-aware, methodology-centric bibliometric and thematic analysis, finding that diagnosis accounts for 71.22% of studies, mammography for 42.11%, explainable AI shows the strongest burst ratio at 1.00, while treatment-response prediction covers only 2.30% and only about 15-25% of studies explicitly report hyperparameter tuning strategies.

Source-provided article image: Artificial intelligence in breast cancer research: a systematic review and bibliometric analysis of emerging trends and future directions
FIGURE 1

FIGURE 1 PRISMA workflow diagram.

· Page 5

Interpretation

Task-level mapping: diagnosis 433 papers (71.22%, average 24.74 citations), prognosis 167 (27.47%, 20.50), treatment-response 14 (2.30%, 3.25), and multi-task prognosis plus treatment 6 (0.99%). Prior bibliometric work often treated the field as one unified domain; this study separates diagnosis, prognosis, and treatment-response prediction and reports citation averages for each. Counts and percentages from the final corpus of 608 publications drawn from Scopus and Web of Science, selected through a PRISMA four-stage process (4,831 records, 2,744 duplicates removed, 1,259 remaining after title and abstract screening, 608 included after full-text eligibility).

Modalities and datasets are highly concentrated: mammography appears in 256 studies (42.11%), roughly 2.5 times ultrasound (100 studies, 16.45%), histopathology in 87 (14.31%), and MRI in only 21 (3.45%); among 168 explicit dataset mentions, MIAS and BreakHis each appear 37 times (22.02%), INbreast 33 (19.64%), CBIS-DDSM 32 (19.05%), and BUSI 29 (17.26%), with mammography datasets together accounting for about 60.7%. By placing modality distribution alongside dataset usage frequency, the study makes the structural pattern of mammography dominance and underuse of genomic and population-scale data quantifiable. Modality counts from the 608-paper corpus and frequency statistics over 168 explicit dataset mentions, accompanied by a table listing dataset sources, approximate sample sizes, analytical software, and access URLs.

Methodological reporting is uneven: among explainable AI methods, general explainability accounts for 31.91% (33 papers), Grad-CAM 14.64% (19), CAM 11.84% (12), saliency maps 5.92% (8), SHAP 3.95% (7), LIME 2.96% (3), attention visualization 1.97% (3), and Integrated Gradients 0.99% (1); for hyperparameter optimization, genetic algorithms lead at 1.97% (12 papers), followed by particle swarm optimization 1.32% (8), Bayesian optimization 0.33% (2), and grid search 0.16% (1), with only about 15-25% of studies explicitly describing tuning methods. Prior bibliometric studies rarely examined explainable AI adoption and hyperparameter optimization reporting as dedicated dimensions; this work quantifies both. Explicit mention counts and percentages across the 608 references, with the authors noting that extraction relied on titles, abstracts, and keywords and therefore carries reporting bias.

Hotspot shift and performance gradient: explainable AI shows a burst ratio of 1.00, multimodal about 0.92, attention 0.94, and Transformer and ViT about 0.89, while deep learning at 0.73 and CNN at 0.76 are classified as mature or stable core; diagnostic models typically report 90-99% accuracy and 0.85-0.99 AUC, prognostic models a C-index of 0.65-0.90, and treatment-response models an AUC of 0.65-0.85 with inconsistent pCR prediction. Using burst ratios rather than raw keyword frequency quantifies a paradigm shift from accuracy-focused models toward interpretable, multimodal, and clinically robust systems, and provides a cross-task performance gradient. Burst ratios computed from keyword frequency and 2023-2026 recent shares, plus performance ranges summarized from the corpus; the authors caution that high accuracies largely come from retrospective, single-center, or public datasets without independent external validation.

Perspective

This analysis covers peer-reviewed English journal articles and reviews indexed in Scopus and Web of Science from 2020 to 2026, focused on imaging modalities such as mammography, ultrasound, MRI, and histopathology, so its conclusions apply to research planning, dataset and tool selection, and reporting standards in imaging-driven breast cancer AI. For researchers, clinical teams, and tool developers who want to understand task distribution, modality concentration, and current explainable AI and hyperparameter optimization practices, the paper provides directly citable counts and tabulated resource lists. The authors propose a clinical translation pathway with three phases: internal validation, prospective and external evaluation, and clinical validation with impact assessment, alongside cross-cutting requirements including data quality, reproducibility, transparency, explainability, fairness, privacy, regulatory compliance, and post-deployment monitoring, which can serve as a reference for future study design and evaluation frameworks.

The loaded text is incomplete and does not include the contents of Figures 1 through 16, so the PRISMA workflow diagram, keyword co-occurrence network, annual output and citation growth curves, journal and author productivity distributions, burst map, and translation pathway figures can only be understood from the body text and their details cannot be verified. The authors also note scope limits: the corpus comes only from selected databases and published indexed studies, potentially missing other repositories, preprints, and non-English work; quantitative extraction relies on titles, abstracts, and keywords, so datasets, frameworks, and tuning strategies may not be consistently recorded; bibliometric analysis cannot determine whether models were validated in real clinical environments; many studies rely on a small set of benchmark datasets such as CBIS-DDSM, MIAS, and INbreast, limiting representativeness of patient populations and imaging conditions; and citation and keyword trends reflect research activity rather than scientific quality or clinical impact. In addition, screening was performed by a single reviewer without prospective registration, and the explainable AI and hyperparameter optimization statistics rest on explicit mentions, which may undercount actual use. Readers may continue to watch: multi-institutional external validation of transformer-based models, the emergence of explainability evaluation standards, data harmonization approaches for multimodal and federated learning, and whether treatment-response prediction, at only 2.30% of studies, develops comparable benchmarks.

Sources