Skip to main content
Back to timeline
Journal of medical Internet researchSource publication:

Fine-Tuning, Retrieval-Augmented Generation, and Hybrid Adaptation of Language Models for Clinical Decision-Making in Health Care: Systematic Review

Synopsis

Following PRISMA 2020 guidelines and searching PubMed/MEDLINE, Scopus, and Web of Science (January 2018 through May 2026), this systematic review screened 1890 records and included 35 studies published between 2024 and 2026, grouping them into fine-tuning or parameter-efficient fine-tuning (7/35, 20%), retrieval-augmented generation (17/35, 48.6%), and hybrid approaches (11/35, 31.4%) for descriptive synthesis, finding that fine-tuning performed strongly on narrow task-specific applications (area under the receiver operating characteristic curve up to 0.912 for cancer detection and area under curve 0.892 for major depressive disorder prediction), that retrieval-augmented generation improved guideline adherence and diagnostic accuracy (from 71.1% to 92.1% and from 78.9% to 94.

Source-provided article image: Fine-Tuning, Retrieval-Augmented Generation, and Hybrid Adaptation of Language Models for Clinical Decision-Making in Health Care: Systematic Review.
Figure 1. ·

PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) flow diagram of study identification, screening, eligibility assessment, and the inclusion process for the systematic review.

PubMed

Interpretation

The review systematically classifies clinical language model adaptation studies by enhancement strategy, mapping the evidence distribution across fine-tuning, retrieval-augmented generation, and hybrid approaches. Comparative evidence on clinical language model adaptation was previously unclear; this work applies a PRISMA 2020 process across three databases and includes 35 studies, specifying the proportions and domain coverage of the three strategy types. Systematic review methodology, screening 1890 records and including 35 studies spanning oncology, neurology, radiology, mental health, cardiology, ophthalmology, and surgical care.

Fine-tuning performed strongly on task-specific applications, retrieval-augmented generation improved guideline adherence and diagnostic accuracy, and hybrid systems generally achieved the strongest performance in complex clinical workflows. The work aligns strategies with task types, proposing fine-tuning for narrow classification, retrieval-augmented generation for guideline-grounded reasoning, and hybrid approaches for complex multimodal tasks. Specific performance values are reported, including area under the receiver operating characteristic curve up to 0.912 for cancer detection, area under curve 0.892 for major depressive disorder prediction, guideline adherence rising from 71.1% to 92.1%, diagnostic accuracy rising from 78.9% to 94.7%, and external validation accuracies exceeding 90% in stroke triage, dermatology, multimodal imaging, and oncology; however, retrieval-augmented generation benefits were inconsistent across larger reasoning-capable models.

The review used PROBAST+AI to assess risk of bias and indicates that the current evidence base remains largely retrospective or benchmark-based. Alongside synthesizing performance evidence, it provides a methodological quality profile: 25 studies at high risk of bias, 9 at unclear risk, and only 1 at low risk. Common concerns included inadequate external validation, lack of calibration assessment, nonrepresentative participant selection, and insufficient reporting of analytical methods; the authors accordingly call for prospective studies with external validation, calibration, and standardized safety reporting.

Perspective

This review is intended for researchers and clinical informatics teams who need to choose posttraining adaptation strategies for clinical diagnosis and decision-support tasks, and its conclusions apply at the level of matching strategy to task type: fine-tuning may be considered for narrow classification, retrieval-augmented generation for guideline-grounded reasoning, and hybrid approaches for complex multimodal tasks. The included evidence comes from 35 studies published between 2024 and 2026 across multiple specialties, so the scope of the conclusions is bounded by the settings of those published studies.

Readers should still watch that retrieval-augmented generation benefits were inconsistent across larger reasoning-capable models, and the specific conditions for that difference remain to be clarified; the included studies are largely retrospective or benchmark-based, with limited standardization of external validation, calibration assessment, and safety reporting, so the transition from these results to prospective clinical application still requires more evidence; in addition, the available content here is at the abstract level and does not include figures or appendix details, so readers who need to verify individual study designs, samples, and statistical methods should consult the full text.

Sources