Skip to main content
Back to timeline

The Use of Ambient Dictation Artificial Intelligence in Clinical Spaces in Surgery: A Scoping Review

Synopsis

Following PRISMA-ScR guidance, this scoping review searched EMBASE, PubMed (MEDLINE), CINAHL, SCOPUS, and Cochrane plus citation searching, and from 252 records included 12 studies, of which only 3 provided original data (one urology primary study, one urology expert commentary, and one narrative review with hand surgery survey data), with the remainder mostly narrative reviews whose cited evidence largely came from nonsurgical specialties, indicating that empirical research on ambient dictation AI in surgical specialties remains sparse.

Source-provided article image: The Use of Ambient Dictation Artificial Intelligence in Clinical Spaces in Surgery: A Scoping Review.
Figure 1 ·

PRISMA‐style flow diagram illustrating the study selection process for the review. Preferred Reporting Items for Systematic Reviews and Meta‐Analyses (PRISMA) flow diagram demonstrating the identification, screening, eligibility assessment, and inclusion of studies evaluating the use of ambient dictation AI tools in surgical clinical environments. Database and citation searches identified 163 unique records. The number of studies included and excluded at each stage of the review process is displayed, resulting in 12 studies meeting eligibility criteria for inclusion in the final review.

PubMed

Interpretation

The review maps the literature landscape on ambient dictation AI ("AI scribes") in surgical clinical environments, including 12 studies: 7 narrative reviews, 1 narrative review with a cross-sectional survey, 1 systematic review, 1 scoping review, 1 primary research investigation, and 1 expert commentary. The authors describe it as the first review to specifically highlight the lack of primary research on AI scribing in the surgical environment, consolidating previously scattered subspecialty discussions into one evidence map. Based on PRISMA-ScR procedures, five databases plus citation searching, independent screening by three team members with consensus resolution, searched on June 6, 2025, covering database inception to that date.

Only 3 studies provided original data: a urology primary study (survey of 20 urologists, 75% identified documentation as a major contributor to burnout, 90% expressed interest in AI scribes; simulated urologic encounters transcribed by 5 AI scribes, with Nabla scoring highest at 68% composite and lowest critical error rate at 28%), a urology expert commentary (reporting a 25% increase in same-day note completion after adoption, while noting AI scribes require editing and do not document physical exam findings), and a narrative review with hand surgery survey data (90% of surveyed clinicians used the tool for discharge summaries, referral letters, and consultation notes, but limited EHR integration and user hesitancy restricted use for surgical planning). These are the only evidence points in the review carrying original data, moving the claim that "AI scribes work in surgery" from opinion toward checkable figures. The primary study is a survey of 20 physicians plus simulated-scenario evaluation of 5 tools; the expert commentary has no review structure; the hand surgery data come from a cross-sectional survey with limited sample and design.

All included review articles expressed positive views on implementing AI scribing in surgical subspecialties, yet 8 of them relied on AI scribing evidence from nonsurgical settings, or on aggregate surgical and nonsurgical data without discrete surgical data. The review separates the surface impression that "surgical literature supports AI scribes" from the underlying sources of that support, showing that supportive citations often do not land in surgical settings. Each review's cited AI scribing sources were checked for surgical data and marked "No" with reasons in the table, constituting an explicit audit of the citation chain.

Included studies were mostly from the United States (n = 9), with additional studies from Canada, Singapore, and the United Kingdom (one each); the surgical fields most represented were orthopedics and urology (four studies each), followed by plastic surgery, pediatric surgery, academic surgery, and generalized surgical applications (one each); publication counts trended upward, with three studies in 2023, four in 2024, and five between January and July 2025. Provides the temporal trend and geographic and subspecialty distribution of the topic, helping identify growth areas and gaps. Descriptive counts based on country, year, and specialty classification of included studies.

Perspective

The review defines the applicable boundary of current evidence: supportive data come mainly from nonsurgical specialties, so conclusions for surgical specialties are inference rather than validation; the authors specifically note that otolaryngology clinics integrate endoscopy and in-office procedures into routine visits, requiring precise anatomic descriptors, laterality, grading scales, and structured findings, and that clinic environments involve background noise from suction devices and endoscopic equipment while patients may present with dysphonia, aphonia, airway compromise, or altered resonance, factors whose effect on speech recognition has not been systematically evaluated in otolaryngology. The review offers a starting point for subsequent prospective, specialty-informed evaluation in surgical specialties, relevant to researchers, department leaders, and informatics teams focused on clinical documentation efficiency and AI implementation.

Readers should still note: included studies varied widely in design and surgical subspecialty focus, preventing pooled analysis; most reviews cited nonsurgical sources, so the judgment that AI scribes work in surgery awaits prospective validation; whether otolaryngology-specific documentation of endoscopy and in-office procedures, noisy environments, and speech clarity affect transcription accuracy is explicitly listed as not yet systematically evaluated; additionally, the review protocol was not registered online and the search closed on June 6, 2025, so primary studies published afterward could change the evidence picture.

Sources