Skip to main content
Back to timeline
medRxivSource publication:

An AI Competency Framework for Emergency Medicine: A Multiphase Consensus Process

Synopsis

Using a Nominal Group Technique session (May 2025, n=6) and a two-round modified Delphi process (April–May 2026, expert panel n=13 across 12 academic medical centers, with 77% Round 1 and 100% Round 2 response rates), this study developed the first specialty-specific AI competency framework for a United States medical specialty, comprising 5 themes (Communicating about AI, Understanding appropriate use cases, Interacting with AI, AI risk management, Cognitive impacts of AI), 5 derived competencies (one-to-one theme-to-competency mapping endorsed by 12 of 13 panelists), 19 subthemes, and 10 retained clinical scenarios, with all themes, derived competencies, and individually rated subthemes meeting pre-specified consensus thresholds and 9 of 10 scenarios reaching consensus while 1 was retain

AI-generated editorial illustration: An AI Competency Framework for Emergency Medicine: A Multiphase Consensus Process

Interpretation

The study produced a clinical-scenario-anchored AI competency framework for emergency medicine with 5 themes, 5 derived competencies, 19 subthemes, and 10 retained clinical scenarios, organized along an arc of the emergency physician's relationship with AI: communicating about it, judging when to use it, working with it in real time, managing what it gets wrong, and protecting one's own reasoning. Previously, ACGME Program Requirements contained no mention of AI, the AAMC cross-continuum competencies were intentionally specialty-agnostic, and a six-domain framework for health care professionals offered only a profession-wide baseline; no specialty had translated such baselines into its own clinical context, making this the first specialty-specific AI competency framework for a United States medical specialty. A Nominal Group Technique plus a two-round modified Delphi with a 13-member panel across 12 academic medical centers, Round 1 response 10/13 (77%) and Round 2 13/13 (100%), with a priori thresholds of ≥70% rating ≥4 with IQR ≤1 for scenarios, ≥75% rating adequate for themes, and ≥70% rating ≥4 with IQR ≤1 for competencies and subthemes, reported per CREDES guidance.

The framework elevates cognitive impacts to a distinct theme with three subthemes—cognitive autonomy and deskilling, anchoring and automation bias, and cognitive burden and alert fatigue—and explicitly names both deskilling and "neverskilling." Prior generalist frameworks acknowledged anchoring and automation bias but did not elevate cognitive autonomy and alert fatigue to distinct AI competency subdomains; this theme was created by combining two candidate gap themes that the Round 1 panel rated as the strongest standalone signals not captured by the original four themes. Theme adequacy mean 4.27 with 91% rating ≥4, and the theme name endorsed by 12 of 13 panelists; subtheme T5b anchoring and automation bias mean 4.36 with 100% rating ≥4, while T5a and T5c each scored 4.18 with 91% rating ≥4.

Two cross-cutting conceptual frames emerged: pre-emptive versus post-hoc AI integration, and tiered competencies as an articulated need rather than a pre-specified answer. Pre-emptive integration (AI surfaces recommendations before the clinician completes independent reasoning) and post-hoc integration (AI reviews completed work) were deliberately paired in scenarios S11 and S17, carrying distinct cognitive implications for anchoring, deskilling, and neverskilling; on tiering, Round 1 showed strong support for differentiation at the medical-student level (7 of 10 panelists endorsed substantial differentiation), but the consolidation meeting resolved to articulate the need without pre-specifying which competencies attach to which stage. S11 pre-emptive AI consultation scored mean 4.18 with 91% rating ≥4 in Round 2; S17 post-hoc AI review scored mean 3.82 with 55% rating ≥4, borderline but retained as a future-state operational model; overall scenario set adequacy was endorsed by 91% of the panel (mean 4.18).

The framework provides a structural target for undergraduate, graduate, and continuing emergency medicine education, while noting that assessment instruments with validity evidence do not yet exist and that the framework defines what assessment would measure rather than prescribing methodology. Themes can be mapped to existing curricular touchpoints (clinical reasoning, evidence-based medicine, professionalism, communication, and informatics) without standalone curriculum development; for graduate medical education the competencies are candidates for residency milestone development; for continuing medical education the framework defines a content scope for faculty development modules. This is a consensus framework-building study without curriculum implementation or assessment instrument validation; the authors state that assessment instruments with validity evidence supporting their use for AI competencies do not yet exist in any specialty, and that workplace-based assessment is a natural extension of established entrustable professional activity frameworks.

Perspective

The framework is intended for emergency medicine educators, residency program directors, medical school curriculum designers, and continuing medical education providers to define targets for curricula, assessment, and faculty development in AI-augmented clinical environments; its setting is U.S. emergency departments with AI applications already implemented or near implementation, spanning the patient journey from triage through disposition, and the 10 clinical scenarios are meant to span the competency landscape sufficiently for educational use rather than to enumerate every encounter.

The Delphi panel of 13 was concentrated in academic medical centers, so additional content validity evidence from community-practice and international emergency physicians will affect external validity; 3 of 13 panelists also participated in the antecedent Nominal Group Technique session, a source of non-independence; Theme 2 received the lowest adequacy rating (mean 3.91), reflecting unresolved debate about the depth of mechanistic understanding emergency physicians need; the framework is current with the 2026 AI deployment landscape, and rapid evolution in large language models, ambient documentation, and patient-facing AI may require re-evaluation on an anticipated 18- to 24-month cycle; the framework articulates competencies but does not assess them, and inter-rater reliability of the operational examples among educators has not been established; additionally, this is a preprint that has not been certified by peer review.

Sources