Early evidence for a multi-agent AI simulator for clinical reasoning practice: performance, consistency, and challenges
Synopsis
In Fall 2024, 175 second-year medical students completed three MAESSCR multi-agent LLM clinical encounters as coursework, and six clinician-educators rated 120 randomly sampled transcripts with a dichotomous tool, finding 92% (110/120) adherence to scripted details, 2.5% (3/120) diagnosis-changing information, 12% (14/120) unrealistic patient portrayal, and 17% (20/120) technical issues, while the most prominent problem was agents interpreting findings before students could, occurring in 28% (34/120) of encounters for history/physical exam agents and 59% (71/120) for diagnostics/management agents, with most disruptions judged minor.
Interpretation
MAESSCR assigns multiple LLM-based agents distinct roles such as patient, physical exam, and diagnostic testing, interacting with learners through a text-based encounter, thereby offering scalable deliberate practice without live instructors or standardized patients. Clinical reasoning training is often limited by continuity, feedback, and scalability; this work positions multi-agent LLM simulation as a response and provides early performance data from a real course setting. 175 second-year medical students completed three encounters as coursework, and six clinician-educators developed a rating tool and evaluated 120 transcripts (40 randomly sampled per case) using a dichotomous yes/no scale supplemented by qualitative narrative review.
The simulation performed well on following scripted details, with 92% (110/120) adherence, and diagnosis-changing information occurred in only 2.5% (3/120). This offers checkable early consistency evidence that multi-agent simulation can reproduce preset cases, beyond mere technical feasibility. Based on 120 randomly sampled transcripts rated by six clinician-educators on a dichotomous scale, with qualitative narrative review.
Agents most often interfered with students' independent reasoning by interpreting findings before students had the opportunity: history/physical exam agents caused interference in 28% (34/120) of encounters and diagnostics/management agents in 59% (71/120), though most disruptions were minor and unlikely to compromise the overall encounter. The work turns the question of whether AI supports rather than replaces clinical reasoning into scorable interference types and proportions, indicating that educational value depends on role stability, contextual fidelity, and learners' opportunity to interpret clinical information independently. Interference proportions come from dichotomous ratings of 120 transcripts with qualitative narrative review; the authors judged most disruptions minor.
Unrealistic patient portrayal appeared in 12% (14/120) of encounters and technical issues in 17% (20/120). These proportions convert challenges of multi-agent simulation from general concerns into quantified early observations, pointing to role performance and platform functionality as design considerations. Also based on dichotomous ratings of 120 randomly sampled transcripts with qualitative narrative review.
Perspective
This work applies to text-based multi-agent clinical encounters used by second-year medical students as coursework, focusing on script adherence, patient portrayal realism, platform functionality, and interference with independent reasoning; its conclusions speak to medical education design and research settings that aim to use LLM simulation to expand deliberate practice.
This is early evidence and the reading scope is limited to the summary, without figures or full methodological detail; readers may still watch whether interference proportions change across courses, agent role configurations, and learner groups, and whether the judgment that most disruptions were minor holds in larger samples.
