Skip to main content
Back to timeline
medRxivSource publication:

StrokeAgent matched expert final decisions in 86% of 100 acute stroke cases, exceeding tool-free LLMs and neurologists

Synopsis

The study presents StrokeAgent, a multimodal agentic decision-support framework built around temporal alignment and continuous information synthesis that processes clinical records, laboratory results, raw electrocardiographic (ECG) waveforms and source multimodal CT images to support emergency triage, the initial thrombolysis decision and final reperfusion planning, and in a curated set of 100 patients with multimodal clinical and imaging data enriched for clinically complex reperfusion scenarios it achieved 86% concordance with expert-adjudicated final decisions, exceeding the mean of three tool-free large language models by 13 percentage points and the mean of six resident and attending neurologists by 17 percentage points, with 76% complete pathway concordance, completing the full deci

Source-provided article image: An auditable multimodal agentic system for sequential decision-making across the acute stroke pathway
Fig. 1 ·

Fig. 1 Overview of StrokeAgent. StrokeAgent processes clinical and imaging information directly from source data as it becomes available through three stage-specific agents for emergency triage, ini- tial thrombolysis decision and final reperfusion planning. At each stage, specialized tools and clinical subagents return structured findings to a shared case record, which the agent reasons over together with guideline-based rules and matched trial evidence to formulate a stage-specific recommendation. The upper panel shows an illustrative case progressing through the three decision stages. The lower panels summarize the agent decision cycle and the processing components available to the system. Illustration created with BioRender.com.

medRxiv · Page 5

Interpretation

StrokeAgent is designed as a multimodal agentic decision-support framework built around temporal alignment and continuous information synthesis, processing clinical records, laboratory results, raw ECG waveforms and source multimodal CT images across emergency triage, the initial thrombolysis decision and final reperfusion planning. Unlike decision support limited to text or a single modality, the framework natively handles raw ECG waveforms and source CT images and formulates and updates recommendations as information becomes available, matching the sequential nature of acute stroke care. The abstract describes the system's input modalities and three decision points, and states that evaluation used 100 patients with multimodal clinical and imaging data deliberately enriched for clinically complex reperfusion scenarios.

Each recommendation is grounded in an auditable chain of supporting findings and evidence. Beyond issuing a recommendation, the framework exposes a traceable evidence chain so the recommendation can be inspected, which distinguishes it from tools that output conclusions alone. The abstract states that the system grounds "each recommendation in an auditable chain of supporting findings and evidence", but the loaded text does not show the concrete form or an example of that chain.

Across 100 patients, StrokeAgent reached 86% concordance with expert-adjudicated final decisions, exceeding the mean of three tool-free LLMs by 13 percentage points and the mean of six resident and attending neurologists by 17 percentage points under matched information content and timing; complete pathway concordance was 76%, with even larger margins over both comparator groups. The comparison was run under matched information content and timing, and results are reported at two levels: the final decision and the full pathway. The abstract reports the two concordance figures of 86% and 76% and the differences of 13 and 17 percentage points, on a sample of 100 curated multimodal cases.

In 1,200 blinded clinician assessments, StrokeAgent received higher overall clinical utility ratings than each LLM (common odds ratios, 10.8 to 16.3) and higher reasoning quality ratings, with fewer outputs judged to pose severe potential harm; on average it completed the full decision pathway 73.9 s faster than the case level median neurologist time. Beyond concordance with expert decisions, the study adds blinded clinician ratings of clinical utility, reasoning quality and potential harm, plus decision time as a further dimension. The abstract reports 1,200 blinded assessments, a common odds ratio range of 10.8 to 16.3, and a 73.9 s time difference, all as quantified results reported by the study.

Perspective

The work targets emergency triage, the initial thrombolysis decision and final reperfusion planning along the acute ischemic stroke pathway, for emergency and neurology clinical teams that must integrate clinical and imaging information under time pressure, especially where access to specialist expertise is limited. What it demonstrates is feasibility-level evidence: concordance with expert-adjudicated decisions and blinded clinician ratings on 100 curated multimodal cases deliberately enriched for complex reperfusion scenarios. To move forward, the abstract states that prospective studies are needed to determine whether physician-supervised use improves clinical care; the data availability statement also notes that de-identified individual-level data are not publicly available because of ethical, privacy and institutional restrictions, and that qualified researchers may request access by submitting a proposal to the corresponding author.

The loaded text is an incomplete reading scope containing only the abstract, competing interest statement, author declarations and data availability statement, so method details, case composition, evaluation workflow and statistical models cannot be verified here. The abstract does not specify the composition and experience distribution of the three tool-free LLMs and six neurologists, nor the specific rating scales behind the common odds ratios or how the 1,200 assessments were allocated. The 86% and 76% figures correspond to the final decision and the complete pathway respectively; the source of that difference, and what the 73.9 s time difference means in a real emergency workflow, still require the methods section. The key open question is the one the abstract itself raises: prospective studies are needed to determine whether physician-supervised use improves clinical care, so the current results should be read as feasibility evidence rather than evidence of clinical benefit.

Sources