Skip to main content
Back to timeline
medRxivSource publication:

MES: A Multi-Agent Evidence Synthesis System for Medical Decision-Making

Synopsis

This work presents MES, a multi-agent framework of six specialized agents (Supervisor, LiteratureMiner, RWD-Analyst, KG-Specialist, Statistician, SafeChecker) that integrates published literature and trial evidence, generates question-specific real-world evidence from real-world clinical data, and queries biomedical knowledge graphs while preserving source traceability; evaluated in two clinical use cases (an Alzheimer disease medication question with no matching literature and a septic shock beta-blocker question with inconsistent randomized evidence) and across 144 clinical queries spanning six evidence-based medicine categories, it produced structured reports that preserved source traceability, identified cross-source disagreement and communicated uncertainty, and its SafeChecker improv

Source-provided article image: MES: A Multi-Agent Evidence Synthesis System for Medical Decision-Making
medRxiv · Page 1

Interpretation

MES separates evidence synthesis into six agents with defined roles, keeping literature evidence, real-world evidence and knowledge-graph evidence distinguishable and traceable in the final report. Relative to automated systems centered on literature retrieval and summarization, the framework brings generation of real-world evidence (constructing cohorts and executing analyses) and knowledge-graph reasoning into one auditable workflow, recording the source, agent and processing step supporting each conclusion. Supported by the architecture description, the agent roles and tool connections in Figure 1 (PubMed, ClinicalTrials.gov, MIMIC-IV, INSIGHT CRN, iBKH) and the narrative of the two use cases, i.e., system-design and illustrative-demonstration evidence.

When matching literature is absent, MES shifts to real-world data to generate question-specific evidence and explicitly reports the literature gap rather than returning a speculative answer. The use case demonstrates query-adaptive integration: when one source is uninformative the system shifts emphasis to available evidence while preserving the evidentiary boundary of observational comparisons. Based on concrete cohorts from MIMIC-IV/INSIGHT data: 43 patients meeting the strict index-matching definition, a broader analytic cohort of 366 patients with persistent weight loss of at least 5%, and 100 most-similar patients; absolute between-pathway differences in weight and safety outcomes had confidence intervals including zero for 263 continuers versus 86 discontinuers, and switching occurred in only 10 of 366 patients, insufficient for a reliable comparative estimate.

When literature exists but is inconsistent, MES supplements the unresolved trial evidence with a bias-controlled real-world analysis and preserves cross-source disagreement in the final synthesis. The system treats an unresolved trial base as a signal to seek additional real-world evidence rather than a dead end, and does not default to the single positive trial. In the septic shock beta-blocker use case, the RWD-Analyst emulated a target trial, addressed immortal time bias by clone-censor-weight methodology, balanced covariates by propensity score matching and estimated 28-day mortality with a Cox model, finding no association with 28-day mortality (hazard ratio near 1.0, non-significant), with the estimate essentially unchanged after excluding patients with any record of cardiac surgery.

SafeChecker, as a claim-level verification layer, improved detection of unsupported claims on the CliniFact benchmark relative to direct prompting. The result supports a dedicated verification layer that decomposes reports into testable claims and checks each claim against cited evidence, rather than relying on citation alone. On an evaluation set of 394 claim-evidence pairs, binary accuracy rose from 0.850 to 0.926, precision from 0.865 to 0.926, recall from 0.936 to 0.975 and F1 from 0.899 to 0.950, with three-class errors reduced from 74 to 44; the theoretical analysis frames SafeChecker as a calibrated screening mechanism that bounds false-claim acceptance under stated assumptions rather than an oracle.

Perspective

The framework targets clinical and research evidence-synthesis settings that require distinguishing trial evidence, observational associations and mechanistic context, for users able to connect literature databases, real-world clinical data and biomedical knowledge graphs and who need source-traceable reports; the paper positions it as a transparent, source-grounded decision-support framework that complements rather than replaces systematic reviews and expert judgement, and states that the present evaluation does not establish the clinical utility or transportability of MES-generated real-world evidence.

Several open questions remain for a careful reader: automated concept mapping and computable phenotype construction were not independently validated for every query and may introduce cohort misclassification; generated real-world evidence lacks external validation across datasets, leaving transportability to be established; the framework does not implement a formal evidence-certainty framework such as GRADE, so the synthesis should not be read as providing a standardized certainty rating or recommendation strength; knowledge graphs may omit relevant relations or encode non-causal associations, and literature retrieval may miss unpublished or negative studies; and synthesis style and expert-rated quality differ across LLM backbones, indicating the architecture does not eliminate model dependence. In addition, calibration of claim verification across dependent claims within a report and principled stopping criteria for evidence-exhausted states remain to be further specified. Because this reading covered the full text, specific numerical details in figures and supplementary materials were not individually cross-checked, so citing particular statistics should be verified against the original figures and tables.

Sources