Skip to main content
Back to timeline
arXivSource publication:

FetalAgents orchestrates specialized fetal-ultrasound models through multiple agents, achieving top performance across eight clinical tasks on external validation and auto-generating structured reports

Synopsis

The work proposes FetalAgents, a multi-agent system built on AutoGen in which a GPT-5-mini-driven Coordinator parses clinical intent and anatomical plane and dynamically dispatches expert agents that combine specialized vision models such as FetalCLIP, nnU-Net, USFM, SAMUS, and AoP-SAM via deterministic fusion rules, while a Summarizer consolidates outputs into structured reports; across eight clinical tasks (standard plane classification, brain plane classification, abdomen and stomach segmentation, AoP estimation, AC and HC measurement, GA prediction) on multi-center external datasets, FetalAgents achieves the best or most competitive results compared with specialized vision models, ultrasound foundation models, and general/medical MLLMs, and supports end-to-end video-stream keyframe ext

Source-provided article image: FetalAgents: A Multi-Agent System for Fetal Ultrasound Image and Video Analysis

Interpretation

FetalAgents organizes fetal ultrasound analysis into three agent types—Coordinator, Experts, and Summarizer. The Coordinator, based on GPT-5-mini, classifies the user query into a task type and identifies the standard ultrasound plane, then rephrases the query into a structured prompt and dynamically dispatches it to relevant experts; expert agents wrap task-specific vision models and produce results through deterministic fusion rules, communicating upstream exclusively via structured JSON. Existing automated models are mostly standalone predictors for a single subtask, leaving clinicians to manually coordinate tool selection and translate predictions into structured documentation; this work brings tool selection, anatomical context identification, and result consolidation into one language-driven coordination framework. The method description is complete, specifying the responsibilities of the three agent types, the concrete integrated models (FetalCLIP, FU-LoRA, RadImageNet-initialized ResNet-50, ViT-B/16, nnU-Net, USFM, SAMUS, AoP-SAM, UperNet, ConvNeXt-Tiny), and the fusion and communication mechanisms; the system is built on the AutoGen framework.

Across eight clinical tasks on multi-center external validation, FetalAgents achieves the best or most competitive results: standard plane classification accuracy 0.927 and Cohen's κ 0.903 (FetalCLIP: 0.914/0.887); brain plane classification F1 0.896 and κ 0.856; abdomen segmentation DSC 0.937 and HD95 2.32; stomach segmentation DSC 0.859 and HD95 10.64; AC measurement MAE 5.53 mm and Acc@5% 84.0%; HC measurement MRAE 1.40% and P95 RAE 3.75%; GA prediction validity rate 84.93% and MAE 1.23 weeks; AoP estimation MAE 5.44° and Acc@5% 67.2%. The paper reports that it consistently outperforms specialized vision models, ultrasound foundation models, and general and medical MLLMs across the eight tasks; for example, GPT-5-mini reaches only 0.605 accuracy on standard plane classification, MedGemma 0.708, and Gemini-3-flash 0.880. Results come from independent external datasets (e.g., multi-center African data, a 511-case HC18 subset), covering 233 standard-plane cases, 569 brain-plane cases, 187 abdomen cases, 253 stomach cases, 64 AoP cases, 187 AC cases, 75 HC cases, and 511 GA prediction cases; metrics include Accuracy, F1, Cohen's κ, AUROC, DSC, IoU, HD95, ASSD, PPV, Sensitivity, MAE, MdAE, MRAE/MdRAE, P95 AE/P95 RAE, Acc@5%, and Validity Rate.

The system supports end-to-end video-stream summarization: a 6-class keyframe identification model fine-tuned on the FetalCLIP encoder backbone automatically extracts diagnostic keyframes from continuous video, dispatches appropriate experts for frame-wise analysis, and aggregates temporal findings into a video-level clinical report; during image caption generation, the system computes growth percentiles using WHO fetal growth charts as a consistency safeguard, triggering a reflection step over expert outputs and replacing outliers when measurements deviate beyond plausible growth bounds. Existing models typically handle a single static frame and a single task, leaving keyframe extraction and documentation to manual work; this work chains keyframe identification, frame-wise analysis, and report aggregation into one pipeline and introduces an automatic consistency check based on growth charts. The keyframe model is fine-tuned on PBF-US1; qualitative evaluation (Fig. 2) shows the system correctly identifying the trans-thalamic plane and generating a report concordant with human experts in image captioning, and autonomously extracting multi-plane keyframes and cross-referencing ultrasound-derived GA against LMP-derived GA in video summarization.

Perspective

The results target prenatal screening and structural assessment in fetal ultrasound, suited to clinical workflows that need to chain plane classification, anatomical segmentation, and biometry into a complete process, and to scenarios that automatically organize continuous ultrasound video into structured reports. For researchers wishing to reuse the framework, the system is built on AutoGen and trains its expert models on public datasets, with source code released, making it a reference implementation for orchestrating specialized vision models into multi-step workflows; for clinical readers, its value lies in demonstrating an automated path from image or video to an auditable structured report, in which WHO growth-chart percentiles serve as an automatic checkpoint for measurement consistency.

The report-generation part (image captioning and video summarization) is presented through qualitative examples, without a systematic quantitative comparison against human expert reports, so how reproducible its report quality is remains to be shown with more evidence. The GA prediction validity-rate metric relies on WHO growth charts as a reference, and their applicability across populations affects how that metric should be interpreted. The system integrates multiple external models and datasets, and how differences in each expert's out-of-distribution performance propagate to the final report is not broken down item by item in the text. The paper lists prospective clinical evaluation and integration of longer-term patient histories as future work, indicating that current evidence comes mainly from retrospective external validation.

Sources