Multi-Agent Collaboration as a Complementary Architecture for AI-Generated Medical Examination Items
Synopsis
Responding to Qian et al.'s finding that a single LLM can generate acceptable knowledge-based questions but struggles with higher-order reasoning, this work proposes MAID, a multi-agent architecture that decomposes item development into specialized agents for drafting, critique, and iterative adversarial refinement, and in a blinded paired-comparison evaluation of items aligned with China's National Medical Licensing Examination standards, 14 experts from 7 medical disciplines compared 25 matched MCQ pairs yielding 350 item-level observations under a two-alternative forced-choice design, with multi-agent outputs receiving 57.7% of expert preferences (95% CI [52.5%, 62.8%]) versus 42.3% for the single-model baseline (95% CI [37.2%, 47.5%]; χ²(1) = 8.33, p = 0.
Fig. 1: The MAID multi-agent item development framework.
PubMedInterpretation
It introduces MAID, a multi-agent item development framework decomposing the process into one Author Agent, three Reviewer Agents, one Editor Agent, and optional Image Generation and Image Verification Agents, following a generation → independent review → editorial synthesis → revision → re-evaluation workflow. Relative to the single-model, single-pass generation paradigm evaluated by Qian et al., this framework splits quality control into specialized stages so that distinct failure modes are handled by dedicated mechanisms rather than a single model's implicit judgment. The paper provides a full description of the architecture and role division and states that complete prompts and implementation details are available in the MAID GitHub repository; the prompt texts themselves are not reproduced in the main text.
In the blinded paired comparison, multi-agent items received 57.7% of expert preferences, significantly exceeding the 42.3% for the single-model baseline, with the direction of advantage consistent across all seven disciplines. This provides quantitative support for the argument that the single-model boundary on higher-order reasoning items is architectural rather than a ceiling on model capability, whereas Qian et al.'s conclusion rested on single-model evaluation. 14 experts, 7 disciplines, 25 matched MCQ pairs, 350 item-level observations, double-blind two-alternative forced-choice design; χ²(1) = 8.33, p = 0.004; paired t-test t(13) = 2.81, p = 0.015.
The framework decouples agent roles from specific language models, with multiple LLMs (including GPT-4o, Claude 3.5, DeepSeek-R1, Qwen, Gemini, and MedSeek) randomly assigned to agent roles across tasks. This design aims to prevent any single model's systematic biases or knowledge gaps from propagating unchecked through the pipeline, a structural difference from single-model item generation. The paper presents this as a critical design principle and states that random assignment occurs across tasks; specific assignment records and implementation details are in the supplementary materials and code repository.
The framework integrates multimodal capability: a dedicated image generation agent autonomously decides whether an item needs visual content and produces medically relevant images (e.g., radiographs, CT scans, fundus photographs), after which a separate verification agent evaluates semantic consistency between visual and textual content. This directly addresses the limitation noted by Qian et al. that current AI cannot reliably generate clinically accurate images integrated with examination items. The paper describes the modular design of the image generation and verification agents; the main text does not report an independent quantitative evaluation of image quality itself.
Perspective
The result is aimed at medical examination item development and assessment contexts, particularly item types requiring higher-order clinical reasoning and multimodal content; the design intent is for multi-agent collaboration to carry structural quality control so that human experts can redirect effort toward finer-grained clinical judgment and pedagogical alignment. The framework is open-sourced to facilitate replication and adaptation across medical specialties and assessment contexts.
The main text notes variation in preference magnitude across disciplines, with significant preferences in some and positive but non-significant trends in others, with details in the supplementary materials; expert characteristics, models and prompts, and evaluation procedure details also reside mainly in the supplementary materials and code repository. The actual quality performance of the image generation and verification modules, and behavior across different examination systems and broader item types, remain open to further validation.
