Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

npj Digital Medicine

Multi-Agent Collaboration as a Complementary Architecture for AI-Generated Medical Examination Items

Responding to Qian et al.'s finding that a single LLM can generate acceptable knowledge-based questions but struggles with higher-order reasoning, this work proposes MAID, a multi-agent architecture that decomposes item development into specialized agents for drafting, critique, and iterative adversarial refinement, and in a blinded paired-comparison evaluation of items aligned with China's National Medical Licensing Examination standards, 14 experts from 7 medical disciplines compared 25 matched MCQ pairs yielding 350 item-level observations under a two-alternative forced-choice design, with multi-agent outputs receiving 57.7% of expert preferences (95% CI [52.5%, 62.8%]) versus 42.3% for the single-model baseline (95% CI [37.2%, 47.5%]; χ²(1) = 8.33, p = 0.