Public articles linked to the same research event.
PLOS digital health The study introduces and evaluates M3, a Python server system built on the Model Context Protocol (MCP) that lets researchers query the MIMIC-IV critical care database in natural language; on 100 answerable and 100 unanswerable EHRSQL 2024 questions, Claude Sonnet 4 reached 94% accuracy and the open-weights gpt-oss-20B, deployable locally on consumer hardware, reached 93% (McNemar p=1.00, no detectable difference), while gpt-oss-20B correctly abstained on 69% of unanswerable questions, with errors mainly reflecting question ambiguity and mismatched schema mappings.
The study introduces and evaluates M3, a Python server system built on the Model Context Protocol (MCP) that lets researchers query the MIMIC-IV critical care database in natural language; on 100 answerable and 100 unanswerable EHRSQL 2024 questions, Claude Sonnet 4 reached 94% accuracy and the open-weights gpt-oss-20B, deployable locally on consumer hardware, reached 93% (McNemar p=1.00, no detectable difference), while gpt-oss-20B correctly abstained on 69% of unanswerable questions, with errors mainly reflecting question ambiguity and mismatched schema mappings.
The study introduces and evaluates M3, a Python server system built on the Model Context Protocol (MCP) that lets researchers query the MIMIC-IV critical care database in natural language; on 100 answerable and 100 unanswerable EHRSQL 2024 questions, Claude Sonnet 4 reached 94% accuracy and the open-weights gpt-oss-20B, deployable locally on consumer hardware, reached 93% (McNemar p=1.00, no detectable difference), while gpt-oss-20B correctly abstained on 69% of unanswerable questions, with errors mainly reflecting question ambiguity and mismatched schema mappings.
The study introduces and evaluates M3, a Python server system built on the Model Context Protocol (MCP) that lets researchers query the MIMIC-IV critical care database in natural language; on 100 answerable and 100 unanswerable EHRSQL 2024 questions, Claude Sonnet 4 reached 94% accuracy and the open-weights gpt-oss-20B, deployable locally on consumer hardware, reached 93% (McNemar p=1.00, no detectable difference), while gpt-oss-20B correctly abstained on 69% of unanswerable questions, with errors mainly reflecting question ambiguity and mismatched schema mappings.