Public articles linked to the same research event.
Nature medicine This work developed and evaluated a fully on-premise clinical AI agent that couples local operational control with a multi-perspective reliability framework to support selective autonomy, achieving 90.04% accuracy on a seven-disease task and 83.8% on a four-disease task across two MIMIC-IV-derived benchmarks, and finding that diagnostic behavioral consistency best discriminated correctness (AUC = 0.860, and AUC = 0.875 under stress testing), with a consistency threshold of 0.90 retaining 49.4% of cases at 98.9% diagnostic accuracy.
This work developed and evaluated a fully on-premise clinical AI agent that couples local operational control with a multi-perspective reliability framework to support selective autonomy, achieving 90.04% accuracy on a seven-disease task and 83.8% on a four-disease task across two MIMIC-IV-derived benchmarks, and finding that diagnostic behavioral consistency best discriminated correctness (AUC = 0.860, and AUC = 0.875 under stress testing), with a consistency threshold of 0.90 retaining 49.4% of cases at 98.9% diagnostic accuracy.
This work developed and evaluated a fully on-premise clinical AI agent that couples local operational control with a multi-perspective reliability framework to support selective autonomy, achieving 90.04% accuracy on a seven-disease task and 83.8% on a four-disease task across two MIMIC-IV-derived benchmarks, and finding that diagnostic behavioral consistency best discriminated correctness (AUC = 0.860, and AUC = 0.875 under stress testing), with a consistency threshold of 0.90 retaining 49.4% of cases at 98.9% diagnostic accuracy.
This work developed and evaluated a fully on-premise clinical AI agent that couples local operational control with a multi-perspective reliability framework to support selective autonomy, achieving 90.04% accuracy on a seven-disease task and 83.8% on a four-disease task across two MIMIC-IV-derived benchmarks, and finding that diagnostic behavioral consistency best discriminated correctness (AUC = 0.860, and AUC = 0.875 under stress testing), with a consistency threshold of 0.90 retaining 49.4% of cases at 98.9% diagnostic accuracy.