Skip to main content
Back to timeline
JMIR Medical InformaticsSource publication:

Locally Deployed Large Language Models for AI-Assisted Outpatient Prescription Review: Crossover Study

Synopsis

This study deployed the open-source Qwen3-14B model on a hospital intranet server using the Ollama framework, supported it with lightweight knowledge augmentation through exact-match injection from a structured knowledge base derived from drug package inserts, and used a 2-period crossover design in which 2 pharmacists independently reviewed the same 213 outpatient prescriptions under unaided and AI-assisted conditions; human-AI collaborative review achieved 97.2% (207/213) accuracy versus 82.6% (176/213) for pharmacist-alone review, sensitivity was 98% (61/62) versus 55% (34/62), the false-negative rate fell from 45% to 2%, knowledge augmentation reduced the model hallucination rate from 19.7% (42/213) to 4.7% (10/213), mean per-prescription review time fell from 2.33 to 1.

Interpretation

A locally deployed open-source large language model can serve as a pharmacist-supervised prescreening tool that improves accuracy and sensitivity in retrospective outpatient prescription review. Most prior LLM prescription review work relies on cloud-based commercial models, whereas this study used the open-source Qwen3-14B on an intranet server and reported paired comparisons between human-AI collaboration and pharmacist-alone review. A 2-period crossover design in which 2 pharmacists independently reviewed the same 213 prescriptions under both conditions, with a reference standard established by independent consensus between two supervising pharmacists and disagreements adjudicated by a deputy chief pharmacist, analyzed with paired McNemar tests and the Wilcoxon signed-rank test.

Building a structured knowledge base from drug package inserts and injecting it through exact matching substantially reduced the model hallucination rate. The study replaced RAG pipelines that depend on text vectorization and vector databases with lightweight exact-match injection, offering an alternative for hospitals with limited IT resources that struggle to build and maintain such pipelines. On the same 213-prescription test set, the hallucination rate fell from 19.7% (42/213) to 4.7% (10/213), an absolute reduction of 15.0 percentage points and a relative reduction of 76.2%, P < .001.

Human-AI collaborative review shortened per-prescription review time while improving accuracy. The study reported both accuracy and efficiency outcomes, showing that the collaborative condition had higher sensitivity and reduced mean per-prescription review time from 2.33 to 1.12 minutes. Under the paired design, the Wilcoxon signed-rank test gave Z=-12.65, P < .001, r=0.87, with an approximately 51.9% reduction in time.

Keeping all prescription data within the hospital network offers a privacy-preserving route to AI-assisted pharmacy decision support. Addressing the data-breach risk of transmitting protected patient information to cloud-based commercial models, the study used an intranet server deployment that architecturally avoids sending prescription data externally. The model was deployed on a hospital intranet server using the Ollama framework, the knowledge base was derived from drug package inserts, and the study is described as a retrospective prescription review setting.

Perspective

The study addresses pharmacist-led retrospective outpatient prescription review and is relevant to hospitals that want to run models within their own network and have limited IT resources for building and maintaining vector databases and RAG pipelines; its knowledge augmentation depends on a structured knowledge base built from drug package inserts and exact-match injection, so the results mainly apply to review scenarios covered by such an insert-based knowledge base. With 2 pharmacists and 213 prescriptions in a paired crossover design, the conclusions pertain to this test set and retrospective review setting and would need separate evaluation before transfer to other hospitals, other prescription types, or prospective real-time review.

A careful reader would still watch whether these single-center retrospective results with 2 pharmacists and 213 prescriptions hold in other hospitals, other pharmacist groups, and prospective real-time review workflows; how well exact-match injection performs when prescriptions involve information not covered by or requiring updates to the package-insert knowledge base; how the model's hallucination behavior compares on other formularies or prescription complexities, since the plain-AI versus knowledge-augmented comparison was on the same test set; and whether subgroup or stratified results in figures and supplementary materials, which are not available at this summary level, would affect judgments about the scope of applicability.

Sources