Skip to main content
Back to timeline
PLOS digital healthSource publication:

Utilization of a HIPAA-compliant large language model chatbot in an academic pediatric medical center

Synopsis

This mixed-methods case study analyzed 14 months of utilization of "InternalGPT," a HIPAA-compliant LLM chatbot at an academic pediatric medical center, finding that 2,149 of approximately 15,788 employees (13.6%) requested access, 52.8% recorded at least one token use, the top 20% of users consumed 69.4% of tokens, and among 461 sustained users, 92 self-report survey respondents indicated a mean 30% productivity gain corresponding to an exploratory perceived productivity value of $6.3M to $18.9M under varying extrapolation assumptions.

Source-provided article image: Utilization of a HIPAA-compliant large language model chatbot in an academic pediatric medical center.
PubMed

Interpretation

The study provides empirical data on usage patterns of a HIPAA-compliant LLM chatbot within a hospital system, showing that 33.6% of employees who requested access never logged in, 13.7% logged in but never used tokens, and the top 20% of users consumed 69.4% of tokens. Prior case studies of similar institutional deployments mainly described implementation successes and challenges, with sparse data on hospital employee engagement; this study fills that gap with 14 months of system logs and survey data. Based on system logs from 2,149 requesters among approximately 15,788 employees, plus 142 non-use/discontinuation survey responses (10.3% response rate) and 92 productivity survey responses (20.0% response rate).

Sign-up motivations aligned with professional roles: operations professionals prioritized administrative automation (55.8%), clinicians focused on clinical augmentation (52.5% of clinical-related reasons), and scientists emphasized research and coding (literature summarization 29.9%, coding and data assistance 21.5%). Using LLM-powered thematic analysis validated by human coding, the study quantified motivation distributions across roles and noted these motivations target core job functions rather than peripheral convenience tasks. Thematic analysis used o3-mini to generate 15 categories and GPT-4o for coding; human-LLM agreement (κ=0.59–0.64) was as good as or better than agreement between two human coders (κ=0.58).

Among 461 sustained users, 92 self-report survey respondents estimated a mean 30% productivity gain (SD 28.1%), corresponding to $13,636 to $40,950 per standard/power user and $6.3M to $18.9M across all 461 users under assumptions that non-respondents gained between one-sixth and the full respondent gains. The study combined self-reported productivity gains with employee cost assumptions to produce an exploratory institutional-level value range, explicitly noting these figures reflect perceived effort reduction rather than measured output or cost savings. Based on 92 voluntary anonymous one-question surveys (20.0% response rate), with a sensitivity analysis for non-respondent differences that the authors describe as exploratory.

The main barriers to non-use and discontinuation were limited time (51.4%) and difficulty using (21.8%), integrating (15.5%), or accessing (9.2%) InternalGPT, rather than AI-feature or performance concerns. The study shifts the adoption barrier framing from model capability to system capacity building, highlighting the importance of user education and technical integration investments. Based on 142 voluntary one-question survey responses (10.3% response rate) with structured multiple-choice and free-text options.

Perspective

The findings apply to voluntary enterprise-level LLM chatbot deployment in an academic pediatric medical center setting, and can help health systems plan user education, technical integration, and token quota strategies; for readers interested in how different professional roles integrate LLMs into core job functions, the study provides role-stratified motivation and usage data.

Readers should still watch: productivity gains are self-reported perceptions rather than measured output, and the 20% response rate may overrepresent satisfied users; low survey response rates may introduce non-response bias; token quotas may have capped demand among the heaviest users, meaning the true utilization distribution could be even more skewed; concurrent other AI tools and EMR deployment may have influenced usage patterns; professional role assignment by LLM is approximate and includes misclassification; future studies could use randomized access and intent-to-treat designs to measure benefit.

Sources