Skip to main content
Back to timeline
Journal of general internal medicineSource publication:

Optimizing Large Language Models for Hospital Discharge Prediction: A Retrospective Cohort Study with Inference-Time Optimization Strategies

Synopsis

In this retrospective cohort study at a single tertiary academic medical center, large language models predicted same-day discharge from clinical documentation in the 30 hours preceding a 06:00 index time for adult inpatients admitted in 2024 with a length of stay between 2 and 14 days, using a randomly generated validation set (n = 860) and test set (n = 886); the baseline GPT-5 prompt achieved an F1 score of 0.48 and sensitivity of 0.

AI-generated editorial illustration: Optimizing Large Language Models for Hospital Discharge Prediction.

Interpretation

The study provides a baseline performance profile for LLM same-day discharge prediction: the baseline GPT-5 prompt reached an F1 score of 0.48 and sensitivity of 0.37 on the validation set, assessed across F1, balanced accuracy, sensitivity, specificity, positive predictive value, and negative predictive value. Prior discharge-prediction work has often reported model performance, but less often characterized the baseline behavior of a general-purpose LLM on the specific task of same-day discharge within a unified multi-metric framework. Based on two independent randomly generated cohorts, a validation set (n = 860) and a test set (n = 886), with explicit sample sizes and predefined metrics, within a retrospective cohort design.

Through clinically grounded qualitative error analysis, the study identified operational workflows as the most common source of model errors, rather than medical knowledge alone. It moves error analysis beyond general accuracy discussion toward classifiable error domains, indicating that errors concentrate at the operational-workflow level. The qualitative error analysis was conducted from June 2025 to December 2025 as a manual categorization of model outputs, with conclusions presented as a distribution of error domains.

Among three inference-time optimization strategies, automated prompt optimization achieved higher F1 and sensitivity than the basic prompt on the hold-out test set, but with lower PPV and specificity; test-time scaling and expert-led prompt engineering showed minimal improvement. It confines optimization attempts to inference time rather than retraining, and directly compares three strategy classes on the same hold-out test set. Strategy comparison was performed on an independent hold-out test set (n = 886), reporting multiple metrics including F1, sensitivity, PPV, and specificity, showing that gains come with trade-offs across metrics.

Automated prompt optimization improved LLM discharge-prediction performance into the range reported for prior discharge-prediction approaches. It situates the performance of a general-purpose LLM after inference-time optimization against the performance range reported in prior discharge-prediction literature. This conclusion comes from comparing the study's hold-out test set results with previously reported ranges, a comparative statement under retrospective single-center data.

Perspective

The study addresses adult inpatients admitted in 2024 with a length of stay between 2 and 14 days at a single tertiary academic medical center, with the prediction task limited to same-day discharge judgment based on clinical documentation from the 30 hours preceding a 06:00 index time; its conclusions therefore apply to inference-time optimization evaluation within this institutional and population setting. For readers wishing to borrow the framework, it offers a reusable metric set (F1, balanced accuracy, sensitivity, specificity, positive predictive value, negative predictive value), a procedure for randomly splitting validation and test sets, and a comparison paradigm for three inference-time optimization strategies, making it possible to test in other institutions or cohorts whether automated prompt optimization likewise raises F1 and sensitivity while lowering PPV and specificity.

Readers will still watch several open questions: the F1 and sensitivity gains from automated prompt optimization come with lower PPV and specificity, and how to trade off this balance across different thresholds and clinical use scenarios remains open; the specific composition of error domains and actionable intervention points behind operational workflows as the most common error source are not elaborated at the abstract level available here; moreover, the text available here is abstract-level content lacking figures and full result tables, so the specific test-set values for each metric, confidence intervals, and the statistical significance of differences between strategies cannot be confirmed here and would require returning to the original.

Sources