Artificial intelligence assessment of Parkland's grading scale in laparoscopic cholecystectomy: a step toward real-world outcome prediction
Synopsis
In routine use of a surgical AI platform, the study automatically assigned Parkland Grading Scale scores across 249 consecutive laparoscopic cholecystectomies, grouped them into Low (PGS 1-2, n=78) and High (PGS 3-5, n=171) severity, compared surgical outcomes, and evaluated model F1, discrimination (AUC 0.932 and 0.896) and calibration against two surgeons' independent double reviews of a 75-video sample, finding that the High group had older age, higher ASA scores, more cholecystitis and urgent surgery, longer operative duration, more intraoperative events and bailouts, longer hospital stay, and more 90-day major complications and readmissions, with operative duration and hemorrhage-related events remaining significant after case-control matching (n=84).
Interpretation
The platform automatically assigned Parkland Grading Scale scores within a real clinical workflow, with 68.7% of cases classified High (PGS 3-5) and 31.3% Low (PGS 1-2). Prior use of PGS was largely manual or retrospective; this shows automated scoring is feasible within routine video capture. Based on consecutive enrollment of 249 cases with explicit group proportions.
Model PGS estimates agreed closely with surgeon ground truth, with F1 of 0.96 for High and 0.93 for Low, discrimination AUC of 0.932 and 0.896, and robust calibration peaking at PGS=3. Provides quantified model performance against two surgeons' independent repeated reviews (n=75) rather than a single rater. Ground truth came from two surgeons reviewing independently and twice, with metrics computed separately against each rater.
High-severity cases were associated with worse perioperative outcomes: operative duration 57.6 vs 35.3 min, intraoperative events 80.1% vs 61.5%, bailouts 10.5% vs 0, hospital stay 2 vs 1 day, and 90-day major complications and readmissions 10.5% vs 2.6%. Links automated grading directly to real-world outcome measures, not only to grading agreement. Multiple comparisons reported p values (mostly p<0.001; complications and readmissions p=0.04).
After case-control matching (n=84) to balance possible confounders, operative duration (mean rank 49.51 vs 35.49, p=0.008) and hemorrhage-related events (mean rank 46.0 vs 39.0, p=0.047) remained significant. Suggests the association between grading and some outcomes is not fully explained by baseline confounders. Matched sample is small, and only two outcomes remained significant.
Perspective
The results apply to laparoscopic cholecystectomy performed under routine video capture, aimed at surgical teams and platform developers who want automated grading to support perioperative risk stratification; the value lies in linking PGS scores to real-world outcomes such as operative duration, intraoperative events, hospital stay, and 90-day complications, providing a basis for prospective validation and clinical framework integration.
Readers would still watch: whether model performance is stable across institutions, video quality, and case mix; how precisely the 84-case matched sample estimates secondary outcomes such as hemorrhage-related events; what the calibration peak at PGS=3 means for clinical decision thresholds; and what further validation is needed to move from correlation toward outcome prediction. This was a summary-level reading without figures or supplementary material, so questions about calibration curve shape and matching details remain open.
