Auditing Sex/Gender Disparities in Emergency Triage with LLM-based Paired Comparisons
Synopsis
The study introduces a domain-agnostic paired-comparison approach that fine-tunes a large language model to emulate documented emergency triage decisions and then compares predictions on sex-swapped pairs in which only sex is flipped while documented clinical content is held constant, finding that otherwise identical presentations were more likely to receive a lower-severity (less urgent) predicted triage score as female than male, with a pooled per-pair rate of about 1.1% (95% CI 0.9–1.3) across more than 140,000 Bordeaux University Hospital admissions and a directionally consistent but larger 2.2% (1.7–2.7) in MIMIC-IV, while a model retrained on sex-neutralized inputs eliminated the between-sex prediction gap, indicating the asymmetry is mediated by explicit sex markers.
Fig. 1: The workflow consists of (1) LLM fine-tuning for emergency triage prediction, (2)
· Page 3Interpretation
It proposes and tests an LLM-based paired-comparison auditing framework that quantifies the sensitivity of documentation-level decisions to sex/gender markers without defining a 'correct' triage standard. Relative to the earlier work that introduced only the core methodological concept, this study develops Directional Triage Skew (DTS), Net Mean Difference (NMD), and Net Asymmetric Triage Shift (NATS) metrics and derives a pooled per-pair rate, making systematic disparities rigorously quantifiable. Applied to the Bordeaux University Hospital ED dataset (n = 72,444 test partition) and the MIMIC-IV test set, with all proportion metrics reported with 1,000-iteration bootstrap 95% confidence intervals.
In sex-swapped pairs, female variants were more likely than male variants to receive a lower-severity (less urgent) predicted triage score, and this direction was consistent across two datasets spanning different countries, languages, and healthcare systems. Prior studies reported that women are assigned lower acuity in acute coronary syndrome, stroke, and abdominal pain, but lacked paired evidence holding clinical content fixed while changing only sex; this paired design controls population-level baseline morbidity differences by construction. Pooled per-pair rate 1.09% (95% CI 0.89–1.29) in Bordeaux and 2.18% (1.67–2.74) in MIMIC-IV; both datasets show DTSF|M > 0 and DTSM|F < 0 in a consistent direction.
The sex/gender asymmetry is mediated by explicit sex markers, and tabular sex fields versus textual gender cues influence predictions in opposite directions. The earlier framework did not distinguish which channel carries gender information; this study's modality-isolated analyses show that editing only text produces predominantly decreases in predicted triage score in both directions (DTSF|M = −10.3%, DTSM|F = −12.3%), while flipping only the tabular sex field produces predominantly increases in both directions (+3.8% each). On MIMIC-IV, retraining with sex-neutralized inputs collapsed the male–female mean predicted score gap from −0.0332 (95% CI [−0.045, −0.021], p = 6.4 × 10⁻⁷) to −0.0005 (95% CI [−0.013, +0.013], p = 0.93), while κw remained comparable (0.603 → 0.627).
The gender-differential pattern varies systematically with triage nurse–patient sex combinations, suggesting the model captures stable features of the recorded data rather than random artifacts. Prior work did not use nurse sex information; this study shows female nurses exhibit the strongest effects in both the M-F and F-F dyads, while male nurses show DTS metrics close to zero regardless of patient sex. Stratified analysis of the four patient–nurse sex combinations in the Bordeaux dataset (M-M: n = 8,398; M-F: n = 29,444; F-M: n = 6,697; F-F: n = 23,820), with 95% confidence intervals reported for each combination.
Perspective
The framework applies wherever decision inputs (structured and/or textual) are available and a well-specified transformation for the protected attribute(s) can be defined; the authors list human resources screening, academic admissions and grading, and justice-related risk assessment and pretrial release as examples. In the emergency triage context, the results apply to documentation-level sensitivity auditing that helps identify which clinical presentations or contexts show the most pronounced sex/gender-associated asymmetries and thereby generate hypotheses; the authors note the results are being shared and discussed with clinicians at Bordeaux University Hospital to help inform triage nurse training protocols, with the broader ambition of collaborating with national bodies such as the French National College of Emergency Physicians (SFMU) toward national recommendations for more equitable emergency care after appropriate validation.
The authors explicitly flag several open questions: an inherent 'multimodality gap' between model inputs and the full set of cues used in real triage, since speech prosody, tone, visible distress, agitation, and overall demeanor are not captured; the LLM necessarily inherits asymmetries present in training data and may amplify or distort biases differently than humans, a phenomenon the authors are investigating through latent space exploration; the binary gender framework does not capture the full spectrum of gender identity and its intersection with other demographic factors; no multi-seed fine-tuning robustness checks were conducted; and the sex-neutralization retraining was performed on MIMIC-IV only, not replicated on the French cohort. In addition, the authors stress that mechanically projecting the documentation-level rate onto France's roughly 20.9 million annual emergency visits to yield about 110,000 additional lower-severity assignments is 'easily over-interpreted,' is not a count of harmed patients, and is not anchored to any tangible care difference. Readers should also note that the supplementary materials, supplementary tables, and supplementary notes referenced in the text were not loaded here, so some metric definition details and stratified percentage breakdowns cannot be verified from this text.
