Sparse Autoencoders Localize the Clinical Triage Format Effect: Medical Features Go Inactive at the Decision Token, Scaffold Features Account for Over 91% of Gemma Attribution
Related research and updatesSynopsis
Using sparse-autoencoder (SAE) features in Gemma 3 4B/12B IT and Qwen3-8B on clinical triage vignettes, this study finds that medical features fire on the shared clinical narrative under both formats but are inactive at the multiple-choice decision token; emergency-tier information is linearly decodable from vignette representations (ROC-AUC 0.95–1.00) with no significant format difference yet is attenuated at the decision token, and scaffold-peaking features account for over 91% of unsigned attribution in both Gemma models, placing the strongest correlates of the format effect at answer selection.
Figure 1: Study overview. Within each input style, the clinical text is held fixed while output format changes (left). The analyses progressively test behavior, error composition, vignette-level representation, decision-token localization, and case-level predictability (right).
arXiv · Page 3Interpretation
Medical features fire on the shared clinical narrative under both response formats but are inactive at the multiple-choice decision token. Prior observations of multiple-choice triage failures stayed at the behavioral level; this work moves the analysis into internal representations, separating case processing from answer mapping. SAE feature analysis in Gemma 3 4B/12B IT and Qwen3-8B, supported by natural-language autoencoder verbalization and top-feature characterization.
Emergency-tier information is linearly decodable from vignette representations with ROC-AUC 0.95–1.00, with no significant format difference, but is attenuated at the decision token. Indicates the format difference does not stem from clinical information being unencoded, but from that information being weakened at the final answer mapping. Linear decoding ROC-AUC range and the absence of a significant format difference.
In a direct linear projection, the identified medical features contribute zero, whereas scaffold-peaking features account for over 91% of unsigned attribution in both Gemma models. Shifts attribution of the format effect from medical content toward the multiple-choice scaffold itself. Direct linear projection attribution analysis across two Gemma models.
Behaviorally, whether multiple choice improves or worsens performance depends on the model; option-order shuffles rule out simple positional bias, and cases that differ between formats are usually one severity tier apart. Provides behavioral boundary conditions for the format effect and rules out positional bias as an alternative explanation. Option-order shuffle experiments and severity-tier comparison of cases differing between formats.
Perspective
This work targets researchers and evaluation designers who assess LLM clinical triage using multiple-choice vignettes, and applies to analysis of models such as Gemma 3 4B/12B IT and Qwen3-8B in clinical triage settings. It places the strongest correlates of the format effect at answer selection, offering a reproducible analysis path for future work separating case processing from answer mapping, with code and data available in the study repository.
The authors explicitly note it remains open whether unmeasured clinical representations also differ by format. In addition, this reading is at the abstract level, lacking sample sizes, case counts, statistical test details, and figure or table information; these gaps affect judgment of effect robustness and generalization scope, so readers needing to assess conclusion strength should consult the original paper and the public repository.
