JEV-like direct-decision models use only 67–76% of the effective ordinal label space across 36 datasets, and BA-LoRA post-training lifts utilization from about 47% to 86%
Synopsis
Analyzing JEV 1.13 and three open KEV models (0.8B/4B/9B), the study finds that on ANLI JEV reaches 74.95% accuracy yet assigns 38.8% of predictions and 51.3% of errors to Neutral, and across 36 ordinal datasets final decisions cover only 67–76% of the effective gold support (versus 87–102% on four nominal tasks); through candidate-order randomization, same-item scale refinement from K=2 to K=14 (utilization falling to 26–75% at K=14), and targeted BA-LoRA post-training (gold-relative utilization rising from roughly 47% to 86% on eight supervised scales), it shows this ordinal scale-utilization bias is a learned and modifiable decision-stage compression of the candidate space.
Interpretation
The paper formulates ordinal scale-utilization bias as a population-level reliability problem: final decisions systematically occupy fewer effective ordinal levels than the gold labels in the same population, measured by an entropy-based effective-support utilization ratio. Prior work on label priors, option order, and judge-specific preferences documented particular preferences such as low scores, Neutral, or one presentation position; this work shifts attention to the shared property of restricted use of model- and task-specific regions of an ordinal scale. Macro-averages over four models on a 40-dataset panel (5,000 items each, stratified by gold label); on 36 ordinal datasets decisions use only 67.2% (KEV-4B) to 75.8% (JEV 1.13) of effective gold support with TVD of 29.8–37.5%, versus 87–102% and TVD 6.0–18.6% on four nominal tasks.
Candidate-order randomization and same-item scale refinement separate this compression from accuracy, gold imbalance, fixed position bias, and candidate count. Randomization weakens but does not remove compression (utilization rising from 55.3–65.0% to 67.7–73.9% on six ordinal datasets while nominal tasks stay at –99.3%), and macro accuracy changes by at most 0.55 points; with items and source scores fixed and gold support and positions balanced, utilization falls for every model as K grows from 2 to 14, reaching 26–75% at K=14. The order intervention covers six ordinal and three nominal datasets with two random permutations per item; the refinement experiment uses 1,400 items from each of three continuous-score sources, with near-balanced gold under equal-frequency binning (99.8% at K=2 and 97.0% at K=14) and DBpedia as a high-accuracy ceiling reference.
The compression is learned and modifiable: targeted BA-LoRA post-training raises gold-relative utilization from roughly 47% to 86% on eight supervised scales while improving accuracy and ordinal error. This turns candidate-space compression from a diagnostic observation into an actionable post-training target and indicates it is not a fixed limitation of the candidate interface. Adapters trained on 32,000 examples whose source IDs and normalized-text hashes are disjoint from the evaluation panel; on training-source datasets utilization rises from 47.4% to 86.8% (0.8B) and 46.4% to 85.6% (4B), TVD falls from 54.7% to 19.7% and 52.9% to 19.0%, and accuracy improves by 13.0 and 15.3 points.
Transfer to unseen sources depends on model size and source similarity, and the authors recommend reporting utilization and TVD alongside accuracy. It separates in-domain recovery from general debiasing: at 4B the other 28 ordinal datasets improve to 81.3% utilization and TVD 27.6%, while at 0.8B the same aggregate moves in the opposite direction (72.7% utilization, TVD 35.1%). Transfer claims use the complete frozen 40-dataset panel; the 15-dataset appendix figure is selected by observed KEV-0.8B gains and is described by the authors as descriptive rather than an unbiased transfer estimate.
Perspective
The results are aimed at researchers and engineering teams using JEV-like direct-decision models for classification, scoring, and automatic evaluation, in settings where candidates form a shared graded scale and adjacent levels are semantically close; in that setting utilization and TVD can serve as routine audit metrics alongside accuracy, and scale refinement can test decision resolution at finer levels. BA-LoRA post-training provides a reproducible path to restore candidate use on supervised scales, with stronger transfer at larger model size, making it most directly usable for teams that hold labeled data for their target scales.
Semantic coupling is operationalized through a binary ordinal/nominal taxonomy, so the account of why ordinal tasks are more compressed rests on task type rather than a graded measure of coupling; the same-item scaling result covers three ordinal sources, and its nominal reference DBpedia is near ceiling for all models rather than a difficulty-matched control; the position result is based on two random permutations per item and all results are point estimates; the utilization ratio compares effective support sizes of predicted and gold distributions and does not imply they assign mass to the same labels, which is why TVD is reported separately; the study covers one hosted system and one open family on a single backbone, and SST-5, DBpedia-14, and IMDb are listed among KEV training sources without an item-level overlap audit, so KEV results on those datasets may partly reflect training exposure.
