Skip to main content
Back to timeline
arXivSource publication:

Semantic entropy flags small-model errors on GSM8K at AUROC 0.871, but its synthetic-arithmetic counterpart is fully matched by a regex difficulty baseline

Synopsis

The work tests semantic entropy as an escalation signal for routing queries from a small to a large language model across three benchmarks and two model families, finding that on GSM8K it reliably distinguishes the small model's mistakes (AUROC 0.871) and improves routed accuracy over random escalation by up to nine points at matched cost, while an equally strong-looking synthetic-arithmetic result is almost exactly matched by a model-free rule based only on how hard the question looks on the page; it then offers a checklist of controls, including a difficulty baseline, three escalation targets, an oracle-headroom pre-flight check, per-generation cost accounting, and a difficulty probe that predicted in advance a retrieval shortcut's collapse from AUROC 0.908 to 0.518.

Source-provided article image: Evaluating Escalation Signals for LLM Routing: Targets, Controls, and Five Ways to Fool Yourself
Figure 1 ·

Figure 1: The cost-quality frontier on GSM8K ( r = 0.56 r=0.56 , γ = 12 \gamma=12 ), drawn from Table 10 together with the cost model of Section 3.4 (which supplies always-small’s cost of 1 1 unit; always-small has no row in Table 10 itself). Always-small and always-large never route, so they carry no AUROC and are shown as cost-only reference bars.

arXiv

Interpretation

On GSM8K, semantic entropy separates the small model's errors and converts into real routing gains: Target 2 AUROC 0.871, Target 1 AUROC 0.734, routed accuracy peaking at 0.528 near a 0.56 escalation rate, exceeding random-at-matched-cost and beating always-large accuracy for escalation rates between roughly 0.17 and 0.56. Semantic entropy was previously developed for hallucination detection; this work evaluates it end to end as an escalation signal in a full routing pipeline and reports both the operationally relevant Target 1 and the fair Target 2. Gemma-3 1B/12B pair on 1000 GSM8K items with 10 samples per query, AUROC with bootstrap intervals; replicated on a second disjoint GSM8K sample, with the two runs agreeing to 0.008 AUROC on Target 2.

On synthetic arithmetic the strong-looking entropy result is nullified by a zero-cost difficulty baseline: regex-recovered difficulty scores Target 2 AUROC 0.839 and 0.830, statistically indistinguishable from entropy's 0.830 and 0.867, with the sign of the difference flipping between runs. The authors state that prior cascade studies do not report a difficulty baseline; this control converted their own initially accepted positive result into a null, and they recommend reporting both raw and calibrated variants and taking the maximum as the bar. Two profiling runs over the same 1000-item synthetic set sharing items and model pair, differing only in re-sampling; entropy's rank correlation with difficulty is 0.98 and its Somers' D against the Target 2 label is 0.97, indicating its discriminative power is largely difficulty tracking.

The choice of escalation target itself determines the verdict: needs_escalation depends on the large model's behavior, which a signal computed only from the small model's samples can reach only indirectly through shared difficulty, so the same signal is strong on Target 2 and weak on Target 1, reaching only 0.613 and 0.666 on synthetic Target 1. The paper presents this target divergence as its central conceptual finding and notes that Target 1's positives concentrate in the middle of the difficulty range, so ranking by raw difficulty is wrong at precisely the end a router acts on. On synthetic arithmetic the Target 1 versus Target 2 gap is about 0.2 AUROC across two runs, exceeding the measured roughly 0.02 noise floor; on GSM8K entropy's Target 3 score of 0.639 is indistinguishable from raw difficulty's 0.633, consistent with a shared-difficulty mediation account.

Live semantic-entropy routing can cost more than simply calling the large model because of its sampling overhead, while a cheaper retrieval signal built from cached outcomes is weaker: for the 1B/12B pair live entropy costs 16.7 units per query against always-large's 12.0, whereas retrieval costs 6.8 units at Target 2 AUROC 0.607. The paper makes explicit the per-generation overhead hidden by a normalized cost model, derives the break-even condition, and uses a model-free difficulty probe to forecast, before running GSM8K, the retrieval Target 3 collapse from 0.908 to 0.518. Costs use parameter count as a proxy, and the authors state that real serving cost also depends on batching, quantization, KV-cache bandwidth, and hardware; retrieval costs essentially the same as random escalation (6.8 versus 6.7 units) while supplying a signal.

Perspective

The work targets cascade routing that decides whether to escalate before calling the large model, in settings where the small/large capability gap is large enough and the large model succeeds where the small one fails. Its positive result concerns GSM8K with a Gemma-3 1B/12B pair, a configuration the authors describe as winning on accuracy but losing on cost, economically sensible only at large capability ratios or substantially smaller sample budgets. The retrieval signal targets settings where queries fall within the knowledge base's coverage and difficulty is not a surface feature of the question, with a reported deployment point of about 6.8 units per query, roughly 57% of always-large cost. The proposed checklist is directly usable for designing evaluations of other escalation signals and model pairs.

Sampling temperature was fixed at 1.0 and never ablated, despite directly setting the sample diversity semantic entropy measures; the NLI entailment threshold was fixed at 0.5 and also not ablated. The knowledge-base profiling did not preserve per-item scores, so the AUROC figures in Tables 9 and 10, including the Target 3 collapse central to the retrieval discussion, carry no bootstrap interval, and the authors list re-running with per-item scores retained as a concrete next step. Both retrieval evaluations are single runs, and the smaller GSM8K differences (such as Target 1's 0.570 versus 0.656) are not individually resolvable at that precision. The relation between embedding-difficulty correlation and retrieval performance rests on only two data points, and whether degradation between them is smooth or threshold-like is unknown. Retrieval assumes queries arrive from a distribution the knowledge base covers, and behavior under sparse coverage is untested. The authors also note that the failure modes cataloged are the ones they happened to detect, so undetected ones may remain.

Sources