Skip to main content
Back to timeline
arXivSource publication:

SaveRouter Cuts LLM Routing's Break-Even Deployment Volume by Up to About 9.5x Using Roughly a Third of the Supervision

Synopsis

The work proposes SaveRouter, a sparse-supervision routing framework that selectively acquires informative query-model feedback and shares capability information across related queries, using only about 33-41% of available training feedback across four routing benchmarks while maintaining competitive or better routing quality and reducing the break-even deployment volume by roughly 1.9-9.5x relative to the fastest conventional fully supervised router.

AI-generated editorial illustration: Routing Should Pay for Itself: Sparse Supervision for Economical LLM Routing

Interpretation

The paper brings upfront supervision expenditure into routing evaluation and introduces two metrics: SA-BEP, the number of deployment queries needed for serving-time savings to recover the upfront supervision cost, and SA-CR, the cost ratio after amortizing that expenditure over a fixed deployment horizon. Prior routing work largely evaluates the post-deployment quality-cost trade-off, assuming the router already exists; this work makes the cost of learning to route part of the economic accounting. The metric definitions appear in Section 3 and Appendix A.2 and are reported across four benchmarks under one cost-accounting protocol; cost fields are benchmark-provided, and the authors state absolute monetary costs are not comparable across benchmarks.

SaveRouter learns routing through adaptive sparse feedback acquisition plus hierarchical capability estimation: it groups queries, scores query-model pairs by both estimated capability and uncertainty, then estimates unobserved model quality via a structured group-model prior, shrinkage with local evidence, and a query-level residual correction. Relative to uniform sampling or capability-only/uncertainty-only acquisition, the paper treats which feedback to acquire and how to estimate capability from sparse feedback as coupled problems while retaining query-level refinement. Ablations show removing query-level residual correction lowers peak score by 1.63 percentage points and raises CR from 0.2320 to 0.3484; replacing task groups with embedding clusters or collapsing to a single global group degrades both quality and cost efficiency.

Across four routing benchmarks, the main setting uses only about 33-41% of available training feedback yet attains the highest peak score and lowest target-quality cost ratio, reducing break-even deployment volume by roughly 1.9-9.5x versus the fastest conventional fully supervised router. The reported gain is not serving-time efficiency alone but earlier payback after amortizing upfront supervision, for example 5.3K versus 11.5K on LLMRouterBench and 42.6K versus 326.4K on RouterBench. Results come from LLMRouterBench, Mixinstruct, MMR-Bench, and RouterBench under a shared 20%/80% train/test split with random seed 42, compared against ten baselines within the ORBIT toolkit; WISERouter, BaRP, and SemiRouter are paper-derived reimplementations.

More supervision is not always more economical: the supervision level that minimizes serving cost need not be the one that pays back earliest, and routing quality often saturates well before all feedback is collected. This turns the intuition that dense supervision may be over-provisioned into a quantifiable trade-off, with a proposition characterizing when additional supervision is worth its acquisition cost. The text reports EmbedLLM and kNN reaching 99% of full-supervision accuracy with only 30% and 60% supervision; in ablations random acquisition pays back earlier (3.8K versus 5.3K) but has lower peak score and higher SA-CR@1M; on MMR-Bench and RouterBench the lowest CR and earliest SA-BEP occur at different supervision budgets.

Perspective

The result applies to supervised routing settings that must acquire query-model quality feedback on historical queries before deployment, especially when the training workload or candidate model pool is large and upfront evaluation cost is substantial. The evaluation counts only the model-execution cost of acquiring routing supervision plus post-deployment model-serving cost, not complete system-level cost, and benchmark cost fields and normalizations differ, so absolute monetary values should not be compared or aggregated across benchmarks. The method assumes training-side query grouping information or frozen query representations are available, and that feedback can be acquired selectively per query-model pair; the Appendix F model-pool expansion study concerns retrained sparse onboarding rather than zero-shot or parameter-preserving online admission. For readers making decisions from this, the paper offers a framework for jointly weighing target quality, deployment scale, and supervision budget rather than a single fixed optimal supervision ratio.

Routing quality is not strictly monotonic in the supervision budget, and the paper explicitly does not interpret small fluctuations as evidence that additional supervision is intrinsically harmful, attributing them instead to changes in the observations used to fit the capability estimator that alter the learned routing frontier, so individual operating points should be read with care. SA-BEP and SA-CR are comparable only within a benchmark, and absolute costs are not comparable across benchmarks. BaRP's amortized cost is sensitive to whether repeated interactions are charged: under cached-feedback accounting its SA-BEP falls by roughly an order of magnitude relative to fresh-feedback accounting, so conclusions depend on the feedback-billing assumption. On Mixinstruct most methods are already near-saturated in quality, so differences show up mainly in supervision expenditure; MMR-Bench's final budget covers only about 92.64% of available training pairs and should not be read as full supervision. In addition, WISERouter, BaRP, and SemiRouter are paper-derived reimplementations rather than the authors' original code, and the InferenceDynamics adaptation uses recorded task names to build task profiles while SaveRouter predicts a query's group from its input, a difference in task-side information. These are scope and interpretation questions rather than faults of the work.

Sources