Skip to main content
Back to timeline
arXivSource publication:

SeLMRoute turns query requirements into a probabilistic state via 16 interpretable semantic probes, reaching 72.08% average accuracy on LLMRouterBench

Synopsis

SeLMRoute splits LLM routing into three stages: a decision model first answers 16 interpretable semantic questions about each query and keeps the probability distributions as a 40-dimensional semantic state, a lightweight supervised regressor then predicts each candidate model's performance, and deployment objectives are applied last; on LLMRouterBench (15 datasets, 20 candidate models, 11,481 queries) it reaches 72.08%±0.45 average accuracy and 72.64% under grouped five-fold out-of-fold evaluation, versus 69.23% for the strongest fixed candidate, and in a 13-model performance-cost setting it improves performance in all five grouped splits with a mean PerfGain of 2.66%.

AI-generated editorial illustration: SeLMRoute: Probabilistic Semantic Evidence for Large Language Model Routing

Interpretation

The paper makes "what a query requires" an explicit, reusable probabilistic semantic state: 16 probes (8 binary Noul, 8 four-level Score) cover math reasoning, code reasoning, formal logic, factual recall, social-affective content, tool interaction, external knowledge, current information, plus domain specialization, reasoning depth, constraint density, context integration, decomposition need, ambiguity, exactness, and answer openness, with raw probability mass forming a 40-dimensional representation. Existing routers mostly learn decisions directly from query embeddings, model representations, preference data, or clusters of similar examples, and their routing representation rarely states what a query actually requires; SeLMRoute gives each coordinate a declared semantic meaning before the performance learner is trained, and contains no candidate identities. The probe table and dimensionality derivation appear in the main text; semantic extraction is performed once per query and reused across all downstream experiments.

The framework decouples semantic evidence extraction, candidate performance learning, and routing objectives, so one semantic state can serve different candidate pools and deployment objectives. The performance learner receives only the semantic state rather than the original text, and model identities are absent from that state; adding a model requires an empirical performance mapping but not re-extraction of existing semantic vectors. The primary implementation uses a multi-output CatBoost regressor to predict the candidate pool jointly; the performance-cost setting uses candidate-specific regressors because candidate-query cells are missing, leaving the semantic representation and routing objective unchanged.

On the LLMRouterBench performance-oriented setting, the probability-mass representation achieves the highest mean performance among the evaluated semantic, dense, lexical, and domain-level representations, and exceeds the published router values listed in the paper. The 40-dimensional probability mass outperforms the 88-feature Full representation and the 16-judgment Hard representation; GTE-Qwen2 embeddings, DomainOnly, and TF-IDF do not surpass it. Across five grouped splits probability mass has the highest mean AvgAcc, the largest Gain@B, and the smallest Gap@O; the paired out-of-fold difference versus Hard has a confidence interval including zero and is not significant at the 5% level; comparisons with published routers use a different split protocol and are treated as external references, not paired statistical claims.

The same semantic representation supports performance-cost routing: in a setting of 13 flagship models, 10 datasets, and 12,446 queries, test performance beats the GPT-5 reference in all five grouped splits, but strict monetary savings are not established. The routing parameter is selected on an inner validation partition and fixed before test evaluation, preventing retrospective selection of a cheaper operating point from favorable test outcomes; configurations that are cheaper and more accurate on test therefore do not count as strict CostSave. PerfGain ranges from +0.52% to +4.58% with a mean of 2.66%; no split achieves positive strict CostSave, and the two configurations meeting the validation accuracy requirement are slightly more expensive than GPT-5.

Perspective

The work targets deployment settings that must assign requests among multiple candidate LLMs, especially where routing evidence should be human-inspectable and one query representation should be reusable across candidate pools and deployment objectives. Because semantic evidence extraction is separated from performance learning, changing prices, quality preferences, or candidate availability does not require redefining the semantic evidence; the open-weight decision-model substitution shows the interface is not tied to one backend. The intended setting assumes historical candidate-query outcomes are available to train the performance mapping and that the extra latency of semantic extraction is acceptable.

The advantage of probability mass over hard judgments has a paired out-of-fold confidence interval that includes zero, so whether it holds on broader benchmarks remains open; the twelve-probe compressed configuration shows a small observed difference but does not establish noninferiority within the predefined margin. Out-of-domain generalization shows two regimes: when a single dataset is held out, semantic and dense representations are close, while when an entire task domain is held out the dense representation leads on the point estimate, indicating that describing an unfamiliar query does not replace evidence about candidate behavior in that region. Semantic entropy ranks difficult queries reasonably within the calibration distribution, but a fixed threshold transferred to an unseen domain yields no meaningful accuracy gain. When a new candidate is added, semantic vectors remain reusable but the performance learner still needs candidate-specific calibration. Semantic extraction is the main latency cost, and its monetary estimate comes from the performance-oriented benchmark, a different query population from the performance-cost benchmark, so the two should not be added directly.

Sources