Skip to main content
Back to timeline
arXivSource publication:

RouteFM turns LLM routing into a reusable pretrained capability, beating the strongest baseline by 2.23 quality points on the unseen multimodal MMR-Bench with only eight observations per candidate

Synopsis

The work introduces RouteFM, reframing large language model routing from environment-specific local fitting into a pretrain-once, route-anywhere foundation-model paradigm: episodic pretraining across heterogeneous routing environments learns a reusable, target-conditioned candidate-comparison capability in which candidates are represented only by anonymous behavioral evidence (historical query embeddings, observed quality, normalized relative cost) with no model names or identity embeddings, and at inference the router stays frozen while adapting to new domains, modalities, candidate pools, and smaller context budgets through context alone, outperforming the strongest non-RouteFM baseline by 2.23 quality points on the pretraining-excluded MMR-Bench with eight observations per candidate.

AI-generated editorial illustration: Pretrain Once, Route Anywhere: Towards a Foundation Model for LLM Routing

Interpretation

The paper reformulates LLM routing as a foundation-model problem: instead of fitting a router per environment, it learns a single shared routing function across environments, with environment differences expressed by conditioning on the current query distribution and candidate pool through behavioral evidence. Existing routers largely follow what the authors call local fitting, where supervision is collected for a particular query workload and candidate pool and environment changes often require collecting new behavioral observations and re-optimizing; RouteFM instead treats routing as a transferable capability, pretrained episodically over varying tasks, candidate-pool composition, candidate ordering, and context size. The paper provides a problem formulation contrasting the environment-specific mapping with a shared routing function, plus concrete episodic pretraining construction including random candidate permutation, disjoint context and target query IDs, and natural, opportunity, and boundary target-episode regimes.

RouteFM's mechanism centers on in-context capability profiling of anonymous candidates and target-conditioned comparison: behavioral observations are encoded into observation tokens and compressed into a fixed set of latent capability tokens as a capability profile, the target query retrieves evidence through parallel profile and raw-context branches, and candidates are compared within the pool by a Transformer encoder over the candidate dimension. Unlike approaches tied to fixed model identities or persistent model embeddings, RouteFM receives no model names, providers, parameter counts, or model-ID embeddings, so candidate representations arise entirely from observed behavior; it also keeps both the compressed profile and the uncompressed observation sequence as complementary views to preserve fine-grained evidence that compression may lose. Ablations show the largest degradation when the candidate-pool Transformer is removed, with a notable drop under limited observations; removing target-to-context attention also consistently reduces performance, particularly under limited observations; replacing the Qwen3-VL query representation with a text-only BGE encoder degrades performance across observation budgets.

A frozen RouteFM achieves the highest routing quality in both in-domain routing and cross-modal transfer, with its advantage concentrated in the low-context regime where behavioral evidence is scarce. The cross-modal evaluation uses MMR-Bench, which is excluded from pretraining and differs from pretraining in both query characteristics and modality, making it a stronger distribution shift; under this setting a small amount of behavioral context suffices for the frozen router to adapt to new candidate behavior without parameter updates. On MMR-Bench, RouteFM attains the highest routing quality across all observation budgets, with the advantage most pronounced at eight observations, exceeding the strongest non-RouteFM baseline by 2.23 quality points; the margin generally decreases as more behavioral evidence becomes available, reaching only a small gap in the large-observation regime. On in-domain RouterEval it likewise achieves the highest quality across observation budgets, and the same frozen model remains competitive across the entire context range.

Deployment changes can be handled through context rather than retraining: new-model incorporation, target-domain shifts, and reduced context budgets all work with frozen parameters. Using a support-controlled protocol, the paper compares training from scratch, fine-tuning after pretraining, and the frozen RouteFM: training the same architecture from scratch consistently underperforms, while fine-tuning the pretrained model on the target domain yields only small changes, marginally better at some budgets and slightly lower at others, indicating that much of the transferable routing capability is acquired during pretraining. In new-model incorporation, RouteFM surpasses all refitted baselines after observing only eight examples of the new candidate, and with sixteen observations routing over the expanded pool outperforms routing over the original incumbent pool; in the context-budget experiment, RouteFM matches the full-context performance of Weighted NN, MLP, and EmbedLLM using only a fraction of available observations, with larger fractions required for stronger full-context references.

Perspective

The result targets deployers who must trade quality against cost across heterogeneous and evolving candidate model pools, especially where behavioral evidence is scarce, candidate pools change frequently, or routing must cross modalities. The intended setting is one where candidates are described by anonymous behavioral evidence (historical query embeddings, observed quality, normalized relative cost), the router stays frozen after pretraining, and deployment preferences are expressed through the quality-cost coefficient in the routing objective. The reported capability comes from a heterogeneous mixture of four pretraining sources (LLMRouterBench, RouterBench, RouterEval, MixInstruct), and the pretraining-composition analysis shows the largest degradation when LLMRouterBench is removed, indicating that transfer depends on the diversity and balance of pretraining environments. Cost is treated as an episode-relative signal rather than an absolute monetary or latency measurement, so predicted cost should not be compared across independently normalized episodes.

Several open questions remain for a careful reader. First, the paper treats cost as an episode-relative feature, and how latency, availability, resource constraints, and other system-level signals could enter a unified decision interface is left as future work. Second, the evaluated candidate-pool sizes and observation budgets have explicit ranges, and behavior in larger, continuously evolving routing ecosystems is framed as a future direction. Third, the cross-modal transfer gain narrows when observations are plentiful, so the advantage sits mainly in the low-context regime and the relative value under rich evidence depends on deployment conditions. Fourth, the loaded text is the full paper body and appendices, but some table values appear as figures or placeholders, so a few specific scores cannot be read directly from the text and should be taken from the original tables.

Sources