CPR routes token-by-token between a base model and its SFT expert, beating the SFT expert by 1.4-5.5% in domain performance while cutting the general-capability drop from 3.4-14.5% to at most 0.5%
Synopsis
The work proposes CPR (Critical-Point Routing), a token-level routing framework between a base model and its SFT expert derivative that trains a lightweight hierarchical router to estimate the per-token expert-call probability based on critical tokens where the base model fails but the expert succeeds, paired with an inference procedure combining momentum smoothing and threshold gating; across diverse model-domain configurations CPR achieves state-of-the-art in all settings, surpassing the SFT expert by 1.4-5.5% in domain performance while recovering its general-capability drop from 3.4-14.5% to at most 0.5%, with minimal overhead from invoking the expert on only one-third of tokens.
Figure 1: Comparison of paradigms against catastrophic forgetting in SFT. (A) Existing methods compress both capabilities into one model, yielding a trade-off. (B) CPR (Ours) decouples them by routing the expert only on critical tokens, preserving both.
arXivInterpretation
It proposes decoupling general and domain capabilities at the model level: keep the original base model for general capability and selectively invoke the SFT expert only when domain-specific knowledge is required, thereby stepping outside the domain-generality trade-off. Existing approaches typically modify the SFT loss to mitigate forgetting and thus inevitably move along a domain-generality trade-off; CPR instead decouples the two capabilities at the model level rather than trading them off. The abstract states it steps "outside this trade-off by decoupling the two capabilities at the model level," keeping the base model and invoking the expert selectively.
It introduces CPR, a token-level routing framework between a base model and its expert derivative, based on critical tokens where the base model fails but the expert succeeds. Routing operates at token granularity and uses the failure/success contrast of critical tokens as the routing signal, rather than switching at sentence or whole-model level. The abstract describes a "token-level routing framework between a base model and its expert derivative, based on critical tokens where the base model fails but the expert succeeds."
It trains a lightweight hierarchical router that estimates the expert-call probability per token, paired with a tailored inference procedure combining momentum smoothing and threshold gating. Routing is implemented as a lightweight hierarchical probability estimator, stabilized at inference by momentum smoothing plus threshold gating. The abstract describes a "lightweight hierarchical router that estimates the expert-call probability per token" and "momentum smoothing and threshold gating."
Across diverse model-domain configurations it achieves state-of-the-art in all settings: domain performance surpasses the SFT expert by 1.4-5.5%, the general-capability drop is recovered from 3.4-14.5% to at most 0.5%, and the expert is invoked on only about one-third of tokens. Relative to the SFT expert it improves domain performance while largely restoring general capability, achieved with expert calls on roughly one-third of tokens and minimal overhead. The abstract reports "state-of-the-art across all settings," "surpassing SFT expert by 1.4-5.5% in domain performance," "recovering its general-capability drop from 3.4-14.5% to at most 0.5%," and "invoking the expert on only one-third of tokens."
Perspective
The work targets LLM deployment settings that need domain adaptation while preserving general capability: the practitioner keeps the original base model for general requests and invokes the SFT expert only when domain knowledge is required. Its design goal is to stop domain performance and general capability from being traded off along a single curve, while controlling overhead by invoking the expert on about one-third of tokens. It is aimed at engineering and research teams that hold both a base model and its SFT expert derivative and can keep both available at inference time.
The abstract does not state the exact criterion for identifying critical tokens, the architecture of the hierarchical router or its training data source, or the parameter settings for momentum smoothing and threshold gating; the number and names of the model-domain configurations and the benchmarks and metrics for domain performance and general capability are likewise not listed in the abstract. These are open questions to confirm in the full text rather than shortcomings of the work.
