Low-rank compression stops using one rank for every token: per-token routing lifts Llama-2-7B downstream accuracy by 7.6 points
Lead
Instead of applying one fixed rank allocation to every token at inference, a lightweight router per Transformer block picks among nested low-rank paths for each token, raising Llama-2-7B average downstream accuracy by 7.6 percentage points at the same average active-parameter budget.
Story
A compressed language model can now spend different amounts of compute on different tokens: within one Transformer block, tokens take different numbers of low-rank components instead of sharing one fixed rank. Earlier low-rank compression assigned the same compute to every token at inference, with ranks fixed once compression finished. On Llama-2-7B at the same average active-parameter budget, the dynamic LRCC-9 reaches 62.7% average downstream accuracy against 55.1% for the best static baseline, a 7.6 percentage-point gap.
The mechanism trains only routers on top of an existing nested low-rank decomposition: one linear router per block makes a hard choice among a small set of paths from the input activations, and the chosen path sets how many ranks each linear layer in that block activates. The paths come from whole-model rank allocations produced by Swift-SVD at different target costs, while the low-rank factors and pretrained weights stay frozen during training and only the routers are updated. Routers are trained on 256 WikiText-2 samples of 2,048 tokens, annealing the target cost from 0.975 down to 0.3 in steps of 0.025, which yields 28 checkpoints in one run.
What to watch
A next step is to test conditional rank selection on nested subspace models, for example by refining factorized weights through fine-tuning before training routers; latency gains are measured only for batch-size-1 decoding, leaving prefill and larger-batch implementations open; the path structure is fixed before router training, so jointly learning path structure and routing policy is another route.
Routers are trained only on WikiText-2, and actual deployment costs stay close to targets across several unseen datasets, but the models and tasks covered so far remain limited; gains shrink beyond five paths while compilation time rises sharply with path count, and 33-path models could not be compiled at all; the prefill implementation still computes at full rank and masks afterward, so the compute savings from routing are not yet fully realized.
