Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

MaskCoFT co-fine-tunes routers and experts with a learnable mask, cutting expert fetches per token by 23.7% on Mixtral-8x7B and 10.1% on DeepSeek-V2-Lite

The work proposes MaskCoFT, a masked co-adaptive fine-tuning method that trains routers and experts together using only the cross-entropy loss: during fine-tuning a learnable binary mask restricts each layer's Top-K routing to a subset of experts so the experts adapt to the redirected tokens, and at inference the learned mask becomes a soft prior that re-ranks experts while every expert remains selectable; with a simulated GPU cache of 4 experts per layer for Mixtral-8x7B and 12 for DeepSeek-V2-Lite, expert fetches per token drop by 23.7% and 10.1% relative to the base model, time per output token in real offloading system serving falls by up to 16.4% and 5.5%, and average accuracy over nine benchmarks stays above the base model by 0.92 and 0.53 points.