Public articles linked to the same research event.
arXiv The work proposes MaskCoFT, a masked co-adaptive fine-tuning method that trains routers and experts together using only the cross-entropy loss: during fine-tuning a learnable binary mask restricts each layer's Top-K routing to a subset of experts so the experts adapt to the redirected tokens, and at inference the learned mask becomes a soft prior that re-ranks experts while every expert remains selectable; with a simulated GPU cache of 4 experts per layer for Mixtral-8x7B and 12 for DeepSeek-V2-Lite, expert fetches per token drop by 23.7% and 10.1% relative to the base model, time per output token in real offloading system serving falls by up to 16.4% and 5.5%, and average accuracy over nine benchmarks stays above the base model by 0.92 and 0.53 points.
The work proposes MaskCoFT, a masked co-adaptive fine-tuning method that trains routers and experts together using only the cross-entropy loss: during fine-tuning a learnable binary mask restricts each layer's Top-K routing to a subset of experts so the experts adapt to the redirected tokens, and at inference the learned mask becomes a soft prior that re-ranks experts while every expert remains selectable; with a simulated GPU cache of 4 experts per layer for Mixtral-8x7B and 12 for DeepSeek-V2-Lite, expert fetches per token drop by 23.7% and 10.1% relative to the base model, time per output token in real offloading system serving falls by up to 16.4% and 5.5%, and average accuracy over nine benchmarks stays above the base model by 0.92 and 0.53 points.
The work proposes MaskCoFT, a masked co-adaptive fine-tuning method that trains routers and experts together using only the cross-entropy loss: during fine-tuning a learnable binary mask restricts each layer's Top-K routing to a subset of experts so the experts adapt to the redirected tokens, and at inference the learned mask becomes a soft prior that re-ranks experts while every expert remains selectable; with a simulated GPU cache of 4 experts per layer for Mixtral-8x7B and 12 for DeepSeek-V2-Lite, expert fetches per token drop by 23.7% and 10.1% relative to the base model, time per output token in real offloading system serving falls by up to 16.4% and 5.5%, and average accuracy over nine benchmarks stays above the base model by 0.92 and 0.53 points.
The work proposes MaskCoFT, a masked co-adaptive fine-tuning method that trains routers and experts together using only the cross-entropy loss: during fine-tuning a learnable binary mask restricts each layer's Top-K routing to a subset of experts so the experts adapt to the redirected tokens, and at inference the learned mask becomes a soft prior that re-ranks experts while every expert remains selectable; with a simulated GPU cache of 4 experts per layer for Mixtral-8x7B and 12 for DeepSeek-V2-Lite, expert fetches per token drop by 23.7% and 10.1% relative to the base model, time per output token in real offloading system serving falls by up to 16.4% and 5.5%, and average accuracy over nine benchmarks stays above the base model by 0.92 and 0.53 points.