Skip to main content
Back to timeline
arXivSource publication:

Intervention experiments across DeepSeekMoE, OLMoE and Qwen3-MoE show post-merge expert routing drift is mostly input-representation-induced yet fails to predict task gains from source-route restoration

Synopsis

The work introduces a routing analysis toolkit and, across DeepSeekMoE, OLMoE and Qwen3-MoE, crosses source and merged router inputs and parameters and runs paired token- and task-level evaluations, finding that most expert reassignments are input-shift-induced rather than parameter-induced, that source-relative routing differences poorly predict next-token likelihood gains from source-route restoration, and that different expert selections can produce directionally similar mixture outputs; it therefore redefines routing failure as task loss recoverable under a specified routing intervention with non-routing parameters fixed, and proposes Selective Router Repair as a case study that finds no reliable evidence that source-specialist token-likelihood advantages identify beneficial local corr

AI-generated editorial illustration: Routing Drift Alone Does Not Diagnose Failure in Merged MoE LLMs

Interpretation

Across DeepSeekMoE, OLMoE and Qwen3-MoE under Average and Task Arithmetic merging, crossing source and merged router inputs and parameters shows that most token–layer expert-set changes are representation-induced, while router-parameter-only changes account for a small share. Prior work such as HARC treats source-to-merged routing mismatch as routing breakdown and realigns the merged router; this work separates input-side from parameter-side origins of routing change and argues that structural change is not itself harm. 256 domain-balanced prompts per setting, all continuation tokens and all sparse layers; attribution ratios recomputed over 10,000 prompt resamples, with redundant and interaction-only cases distinguished.

Source-relative routing differences predict next-token negative log-likelihood gains from source-route restoration near chance, with full-distribution JS divergence AUROC near 0.5; at task level, replacing routes across question–choice sequences yields correct-choice margin changes whose 95% confidence intervals span zero. It demotes routing disagreement from a failure label to a structural description awaiting test, and supplies a task-consequence criterion instead. Structural prediction is evaluated on changed-route events in the final five sparse layers, and no primary test survives Holm correction; task evaluation is paired on fixed items, with all 12 margin and accuracy intervals including zero.

At fixed merged hidden states and expert parameters, source-route and native merged-route mixtures have higher cosine similarity than the best expert pair, and mixture cosine exceeds the expert-pair maximum in all 727 sampled events; observed routes exceed 32 random alternatives matched on expert count, overlap and routing weights, with positive paired prompt-bootstrap intervals in every setting. It provides mixture-level functional-redundancy evidence showing that different expert selections can preserve aligned mixture outputs, helping explain why routing drift need not imply behavioral degradation. Event-level comparisons across six settings with matched random controls preserving cardinality, shared subset, BF16 source-weight multiset and routed mass; the authors note cosine measures direction, not magnitude or task recovery.

Operationalizing routing failure as task loss recoverable under a specified routing intervention with non-routing parameters fixed detects recoverable loss under deliberate OLMoE router-logit permutation (about 12–14 percentage points recovery at four layers and about 27–33 pp at full layers, with eight comparisons passing Holm correction), whereas source-route restoration and frozen LC route replay do not establish reliable task benefits in the evaluated merged models; the Selective Router Repair case study likewise finds no reliable evidence that source-likelihood advantages identify beneficial corrections or improve average task performance. It separates candidate construction from demonstrated recovery: fitting source-derived targets, or outperforming another repair method, does not substitute for evidence of improvement over the unrepaired parent. Positive controls are paired on 256 fixed ARC items; SRR is reported on 35,326 shared items over five evaluation runs, with five of eight OLMoE/Qwen3-MoE settings showing positive average-score changes, though intervals condition on fixed checkpoints and 440 matched local diagnostics show no higher utility for source-direction contrasts.

Perspective

The results target researchers and engineering teams studying merged MoE language models, in diagnostic settings covering DeepSeekMoE, OLMoE and Qwen3-MoE under Average, Task Arithmetic, TIES and WUDI-Merge merging. What it offers is a reusable procedure: first cross source and merged router inputs and parameters for local attribution, then use paired task evaluation with non-routing parameters fixed to test whether a specified routing intervention recovers loss, and only then judge whether repair is warranted. The analysis toolkit and SRR code are released, allowing candidate repair construction to be evaluated separately from its task-level effects.

Task-evaluation intervals condition on fixed checkpoints and do not cover variation across repair seeds, calibration samples or training runs; the null result for source-route restoration establishes neither harmlessness nor optimality of merged routing, and the authors note the local candidate comparison is not an exhaustive route search. SRR's fitting objective matches selected router-logit residuals rather than task loss, and during joint execution the input response of later layers changes the realized correction, so fixed-input corrections need not reproduce under joint execution. In addition, the Qwen3-MoE extended diagnostics were added after earlier diagnostics and its correlation intervals are pointwise; some tables and figures appear as placeholders in the loaded text, so specific values should be checked against the original and its appendices.

Sources