Skip to main content
Back to timeline
arXivSource publication:

ACLArena quantifies catastrophic forgetting across a four-stage post-training pipeline and proposes a LoRA expert mixture that lifts AIME26 from 10.21 toward independently trained experts

Synopsis

The work builds ACLArena, a four-stage sequential post-training pipeline spanning math, search, e-commerce, and instruction following; diagnoses forgetting and transfer at both the model level and the token level; systematically compares multi-teacher on-policy distillation (MMOPD), self-distilled fine-tuning (SDFT), and model merging; and proposes Mixture of Low-Rank Experts (MLE), which first consolidates multi-domain trajectories into a shared backbone via SDFT, then freezes the backbone and trains one RL-optimized LoRA expert per stage routed by environment context, improving both in-domain and out-of-domain performance on four reasoning and agentic tasks.

AI-generated editorial illustration: ACLArena: Agent Continue Learning in Multi-stage Post-training

Interpretation

Capability evolution under sequential post-training is non-monotonic and highly task-dependent: early stages show positive transfer, middle stages introduce severe interference, and the final stage only partially restores what was lost. Prior industrial agent training reports rarely compare alternatives under controlled settings; this work makes forgetting and transfer explicit using a fixed curriculum (math, then search, then e-commerce, then instruction following). Table 1 reports that Seq-Math raises NQ from 13.3 to 22.0, that Seq-E-commerce drops AIME from 23.33 to 6.04, NQ from 45.2 to 14.6, and multi-hop search from 37.4 to 9.4, and that Seq-IF recovers NQ to 33.5 and multi-hop to 25.0 while reaching 84.8 on IF-Eval; the final checkpoint ends below earlier peaks on AIME26, NQ, and -Retail.

The mechanism behind forgetting is that task-specific optimization directions are only partially aligned, and that token-level rewriting is concentrated at high-entropy positions. The authors localize interference from two complementary perspectives, model-level (PCA projection of parameter displacement, task-vector cosine similarity) and token-level (top-1 prediction consistency under a fixed prefix, binned by entropy), rather than reporting performance curves alone. The four single-task oracles occupy clearly distinct directions in the PCA projection; after removing the shared prefix, each domain's own RL increment is nearly orthogonal to the others; token-level analysis shows low-entropy positions remain essentially unchanged across all later stages while prediction changes are almost entirely carried by a small high-entropy minority, with consistency falling as entropy rises and along the training order.

The three consolidation paradigms each trade off differently, and none improves a shared checkpoint consistently across all domains. Within one framework the authors group OPD, SDFT, and model merging as online behavior-space, offline behavior-space, and parameter-space reuse respectively, and tune each route to its best attainable configuration rather than evaluating strawman implementations. MMOPD raises AIME26 from 10.21 to 21.25, NQ from 33.5 to 45.2, single-hop from 44.9 to 53.7, and multi-hop from 25.0 to 32.8, while MMLU and IF-Bench dip slightly; SDFT reaches 48.3 on NQ, 55.6 on single-hop, and 38.0 on multi-hop but cuts -Telecom from 45.9 to 21.7 and IF-Eval from 84.8 to 53.6; model merging achieves the strongest MMLU at 80.3 yet performs worse on most e-commerce and IF benchmarks.

MLE places transferable capabilities in a shared backbone and interfering stage-specific residuals in parameter-isolated LoRA experts, approaching independent experts in-domain and posting several top out-of-domain scores. The recipe follows directly from the preceding diagnosis: SDFT rapidly consolidates heterogeneous trajectories into a shared capability centroid, RL updates are far smaller than SFT updates and thus suited to lightweight residuals, and parameter isolation prevents later stages from overwriting acquired specialization. MLE approaches the performance of independently trained task experts across all in-domain benchmarks and achieves the highest reported scores among evaluated methods on GPQA, single-hop search, multi-hop search, and the mock domain; at inference it routes by the interaction environment observable in the agent execution context, without benchmark identities or ground-truth task labels.

Perspective

The framework targets general-purpose agent post-training settings where multiple capabilities must be acquired sequentially, and applies when environment context (system instruction and available tool schemas) is observable at deployment while benchmark identities and ground-truth task labels are not. For practitioners it offers a reusable diagnostic procedure: first use sequential training as a microscope on forgetting and transfer, then choose a consolidation route in behavior space or parameter space; MLE additionally assumes stage-specific residuals can be represented by low-rank adapters and that the routing signal comes from the execution environment itself. For researchers, ACLArena serves as a controlled testbed for comparing new consolidation strategies, since its four-stage curriculum, unified tool-call format, and verifiable rewards make differences between methods easier to attribute.

Experiments are mainly based on Qwen3-8B-Base and one fixed four-stage curriculum, so whether the observed forgetting patterns and method trade-offs reproduce across other model families, longer training sequences, and different curriculum orders remains an open question. MLE relies on observable environment context to select an expert; its effectiveness when that context is ambiguous, changes during an interaction, or requires combining multiple specializations within a single episode is untested, and the authors propose learned routing, dynamic expert composition, and adapter consolidation as future directions to be evaluated jointly for capability retention, transfer, memory usage, and inference latency. In addition, some result tables in the loaded text render their numeric cells as blanks, so the main-result figures cited here come primarily from the body prose; readers who need to verify individual entries should consult the original tables and appendices.

Sources