Skip to main content
Back to timeline
arXivSource publication:

One token-level interaction decomposition links data selection, collision-versus-erosion forgetting, and plasticity loss into a single learning-dynamics account

Synopsis

The work derives a token- and layer-wise decomposition of how learning from one token changes another prediction, splitting it into a direct force-alignment channel (CH1) and a diffuse channel mediated by shared readout geometry (CH2), with a forward-computable approximation using only logits, hidden states, and the readout; following this interaction over time, the authors report that positive interaction supports data selection, negative interaction splits into concentrated collision and accumulated erosion with distinct controls, and that declining task-conditioned readout transmission during long-horizon training tracks declining future learnability, with readout restoration partially recovering plasticity.

AI-generated editorial illustration: Learning Dynamics of Continual Learning: A Unified View of Data Attribution, Forgetting, and Plasticity Loss

Interpretation

The paper unifies three continual-learning questions — which experience to learn from, what an update changes, and whether the model can still learn next — as one evolving update–behavior interaction, and gives a token- and layer-wise decomposition: separating the softmax force, shared readout geometry, and residual connections exposes a direct force-alignment channel CH1 and a diffuse channel CH2 coupled through the shared readout, which under identity-path and linear-block approximations becomes forward-computable. Relative to the earlier example-level learning-dynamics framework (Ren & Sutherland, 2025), the granularity is refined to tokens and layers and the model structure is opened explicitly, so the interaction can be probed repeatedly without full-parameter Jacobians. The decomposition is derived (Proposition 2.1), and the paper reports that CH1 nearly matches measured log-probability change under readout-only updates; CH2's numerical fidelity is characterized separately in Appendix C.

Positive interaction serves directly as a selection criterion: on controlled attribution benchmarks CH1 nearly saturates sentence-transformation and mathematical attribution, while in a cross-lingual setting that weakens output-vocabulary overlap, adding CH2 raises Qwen2.5-1.5B retrieval accuracy from 0.488 to 0.625 on MMLU and from 0.445 to 0.600 on GSM8K; in multilingual agent-memory retrieval, Llama3.2-3B top-1 retrieval accuracy rises from 0.861 to 1.000 with downstream answer accuracy gains. Unlike influence-function or gradient-alignment attribution, the score needs only forward quantities and no per-example backpropagation; the paper reports that fine-tuning on GSM8K subsets selected by CH1+CH2 at a 1% budget outperforms random selection and LESS. Evidence comes from controlled attribution benchmarks, cross-lingual retrieval, and subset fine-tuning across Qwen2.5-1.5B, Qwen3-4B, and Llama3.2-3B; scores are computed after a 100-step in-distribution warm start while downstream fine-tuning starts from the original base checkpoint.

Negative interaction splits into two mechanisms: collision, concentrated interference from a few high-energy updates, detectable via a token-energy criterion (Proposition 4.1 gives two-way bounds between extreme energy and a large negative force) and responsive to energy masking; and erosion, the coherent accumulation of many weak interactions, which can be redirected by changing the response format of incoming supervision without preventing the new task from being learned. Forgetting was often treated as a single phenomenon or explained through representation or task-subspace overlap; here signed local interactions place abrupt capability loss and gradual behavioral drift in one framework with mechanism-specific controls. The paper reports energy masking helps in the extreme-energy regime while overly aggressive thresholds sharply degrade downstream learning; in format-separation experiments GSM8K is learned in all settings while instruction-following degradation varies with format matching, with a complementary check on instruction-tuned models.

Long-horizon training reshapes the shared readout geometry: task-conditioned readout transmission (a Rayleigh quotient along task-induced force directions) declines over roughly 100M PubMed tokens for three held-out OOD tasks (GSM8K, MBPP, Dolly-QA) while in-distribution PubMed shows no comparable systematic decay; the decline coincides with slower subsequent adaptation, restoring the readout to its base value substantially improves validation loss, and the degree of degeneration correlates strongly with the benefit of restoration (Spearman correlation reported on OOD tasks). This turns plasticity loss from a separate failure mode into an outcome of the same evolving interaction geometry and provides a measurable diagnostic rather than a post hoc observation that fine-tuning got harder. Evidence comes from checkpoint measurements along a long-horizon PubMed trajectory, a controlled sequential-SFT readout-reshaping experiment, and readout-reset comparisons across models, tasks, and checkpoints; the paper explicitly treats the readout as important but not the sole source.

Perspective

The framework targets the stage where experience is consolidated into parameters, for Transformer language models with residual connections and a shared readout; the authors note that when a system relies mainly on in-context learning or external memory, the cycle of selection, interference, and sustainability still appears but with different mechanisms. Actionable uses include ranking candidate experience by the interaction score, screening high-conflict updates with the energy criterion, redirecting erosion through response-format separation, and using readout transmission as a diagnostic of long-term learnability to decide whether to restore the readout. The authors stress that readout restoration is a mechanistic intervention rather than a deployment strategy, and that preserving capabilities acquired during long-horizon training while recovering plasticity is outside the scope of this diagnostic.

The authors describe the analysis as fundamentally local in time: the interaction analysis concerns a single update, whereas forgetting and plasticity loss emerge over many updates, so the mechanistic argument extrapolates local structure along a training trajectory and its predictions are validated empirically rather than derived from an exact multi-step theory. On approximations, the first-order expansion, identity-path, and linear-block approximations jointly determine fidelity; the paper reports that interaction magnitude is reliable even at a cold start, but signed fidelity can fail at initialization and recovers after a small amount of in-distribution adaptation, with signed error accumulating with residual depth. On optimization, the derivation uses SGD as the canonical setting; adaptive optimizers change channel weights and rankings but not the origin of the channels, and decoupled weight decay does not affect candidate ranking at a fixed observation. On empirical scope, the authors frame results as mechanistic patterns rather than universal quantitative laws, and list higher-order residual paths, optimizer-state evolution, longer-term feedback, and broader architectures as open directions. In addition, some tables in the loaded text appear with empty cells, so verifying specific numbers still requires the original appendices.

Sources