ASPD grounds weight decomposition in activation space, recovering the IOI circuit and tracing semantic transformations in Qwen-3-8B
Related research and updatesSynopsis
The work introduces Activation-Supported Parameter Decomposition (ASPD), which jointly decomposes activation and parameter spaces and grounds each learned weight component in the activation features it reads or writes, constraining otherwise non-unique parameter decompositions with the model's internal activations while an internal reconstruction objective provides a local learning signal at the weight matrix being analyzed, thereby enabling scalable, interpretable, and causally editable parameter decomposition in pretrained large language models (demonstrated on Qwen-3-8B); the learned read-write components can also be composed into parameter-level mechanism circuits, used to recover mechanisms underlying the classic IOI circuit and to trace semantic transformations through model weights.
Figure 1: Parameter-level mechanisms recovered for the IOI circuit ( Wang et al., 2022 ) . ASPD identifies weight components implementing known head-level computations and connects them through read–write interactions across heads. For example, the duplicate-name component 224 O 224O of H3.0 and the induction component 228 O 228O of H5.5 provide inputs to the downstream S-inhibition mechanism in H8.6.
arXivInterpretation
Introduces ASPD, which jointly decomposes activation and parameter spaces and grounds each learned weight component in the activation features it reads or writes. Existing interpretability methods largely study the two spaces separately; this work connects represented information with parameter-level computation, addressing that gap. The abstract presents the method design and motivation and states that this grounding constrains otherwise non-unique parameter decompositions using the model's internal activations; the demonstration model is Qwen-3-8B.
Adds an internal reconstruction objective that provides a local learning signal at the weight matrix being analyzed, making the decomposition scalable, interpretable, and causally editable. Compared with decompositions relying only on global or external objectives, this provides a local learning signal at the weight-matrix level, supporting scalability and causal editing. The abstract states this objective and these properties and reports demonstration on pretrained large language models; no quantitative metrics are given.
The learned read-write components can be composed into parameter-level mechanism circuits and used to recover mechanisms underlying the classic IOI circuit and trace semantic transformations through model weights. Raises parameter decomposition results to the level of mechanism circuits, letting weight-level components correspond to known circuits and semantic transformations. The abstract reports demonstrations on Qwen-3-8B using the IOI circuit and semantic transformation tracing; no numerical circuit-match or editing-effect values are provided.
Perspective
The work targets parameter-level interpretability and editing of pretrained large language models, with the demonstration setting being Qwen-3-8B; the method grounds weight components in activation features and provides a local learning signal via an internal reconstruction objective, suited to analyses that map weight components to specific read-write features and compose them into parameter-level mechanism circuits. The cases described in the abstract are recovering mechanisms underlying the classic IOI circuit and tracing semantic transformations through weights.
The abstract provides no quantitative metrics or comparison conditions for decomposition quality, causal editing effects, or circuit recovery, nor does it state computational cost or hyperparameter sensitivity; these are open questions to confirm in the body. In addition, the loaded text is abstract-level and contains no figures or experimental details, so the practical strength of scalability and editability remains a question to verify.
