Public articles linked to the same research event.
arXiv The work introduces Activation-Supported Parameter Decomposition (ASPD), which jointly decomposes activation and parameter spaces and grounds each learned weight component in the activation features it reads or writes, constraining otherwise non-unique parameter decompositions with the model's internal activations while an internal reconstruction objective provides a local learning signal at the weight matrix being analyzed, thereby enabling scalable, interpretable, and causally editable parameter decomposition in pretrained large language models (demonstrated on Qwen-3-8B); the learned read-write components can also be composed into parameter-level mechanism circuits, used to recover mechanisms underlying the classic IOI circuit and to trace semantic transformations through model weights.
The work introduces Activation-Supported Parameter Decomposition (ASPD), which jointly decomposes activation and parameter spaces and grounds each learned weight component in the activation features it reads or writes, constraining otherwise non-unique parameter decompositions with the model's internal activations while an internal reconstruction objective provides a local learning signal at the weight matrix being analyzed, thereby enabling scalable, interpretable, and causally editable parameter decomposition in pretrained large language models (demonstrated on Qwen-3-8B); the learned read-write components can also be composed into parameter-level mechanism circuits, used to recover mechanisms underlying the classic IOI circuit and to trace semantic transformations through model weights.
The work introduces Activation-Supported Parameter Decomposition (ASPD), which jointly decomposes activation and parameter spaces and grounds each learned weight component in the activation features it reads or writes, constraining otherwise non-unique parameter decompositions with the model's internal activations while an internal reconstruction objective provides a local learning signal at the weight matrix being analyzed, thereby enabling scalable, interpretable, and causally editable parameter decomposition in pretrained large language models (demonstrated on Qwen-3-8B); the learned read-write components can also be composed into parameter-level mechanism circuits, used to recover mechanisms underlying the classic IOI circuit and to trace semantic transformations through model weights.
The work introduces Activation-Supported Parameter Decomposition (ASPD), which jointly decomposes activation and parameter spaces and grounds each learned weight component in the activation features it reads or writes, constraining otherwise non-unique parameter decompositions with the model's internal activations while an internal reconstruction objective provides a local learning signal at the weight matrix being analyzed, thereby enabling scalable, interpretable, and causally editable parameter decomposition in pretrained large language models (demonstrated on Qwen-3-8B); the learned read-write components can also be composed into parameter-level mechanism circuits, used to recover mechanisms underlying the classic IOI circuit and to trace semantic transformations through model weights.