Skip to main content
Back to timeline
arXivSource publication:

AWT uses activation-derived diagonal preconditioning to close 12–60% of the tensor-network compression perplexity gap on Llama, Ministral, and Qwen

Synopsis

The authors propose Activation-aware Weight Tensorization (AWT), a training-free, solver-agnostic calibration-time preconditioner that applies an activation-derived diagonal equivalent reparameterization to each weight matrix before an unchanged TT/TTN solver and deploys it with only an input-side elementwise rescaling. Across Llama 3.1 8B, Ministral 8B, and Qwen2.5 7B at 2–6 compression, AWT consistently improves vanilla tensorization: it closes 12–35% of the WikiText perplexity gap under single-operator replacement and 27–60% under Llama multi-operator suffix replacement, with gains transferring to HellaSwag and ARC-Challenge.

AI-generated editorial illustration: Activation-Aware Weight Tensorization: A Calibration-Time Preconditioner for Tensor-Network LLM Compression

Interpretation

AWT transfers activation-aware equivalent reparameterization from quantization to tensor-network decomposition: it computes a per-channel diagonal scale from cached input activations, right-preconditions the weight matrix, hands it to an unchanged TT/TTN solver, and applies the inverse scaling to the layer input at deployment. Post-training tensorization was previously activation-blind, with standard TT/TTN backends minimizing an isotropic weight-space reconstruction objective; AWT injects activation information without changing the tensorization shape, topology, rank budget, or solver. Validated on Llama 3.1 8B, Ministral 8B, and Qwen2.5 7B across TT, TTN4, and TTN8 backends, seven dense Transformer operators, and 2–6 compression; calibration and metric caches are separated, and calibration examples are disjoint from downstream evaluation examples.

Diagonal preconditioning is equivalent to a column-weighted Frobenius reconstruction proxy, so input channels with stronger activations are fit more accurately, lowering activation-conditioned output error. The paper provides an identity linking tensorization error to activation-conditioned functional distortion, explaining why the scaling strength that optimizes functional fidelity need not minimize weight-space reconstruction error. At matched parameter budgets, moderate activation reweighting reduces relative output error; weight error keeps falling as scaling strength grows, but output error is best at moderate strength, and changes in output error are more predictive of perplexity changes than changes in weight error.

AWT consistently beats vanilla tensorization in single-operator and multi-operator suffix replacement, with larger absolute gains when several operators are replaced jointly, indicating that activation-aware preconditioning helps control accumulated tensorization error. Gains hold across three model families, three compression ratios, seven operator types, and TT/TTN backends, and transfer to HellaSwag and ARC-Challenge downstream tasks. Single-operator replacement closes 12–35% of the WikiText perplexity gap; Llama suffix replacement closes 27–60% across attention-group and all-seven-matrix settings; downstream absolute gains are small, on the order of roughly 0.1–0.5 accuracy points for HellaSwag and roughly 0.1–0.3 points for ARC-Challenge.

The diagonal restriction is a robustness–modularity tradeoff rather than a diagonal-covariance assumption: a dense full-covariance oracle wins its own weighted objective in 80/81 cases, yet diagonal AWT gives better held-out functional fidelity in 53/81 cases. Covariance analysis shows activations are strongly non-diagonal (diagonal covariance mass only 0.0320/0.1415/0.2122; off-diagonal Frobenius ratio 0.9838/0.9255/0.8850), so diagonal AWT gives up optimality on that objective in exchange for the original channel layout, no dense deployment transform, and a lighter statistical estimation burden. Offline operator-level analysis covers 81 cases; by model, full-covariance versus diagonal wins on held-out output error are 13 vs 14 for Llama, 10 vs 17 for Ministral, and 5 vs 22 for Qwen.

Perspective

AWT targets TT/TTN post-training compression pipelines that use a fixed tensorization shape, topology, and rank budget, and suits engineering and research settings that want better functional fidelity without retraining or replacing the solver. It needs a small unlabeled calibration set to estimate per-channel activation statistics, and deployment adds only an input-side elementwise scaling that can be fused into surrounding kernels. Multi-operator suffix replacement shows that relative gains are larger when several operators are compressed jointly and errors accumulate, making AWT a natural default preconditioning step in multi-operator pipelines. The method is also complementary to quantization, pruning, and matrix factorization: one could apply AWT before tensorizing weights and then further quantize the resulting tensor-network cores.

Tensorization quality depends on design choices such as dimension factorization, mode ordering, rank allocation, and TTN topology, which the paper keeps fixed or only partially sweeps; jointly optimizing these with activation-aware preconditioning remains open. The runtime analysis isolates AWT's own overhead, but end-to-end latency depends on optimized TT/TTN kernels and hardware-specific implementations. Calibration statistics come from a small input set and may not capture all downstream distributions. Multi-operator suffix experiments move beyond isolated substitutions, but full-model compression still requires coordinated choices across layers and operators. In addition, some per-operator error tables in the appendix have missing values in the loaded text, so those fine-grained comparisons cannot be fully restated here.

Sources