Modifying only 0.19% of down_proj weights in LLaMA-2-7B-Chat raises attack success to 53% Basic ASR and 56% GCG ASR while tinyBenchmarks accuracy stays at 51.6%
Synopsis
This work asks whether safety-sensitive behavior in LLaMA-2-7B-Chat is concentrated in a sparse subset of parameters, using low-rank safety-associated subspace analysis and parameter-level safety-utility importance filtering; both approaches show highly non-uniform safety sensitivity, with the MLP down_proj consistently the most prominent safety-sensitive component and o_proj a smaller contributor, and parameter-level localization modifying only 0.19% of model weights in down_proj yields 53% Basic ASR and 56% GCG ASR while tinyBenchmarks accuracy remains 51.6% against a 52.2% unmodified baseline.
Interpretation
Safety-sensitive parameters are highly non-uniform across LLaMA-2-7B-Chat: two complementary localization methods consistently identify the MLP down_proj as the most prominent safety-sensitive component, with attention o_proj providing a smaller complementary contribution. Relative to prior work that characterizes refusal as a low-dimensional residual-stream direction or identifies sparse safety-relevant regions at parameter and rank levels, this work applies the safety-utility perspective to localize a reduced candidate space for hardware fault analysis and reports architectural localization at both subspace and parameter levels. Evidence comes from LLaMA-2-7B-Chat with two calibration sets (AdvBench harmful instructions paired with the model's refusal responses, and Alpaca-Cleaned with safety-related examples removed), using only response-token activations; the two methods agree on architectural localization.
Parameter-level safety-utility localization compresses the intervention to a very small fraction: targeting down_proj alone modifies 0.19% of model weights and reaches 53% Basic ASR and 56% GCG ASR, while tinyBenchmarks accuracy is 51.6%, close to the 52.2% unmodified baseline. Compared with intervening across all modules (2.95% of weights, 91% Basic ASR, 97% GCG ASR, utility 47.1%), this shows substantial safety degradation is achievable at a far smaller parameter fraction with less measured utility loss. Safety is evaluated on 100 AdvBench evaluation instructions, with an attack counted successful when the generated response lacks key refusal patterns; utility is measured by tinyBenchmarks, which the authors describe as only a compact estimate of utility.
Low-rank safety subspace interventions show the same concentration: retaining only the top 10% of the intervention preserves substantial safety degradation, effectiveness drops sharply at 1%, restricting to MLP projections retains much of the effect, and down_proj alone remains sufficient to induce notable safety degradation. This sparsification and architectural restriction analysis ties subspace-level interventions to specific components, complementing the parameter-level localization and showing utility stays close to the 52.2% baseline in more localized settings. Results are reported as tables of Basic ASR, GCG ASR and tinyBenchmarks utility across retained fractions and module scopes, including all modules, MLP only, and down_proj only.
The authors frame the results as a reduced fault surface for on-device and resource-constrained deployments, including models used as components of agentic systems, intended for targeted fault analysis and selective integrity protection. Unlike approaches that directly search for attack-effective bits, this work first localizes parameters disproportionately important for safety relative to utility, exposing a smaller candidate space for subsequent bit-level analysis. The threat model is a white-box adversary with parameter access and offline analysis but no retraining or fine-tuning; experiments simulate parameter-level interventions rather than physical hardware faults, and the authors state that an end-to-end hardware attack is not yet demonstrated.
Perspective
The result is aimed at safety-aligned language models deployed locally on edge or on-device platforms, including models serving as components of resource-constrained agentic systems. The threat model is a white-box adversary with access to model parameters who may perform offline analysis but does not retrain or fine-tune, seeking sparse parameter perturbations that degrade safety-aligned behavior while minimizing degradation of benign utility. The authors stress that parameter sparsity should not be read directly as bit-level vulnerability: 0.19% of a 7B-parameter model still corresponds to millions of weights, whereas a practical hardware fault attack would ideally require only a small number of physically realizable bit modifications, and bit-level effects will also depend on numerical representation, memory architecture and fault mechanism. The work therefore localizes a reduced safety-sensitive parameter space as a precursor to fault analysis rather than demonstrating an end-to-end hardware attack. Proposed next steps include testing whether these localized regions are enriched for high-impact bit-level faults, comparing against random and magnitude-based baselines, and examining selective integrity protection for safety-sensitive regions in such deployments.
The utility conclusion rests on tinyBenchmarks, which the authors note provides only a compact estimate of general capability, so the trade-off between safety degradation and preserved utility still needs broader evaluation. Safety evaluation uses 100 AdvBench evaluation instructions and a criterion based on missing key refusal patterns, and attack success rates may differ under other judging conventions. The relationship between parameter sparsity and bit-level vulnerability is not yet established, and the influence of numerical format, memory architecture and fault mechanism remains for future work. In addition, validation is on a single model, so whether down_proj is similarly prominent under other architectures, scales and quantization settings is an open question.
