RASteer erases concepts in diffusion models via retain-aware activation steering, matching or outperforming evaluated activation-steering and weight-editing baselines across backbones and benchmarks
Related research and updatesSynopsis
The work proposes Retain-aware Activation Steering (RASteer), a training-free method that first builds a retain subspace from the concepts to preserve, then uses Retain-Orthogonal Steering (ROS) to remove components aligned with this subspace from the erasure direction, and further introduces Overlap-Adaptive Calibration (OAC) to control, at each layer and denoising step, how much of each shared component is removed based on the overlap between the erasure direction and the retain subspace, thereby better balancing target erasure and concept preservation; experiments on unsafe-content, instance, and artistic-style erasure across multiple backbones and benchmarks show RASteer matches or outperforms the activation-steering and weight-editing baselines evaluated.
Figure 1: Erasure–preservation trade-offs in steering. (a) Collateral damage from original-direction steering (top) and incomplete erasure under full projection (bottom). (b) Fractions of activations above each threshold. (c) Target and mean retained CLIP accuracy for Snoopy erasure.
arXivInterpretation
It introduces RASteer, a training-free method that shifts concept erasure from building an erasure direction mainly from the target concept to explicitly modeling the concepts to preserve, mitigating suppression of non-target content by shared components in the erasure direction. Existing activation-steering methods build an erasure direction mainly from the target concept and adjust activations along it at inference time, whereas RASteer additionally builds a retain subspace from the concepts to preserve, making steering more specific to the target. Method description and experimental statements at the abstract level; reported to match or outperform the evaluated activation-steering and weight-editing baselines on unsafe-content, instance, and artistic-style erasure across multiple backbones and benchmarks.
It proposes Retain-Orthogonal Steering (ROS), which removes components aligned with the retain subspace from the erasure direction so that steering is more specific to the target. Compared with steering directly along the original erasure direction, ROS explicitly reduces components shared with retained concepts through orthogonalization. Mechanism description given in the abstract; specific ablations and quantitative results are not provided in the abstract text.
It introduces Overlap-Adaptive Calibration (OAC), which at each layer and denoising step uses the overlap between the erasure direction and the retain subspace to control how much of each shared component is removed, balancing erasure strength and concept preservation. Addressing the observation that fully removing shared components can weaken erasure, OAC uses overlap as a modulating signal rather than removing all shared components uniformly. Mechanism description given in the abstract; implementation details and hyperparameters for the per-layer and per-step adjustment are not presented in the abstract text.
Across multiple erasure scenarios, backbones, and benchmarks, RASteer matches or outperforms the evaluated activation-steering and weight-editing baselines and achieves a better balance between erasure and preservation. The evaluation spans three erasure categories—unsafe content, instance, and artistic style—and compares against both activation-steering and weight-editing baselines. Experimental conclusion stated at the abstract level; specific metric values, benchmark names, and statistical details are not given in the abstract text.
Perspective
The work targets concept erasure in pretrained text-to-image diffusion models, suited to deployment and governance needs where a target concept (such as a copyrighted style, a recognizable character, or unsafe content) must be removed while preserving the ability to generate other content. The method is training-free inference-time activation steering, so users can apply it without retraining the model; the evaluation described in the abstract covers unsafe-content, instance, and artistic-style erasure and compares against activation-steering and weight-editing baselines across multiple backbones and benchmarks. Its conclusions apply to the settings formed by these evaluated backbones, benchmarks, and concept types.
The abstract gives no specific evaluation metrics, numerical results, backbone or benchmark lists, and does not report separate ablations for ROS and OAC, so it is hard to judge how much each mechanism contributes. How the per-layer and per-step overlap is computed, how the removal proportion is parameterized, and the trade-off curve between erasure strength and preservation quality still need to be confirmed in the main text. The abstract also does not discuss performance on unseen concept types or different diffusion architectures, which can be directions for further observation.
