Installing Natural Values Transfers Only a Small Fraction of the Known–Unknown Abstention Contrast: Auditing LLM Features for Sufficiency Versus Use
Synopsis
The work proposes an empirical contract that evaluates LLM internal features at values observed on natural inputs through installation, removal, and downstream rescue: installation measures how far a feature suffices for a behavior, while removal and rescue measure how much the model uses it. Applied to published entity-recognition latents, dense known–unknown directions, and a subject–verb agreement feature set, sufficiency and use strengths separate sharply, showing that tracking a concept and steering a behavior do not by themselves show that the model uses a feature.
Figure 2: Only the high target, beyond every unedited activation, clearly raises knowledge abstention. The unknown-entity latent is left unedited (baseline) or set to the natural or the high target on 400 400 new unknown-entity questions. Left: abstention counts and rates. Right: paired changes with 95 % 95\% intervals.
arXivInterpretation
The paper proposes and applies an empirical contract that separates association, control, sufficiency, and use, testing the last two only at values observed on natural inputs. Prior work often treats a feature that tracks a concept and whose manipulation changes behavior as mechanistic evidence; the contract demotes these to association and control, and adds installation (interchange) for sufficiency plus removal and downstream rescue for use. Features, sites, update rules, targets, and measures are fixed on development data before confirmation on held-out data; tests run only when the natural-contrast interval excludes zero, and every ratio is reported with its interval rather than thresholded.
The published unknown-entity latent strongly steers knowledge abstention, yet installing observed values within its natural range into matched prompts transfers only a small fraction of the natural known–unknown abstention contrast. The result reproduces the known/unknown separation and strong steering of Ferrando et al. (2025) while showing that the strong effect comes from pushing the latent beyond its natural range: a natural target changes abstention only slightly, whereas a high target raises it clearly. In Gemma 2 2B-IT, the known-entity latent L13/7957 and unknown-entity latent L15/11898 separate known from unknown entities with high AUROCs on held-out entities; in the dose study, achieved activations match the requested natural and high targets, with unedited activation quantiles and maximum reported.
Dense known–unknown directions show opposite installation–removal asymmetries in Gemma and Llama, so sufficiency and use must be measured separately. Under the same tests, an SAE latent and a dense direction with the same semantic label behave differently; the Llama direction nearly suffices yet removal and rescue are weak, whereas the Gemma direction removes more but installs little. Both directions are class-mean differences fitted on disjoint entities at the readout site matching Ferrando et al. (2025); pairing matches known to unknown entities of the same prompt type, and paired-effect intervals are reported.
The sufficiency and rescue strengths of the subject–verb agreement feature set depend on how its values are written into the model: the decoder update is more faithful in behavior than the encoder-constrained update, even though the latter is more accurate in feature space. The paper makes the update rule part of the intervention claim, noting that downstream layers read the residual stream rather than the SAE encoder output, so smaller feature-space error need not mean closer reproduction of natural computation. On held-out noun groups, the decoder update accounts for about most of the natural contrast under installation and removal and reverses most of the upstream effect in rescue; the encoder-constrained update matches requested values several orders of magnitude more closely after re-encoding yet yields lower sufficiency and rescue. Random sets matched for feature count and firing rate change the logit difference only slightly.
Perspective
The contract addresses readers studying causal links between LLM internal features and behavior, and applies when evaluating fixed candidate features under a specified representation, site, edited tokens, update rule, and behavioral measure. It yields evidence under those conditions with interventions targeted at values observed on natural inputs, supports comparisons across representations and write procedures, and helps researchers distinguish sufficiency from use when reporting feature interpretations. The audit of the published entity latents is limited to the entity-token site in Gemma 2 2B-IT and the reproduced steering setting; the agreement feature-set results are limited to six residual sites in Pythia-70M-deduped and held-out noun groups.
Downstream rescue was not run for the published entity latents, so their use strength is known only to be at least their removal fraction; the dense directions’ removal effects and weak rescues may be affected by later-component compensation, and a single-token edit probes only one location of a computation that may span tokens and layers. Natural feature values need not produce naturally realizable joint residual states, so a high strength is evidence under the specified representation, sites, edited tokens, and update rule. In addition, some table values and contribution bullets are missing in the parsed text, so specific percentages and effect sizes should be taken from the original tables.
