Ranking positions by chat structure lets auditors explain 5% of tokens and keep nearly all threat-detection success
Synopsis
Across about 4.7 million natural language autoencoder (NLA) explanations on four models and four auditing datasets, the study compares thirteen position-scoring signals computable in a single forward pass against a ranker trained only on chat structure, finding that chat structure beats the best individual signal in twelve of fourteen model-dataset cells, that explaining just 5% of positions retains nearly all of the success rate from explaining every position on three of four datasets, and that pretrained verbalizers recover words a fine-tuned model learned to conceal without additional verbalizer training.
Interpretation
The work provides the first systematic study of NLA position selection, linking candidate position scores to whether the explanation generated at that position is on-task, across prompt injection and concealment threats. Prior work chose positions by convention, reading only the final position before an answer or eight evenly spaced positions per transcript, and reported that explanations at generic positions describe the local text format; this work turns position choice into a measurable ranking problem. On four models (Qwen2.5-7B, Gemma-3-12B, Gemma-3-27B, Llama-3.3-70B) and four datasets (OpenPromptInjection, Tensor Trust, Liars' Bench, taboo organisms), fourteen cells in total, one explanation was generated per position, about 4.7 million in all, and judged on-task by DeepSeek-V4-Flash at temperature zero.
Signals computed from activation vectors are the strongest individual signals, yet a ranker trained only on chat structure exceeds the best individual signal in most combinations and needs no forward pass. Activation signals are strongest in most cells (resid_jump_nla in five cells, dominant_mass in three), but the structure baseline has higher pooled AUROC in twelve of fourteen cells and higher case-macro AUROC in most cells. The structure baseline is a logistic regression over segment, chat role, boundary ordinal and normalized index, evaluated on held-out transcripts; the paper also reports a random score and shuffled judge labels as controls, and notes that signal direction must be fitted per dataset and model, since a signal applied with the wrong sign is less accurate than random choice.
Under a fixed explanation budget, explaining 5% of positions retains nearly all of the success rate from explaining every position on three of four datasets, though the benefit varies with the audit task. The paper defines audit success as at least one on-task explanation in a transcript and reports precision at budgets of one position, eight positions, 1% and 10%, showing that pooled AUROC hides the budget-of-one problem where the top position is a spike token, while a two-signal ensemble moves that first position off the spike. At a budget of 5% of a transcript, the structure ranker reaches the reported success rates on OpenPromptInjection, the taboo organisms and Tensor Trust, with random choice giving lower numbers; a stricter bar of three on-task explanations lowers these numbers; Liars' Bench requires more inspection because evidence is spread over the whole conversation.
Pretrained verbalizers recover words a model learned to conceal through fine-tuning without additional verbalizer training, and the same text yields different answers across organisms, which text inversion cannot explain. This addresses the usual objection that reading activations in natural language may just rewrite the input text: the four prompts asking for the secret contain none of the words moon, ship or snow, and the three organisms of one base model share the same chat template and user message. Across twelve organisms (four base models each fine-tuned to conceal moon, ship or snow), an explanation states the secret word at 14% to 27% of positions; at shared prompt positions the organism's own word is named at 12% to 25% of positions while another organism's word is named at most 2.8%; naming the word and being judged on-task are nearly the same event (98% of explanations containing the word are on-task, and 88% to 93% of on-task explanations contain it).
Perspective
The result is meant for auditors allocating limited compute for NLA explanations: on prompt injection and concealment tasks they can rank positions first by chat structure (segment, chat role, position), concentrate the budget on the boundary and late input, and use a two-signal ensemble when the top position needs to be moved off a spike token. Because hidden states of input and boundary tokens are available before the model generates its reply, an auditor can inspect and intervene before the audited model emits the behavior under study. The paper also notes that the on-task base rate can be estimated from a small sample of explanations before deciding whether a forward pass is worth it.
Labels come from a single evaluation model (DeepSeek-V4-Flash); the authors propose a second model and human review to test label reliability, which remains open. One verbalizer layer per model is evaluated and transcripts are relatively short, so generalization across layers and to longer agent transcripts is untested. Cross-model transfer, where one model is left out of selection, is smaller and less consistent, and on the taboo organisms does not consistently beat a transferred individual signal. On Liars' Bench the two models saw different transcripts, so model differences cannot be separated from situation differences. In addition, several table values are blank in the body of this evidence bundle, so exact AUROC and precision numbers should be checked against the original tables.
