Researchers distinguish faithful from unfaithful self-reporting models via attribution similarity, 0.34 versus 0.08 on average
Related research and updatesSynopsis
In a controlled setting, the authors trained low-rank adapters to make binary decisions for fictitious characters according to latent linear preference functions, found that accurate self-reporting of learned preferences emerges from decision-task fine-tuning alone without explicit self-report supervision, and showed that faithful models have significantly higher attribution similarity between the decision and self-report tasks (mean 0.34 versus 0.08; paired-bootstrap 95% CI [0.16, 0.36]), allowing faithful and unfaithful self-report to be distinguished without understanding the report content.
Figure 1: Overview of our setting. (a) We train a model to make decisions according to 100 different latent linear preference functions (here, 𝒑 ∈ ℝ 5 \bm{p}\in\mathbb{R}^{5} ), each associated with a different character (here, “Gregor Samsa” choosing between washing machines). At test time, we present the model with multiple decisions, and use logistic regression to recover a preference vector ( 𝒑 ^ \hat{\bm{p}} ) from its choices. We also ask the model to report on its decision making process, and average the results ( 𝒑 ~ \tilde{\bm{p}} ). (b) This allows us to measure the model’s decision performance ( corr ( 𝒑 ^ , 𝒑 ) \text{corr}(\hat{\bm{p}},\bm{p}) ) and faithfulness ( corr ( 𝒑 ^ , 𝒑 ~ ) \text{corr}(\hat{\bm{p}},\tilde{\bm{p}}) ). We observe an example of delayed generalization: at checkpoint 1000 , the model has relatively high (0.82) decision performance, but low (0.25) faithfulness. By checkpoint 3000 , the model’s decision performance has improved moderately to 0.92, but its faithfulness has risen sharply to 0.83, even with no explicit training for self-report (see Section 3 ). These two checkpoints are our unfaithful and faithful models. This allows us to ask: are there structural changes to the model’s internals that accompany the emergence of faithful self-report (Section 4 ), and can they tell us whether that self-report is grounded (Section 5 )?
arXivInterpretation
Faithful self-report emerges from decision-task training alone and lags well behind decision performance. Prior work observed behaviorally that models can describe implicitly learned internal states; this work constructs contrast pairs of faithful and unfaithful models on the same task: on the 32B model, checkpoint 1000 shows decision performance 0.82 with faithfulness about 0.25, while checkpoint 3000 shows decision performance 0.92 with faithfulness jumping to 0.83, all without explicit self-report supervision. Rank-8 LoRA training on 100 characters across Qwen3 models from 0.6B to 32B; faithfulness is measured as agreement between the self-reported preference vector and a behaviorally inferred preference vector from logistic regression, with the contrast pair drawn from different checkpoints of one training trajectory.
Faithful self-report is accompanied by a measurable structural change: preference representations shift to earlier layers over training. Weight ablations show the faithful checkpoint responds to ablations 5-6 layers earlier than the unfaithful checkpoint; layer-freezing experiments on the 40-layer Qwen3-14B show that training only the first 10, 15, or 20 layers (centered around layers 10-20) yields markedly more faithful self-report than training all layers, and a reversed freezing experiment rules out trainable parameter count as the explanation. Ablation and freezing experiments on Qwen3-14B and 32B, with the training sweep and layer-ablation analysis repeated on Gemma-4 E4B and 31B, where the tested faithful 31B adapter shows the same early-layer localization.
Attribution similarity serves as a mechanistic criterion for faithful self-report without requiring understanding of the report content. Prior diagnoses of self-report were behavioral or required a known target behavior; this work proposes cosine similarity between attribution vectors on the decision and self-report tasks as a measure of shared computation, yielding a mean of 0.34 for faithful models versus 0.08 for unfaithful models across 32 adapter pairs, with a paired-bootstrap 95% CI of [0.16, 0.36]. Node attribution patching with integrated gradients scores LoRA contributions, and cross-task weight-activation interventions provide causal validation: activating the most-attributed weight matrices recovers a larger fraction of the full adapter's effect for faithful self-reporters, even at high sparsity.
Layer-wise attribution profiles also differ between faithful and unfaithful models. When attributions are summed by layer, three of the four combinations peak at the same layer (38); the exception, unfaithful adapters on the decision task, peaks 11 layers later (49), indicating that faithful adapters align decision and self-report at similar layer positions. Layer-wise aggregation and visualization of attributions over the same 32 single-character adapter pairs, serving as a mechanistic complement to the attribution-similarity result.
Perspective
The result applies to a controlled linear-preference setting: preference functions are linear, five-dimensional, and explicitly constructed; the models are LoRA adapters on the Qwen3 and Gemma-4 families rather than full models; and training and evaluation use separate context windows. Within this scope it provides a reusable experimental scaffold for studying honest self-report: constructing faithful and unfaithful contrast pairs from different checkpoints and using attribution similarity as a measure of shared causal mechanism. For settings where model self-testimony must be assessed, such as outputs too long or complex for humans to understand, described behaviors too rare to observe reliably, or claims about purely internal reasoning, the work suggests a detection approach that does not depend on report content. The method targets group-level discrimination, suited to comparing training or deployment configurations of two model populations rather than settling a verdict on a single model.
The attribution-similarity distributions overlap (SD 0.26 for faithful models and 0.10 for unfaithful models), so high similarity is sufficient evidence for faithfulness while low similarity is not strong evidence against it, and individual-level judgments warrant caution. The reason for preference migration to earlier layers is not explained, and the negative faithfulness scores in the 4B and 8B models remain unexplained. The authors offer one interpretation: if attribution similarity measures grounding rather than faithfulness per se, a faithful adapter with low similarity might be a correct confabulator, accurately self-reporting through an ungrounded mechanism. A complementary within-backbone analysis across 100 characters finds only a weak positive relationship, which the authors treat as a sign of life rather than a standalone finding; the reasoning-mode evaluation is a single run without uncertainty bands. The controlled task uses linear, five-dimensional, explicitly constructed preference functions, and whether this generalizes to complex, unverifiable reports remains open.
