Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

Researchers distinguish faithful from unfaithful self-reporting models via attribution similarity, 0.34 versus 0.08 on average

In a controlled setting, the authors trained low-rank adapters to make binary decisions for fictitious characters according to latent linear preference functions, found that accurate self-reporting of learned preferences emerges from decision-task fine-tuning alone without explicit self-report supervision, and showed that faithful models have significantly higher attribution similarity between the decision and self-report tasks (mean 0.34 versus 0.08; paired-bootstrap 95% CI [0.16, 0.36]), allowing faithful and unfaithful self-report to be distinguished without understanding the report content.