Public articles linked to the same research event.
arXiv In a controlled setting, the authors trained low-rank adapters to make binary decisions for fictitious characters according to latent linear preference functions, found that accurate self-reporting of learned preferences emerges from decision-task fine-tuning alone without explicit self-report supervision, and showed that faithful models have significantly higher attribution similarity between the decision and self-report tasks (mean 0.34 versus 0.08; paired-bootstrap 95% CI [0.16, 0.36]), allowing faithful and unfaithful self-report to be distinguished without understanding the report content.
In a controlled setting, the authors trained low-rank adapters to make binary decisions for fictitious characters according to latent linear preference functions, found that accurate self-reporting of learned preferences emerges from decision-task fine-tuning alone without explicit self-report supervision, and showed that faithful models have significantly higher attribution similarity between the decision and self-report tasks (mean 0.34 versus 0.08; paired-bootstrap 95% CI [0.16, 0.36]), allowing faithful and unfaithful self-report to be distinguished without understanding the report content.
In a controlled setting, the authors trained low-rank adapters to make binary decisions for fictitious characters according to latent linear preference functions, found that accurate self-reporting of learned preferences emerges from decision-task fine-tuning alone without explicit self-report supervision, and showed that faithful models have significantly higher attribution similarity between the decision and self-report tasks (mean 0.34 versus 0.08; paired-bootstrap 95% CI [0.16, 0.36]), allowing faithful and unfaithful self-report to be distinguished without understanding the report content.
In a controlled setting, the authors trained low-rank adapters to make binary decisions for fictitious characters according to latent linear preference functions, found that accurate self-reporting of learned preferences emerges from decision-task fine-tuning alone without explicit self-report supervision, and showed that faithful models have significantly higher attribution similarity between the decision and self-report tasks (mean 0.34 versus 0.08; paired-bootstrap 95% CI [0.16, 0.36]), allowing faithful and unfaithful self-report to be distinguished without understanding the report content.