Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

Post-training slices a language model's belief state into a causal quotient: reward-null information is mostly rerouted or rescaled, and only sustained weight decay erases it

In hidden-Markov worlds where pretraining recovers Bayesian belief states, the authors define an exact reward-null kernel from a reward that reads only a coarse function of the hidden state, and separately measure whether reward-null information stays linearly decodable, whether decisions causally depend on it, and how much activation variance it occupies; post-training mostly reroutes or rescales this information while leaving it decodable, a KL anchor keeps decisions using it, unanchored objectives let decisions stop using it, and only prolonged weight decay erases distinctions that neither reward nor next-token prediction can see, with open language models showing the same dissociation in in-context belief geometry.