Public articles linked to the same research event.
arXiv In hidden-Markov worlds where pretraining recovers Bayesian belief states, the authors define an exact reward-null kernel from a reward that reads only a coarse function of the hidden state, and separately measure whether reward-null information stays linearly decodable, whether decisions causally depend on it, and how much activation variance it occupies; post-training mostly reroutes or rescales this information while leaving it decodable, a KL anchor keeps decisions using it, unanchored objectives let decisions stop using it, and only prolonged weight decay erases distinctions that neither reward nor next-token prediction can see, with open language models showing the same dissociation in in-context belief geometry.
In hidden-Markov worlds where pretraining recovers Bayesian belief states, the authors define an exact reward-null kernel from a reward that reads only a coarse function of the hidden state, and separately measure whether reward-null information stays linearly decodable, whether decisions causally depend on it, and how much activation variance it occupies; post-training mostly reroutes or rescales this information while leaving it decodable, a KL anchor keeps decisions using it, unanchored objectives let decisions stop using it, and only prolonged weight decay erases distinctions that neither reward nor next-token prediction can see, with open language models showing the same dissociation in in-context belief geometry.
In hidden-Markov worlds where pretraining recovers Bayesian belief states, the authors define an exact reward-null kernel from a reward that reads only a coarse function of the hidden state, and separately measure whether reward-null information stays linearly decodable, whether decisions causally depend on it, and how much activation variance it occupies; post-training mostly reroutes or rescales this information while leaving it decodable, a KL anchor keeps decisions using it, unanchored objectives let decisions stop using it, and only prolonged weight decay erases distinctions that neither reward nor next-token prediction can see, with open language models showing the same dissociation in in-context belief geometry.
In hidden-Markov worlds where pretraining recovers Bayesian belief states, the authors define an exact reward-null kernel from a reward that reads only a coarse function of the hidden state, and separately measure whether reward-null information stays linearly decodable, whether decisions causally depend on it, and how much activation variance it occupies; post-training mostly reroutes or rescales this information while leaving it decodable, a KL anchor keeps decisions using it, unanchored objectives let decisions stop using it, and only prolonged weight decay erases distinctions that neither reward nor next-token prediction can see, with open language models showing the same dissociation in in-context belief geometry.