Skip to main content
Back to timeline
arXivSource publication:

Post-training slices a language model's belief state into a causal quotient: reward-null information is mostly rerouted or rescaled, and only sustained weight decay erases it

Related research and updates

Synopsis

In hidden-Markov worlds where pretraining recovers Bayesian belief states, the authors define an exact reward-null kernel from a reward that reads only a coarse function of the hidden state, and separately measure whether reward-null information stays linearly decodable, whether decisions causally depend on it, and how much activation variance it occupies; post-training mostly reroutes or rescales this information while leaving it decodable, a KL anchor keeps decisions using it, unanchored objectives let decisions stop using it, and only prolonged weight decay erases distinctions that neither reward nor next-token prediction can see, with open language models showing the same dissociation in in-context belief geometry.

Source-provided article image: Erased, Rerouted, or Rescaled? Post-Training and the Causal Quotient of a Language Model's Belief State
Figure 1 ·

Figure 1: From a belief state to its fates. I. A hidden Markov process with a task and a nuisance variable generates tokens, and a pretrained transformer represents the belief state b t b_{t} . The reward reads only the class marginal b κ b^{\kappa} (large simplex). Beliefs that share it form a reward-null fiber (small simplices), with one direction the next token can see and one it cannot. Each card shows one fiber in the representation (plane) and the decisions its points map to (track; dashed: the base model): erased (collapsed), rerouted (intact, every point mapped to one decision) or rescaled (shrunk, mapped to the same decisions as before), with the fate’s signature in R R , U U and A A . II–IV. Three results in preview (open circles in III: long training with weight decay).

arXiv

Interpretation

The paper separates representation compression into three separately measurable fates: erasure (decodability falls), rerouting (still decodable but the decision no longer causally depends on it), and rescaling (still decodable and still used, but occupying less variance), with a closed-form reward-null kernel giving the criterion when the reward reads only a coarse function of the hidden state. Prior geometric and spectral studies of post-training could observe that representations changed but could not say whether the change was information loss, a change in which directions the policy reads, or an invertible rescaling; the reward-null kernel, computable before training, separates the three. In exactly solvable hidden-Markov worlds, four worlds and several action designs are measured with probes, twin-exchange interventions and effective-rank statistics along the three coordinates, with subspaces, criteria and predictions pre-registered.

The KL anchor decides whether decisions keep using reward-null information: the anchored objective preserves the reference policy's log-odds among equally rewarded outputs, trained networks reach that optimum at practical coefficients (slope 0.93 to 1.02), while supervised fine-tuning and unanchored policy gradient let the decision stop reading a component that remains decodable. It derives what the known KL-regularized optimum implies for the belief state, verifies it in trained networks, and shows what replaces it when the anchor is removed. At matched reward KL-RL and SFT reach the same reward (0.691 and 0.693) yet respectively keep and lose decision sensitivity; without the anchor policy gradient collapses each class onto one token while the next-token-visible component stays decodable.

Genuine erasure needs pressure beyond the objective: under a stress test of weight decay 1 over 50,000 steps, SFT and fresh-token objectives take the next-token-invisible component to 0.001 to 0.012 of the base, while the anchored policy under the same pressure keeps the visible component intact, still reproduces the reference's log-odds, and retains 0.64 of the invisible component. It shows that the reward's indifference alone erases nothing, and that the line between kept and erased follows whether any term of the objective needs the distinction. Erasure is late (under SFT 0.69, 0.30 and 0.001 after 10, 30 and 50 thousand steps), and a recurrent model erases along the same line as the transformer in the weight-decay corner.

Open language models reproduce the dissociation: across three pairs of Qwen base and post-trained checkpoints, peak decodability of the belief state and of the next-token-invisible component changes by at most 0.05 and 0.06, while the effective rank of the later half of the network falls by 12 to 19%; in controlled coarse-class GRPO the anchored run returns to the reference's log-odds within 1,000 steps and the unanchored run amplifies them ten to twenty times with no collapse. It carries the three-coordinate framework and the role of the anchor from exact worlds to pretrained language models on in-context synthetic streams, and shows small transformers and language models leave the anchor identity in opposite directions. The result holds across three checkpoint pairs, two model sizes, several parameterizations and training lengths, with both kernel components staying decodable in every run.

Perspective

The framework applies to hidden-Markov worlds where pretraining recovers the Bayesian belief state and the decision is one-step with the action not fed back into the context, so the reward-null kernel can be computed in closed form before training and the three coordinates measured separately. For a reader, it supplies a diagnostic checklist: when assessing whether post-training has 'forgotten' something, check decodability, causal sensitivity of the decision, and variance share separately rather than reading spectra or effective rank alone; for reinforcement-learning recipes that use a KL anchor, the anchor coefficient decides whether the decision keeps reading the reference's preferences among equally rewarded outputs, while pressures outside the objective such as weight decay decide which components are actually erased. The language-model conclusions apply to the tested 1.7B and 4B checkpoints and in-context synthetic-stream setting.

In multi-step tasks the policy also changes which histories are visited, and whether that moves the causal quotient or only the support of the belief state remains to be studied. The language-model experiments cover three checkpoint pairs from one model family and one controlled objective, and the size of either effect on natural tasks is unmeasured. Represented information is measured with linear probes, so a distinction surviving only in a nonlinear code is counted as lost. Small transformers and language models leave the anchor identity in opposite directions, and neither parameterization nor training budget accounts for the difference; what makes amplification the cheaper direction in a pretrained language model is open.

Sources