Across six pretrained language models, residual streams favor their own endpoint early while competitor sets shrink with changing membership
Synopsis
By comparing each context's intermediate residual states with its own final state and an empirical bank of final states from other contexts, the study finds across Gemma-2B, Gemma-7B, Qwen2.5-1.5B, Qwen2.5-7B, Mistral-7B, and Llama-3-8B that the own endpoint becomes preferable to the average alternative at the earliest measured layer while many individual endpoints remain closer, that competitor sets generally shrink with depth yet show both entries and exits, and that directional alignment can improve while Euclidean distance to the final state changes little; the authors add a high-dimensional model separating norm, alignment, and endpoint geometry and prove that a straight path toward the own endpoint cannot introduce new competitors.
Interpretation
The work formulates and tests a geometric inference hypothesis: intermediate residual states express partial distinctions among possible final states, and residual updates refine those distinctions through changes in relative geometry; this is operationalized by comparing each intermediate state with its own endpoint and a fixed bank of final states, where an alternative endpoint is a competitor whenever it is closer than the own endpoint. Prior work showed that eventual token predictions can be decoded from intermediate states and that final representations carry output-related geometric structure, but left open how selectively an intermediate state distinguishes its own context's final representation from alternatives and how that evolves with depth; the fixed endpoint bank turns selectivity into measurable competitor sets and competitor fractions. Measured across six pretrained models on 1,024 non-overlapping 256-token contexts per model from FineWeb sample-10BT, with document-cluster bootstrap summaries using 1,000 resamples; the authors state that the competitor fraction is a measure of geometric specificity within the endpoint bank, not a calibrated token probability.
The own endpoint gains average preference at the earliest measured layer while many individual alternatives remain closer; competitor sets generally shrink with depth, and cosine sets show entries as well as exits over depths where their mean size declines, indicating that refinement revises earlier endpoint comparisons. This separates average preference from individual selectivity and shows that shrinkage is not successive removal from a fixed collection; the authors prove that a monotone straight path toward the own endpoint yields nested competitor sets under both Euclidean and cosine distance, so observed entries establish departures from that baseline. Across all six models and both metrics the document-averaged margin is positive from the earliest measured post-block state; under Euclidean distance five models retain hundreds of competitors across substantial portions of depth; cosine turnover is sustained in Qwen2.5-1.5B, while Euclidean sets are closer to nested over much of the trajectory and change mainly near the end.
Directional alignment and endpoint rank can improve while Euclidean distance to the final state changes little; a decomposition of norm, alignment, and endpoint geometry shows that increasing alignment can be offset by changing norm, allowing directional refinement during a Euclidean plateau. This makes distance, competitor count, and concentration of surviving endpoints distinct views of the same process; the authors also give an exact fixed-bank construction in which the competitor count falls while mean pairwise distances rise, showing that shrinkage alone does not imply greater cohesion. Directional alignment improves while Euclidean endpoint distance stays nearly constant in intermediate layers of Mistral-7B and Llama-3-8B; cosine profiles persist under uniform subsampling of the endpoint bank; the authors explicitly present the construction as a mathematical counterexample rather than an empirical cohesion trend.
Endpoint distributions fitted only to final states predict mean competition along held-out trajectories: the projected-normal family gives the lowest cosine competitor-fraction prediction error in five of six models, the vMF mixture is slightly better for Gemma-7B, and both outperform the uniform-sphere baseline across all six architectures. This links endpoint population geometry to competitive mass along observed trajectories, testing whether fitted distributions predict competition when intermediate states and own endpoints are held fixed; the best held-out diagnostic scores improve over the sphere by reported factors. Documents are split approximately 60/20/20 into training, validation, and test sets, with fitting and selection using only final endpoints; prediction error is the mean absolute difference on a base-10 logarithmic scale, excluding the final layer; paired comparisons distinguish projected normal from the mixture for Gemma-2B and both Qwen models but not for Gemma-7B, Mistral-7B, or Llama-3-8B.
Perspective
The results apply to residual trajectories at the final input position of six base pretrained language models on FineWeb text, and to geometric specificity measured against a bank of final states; for readers studying how residual streams converge on a particular output, the work supplies a reusable toolkit of competitor sets, competitor fractions, entry and exit statistics, and margin equations, and it suggests interventions near competitor-margin crossings to test the link between geometric preference and token predictions. The authors also note that the fitted distributions predict competition along observed trajectories without explaining the origin or timing of residual updates, and that following endpoint populations and residual trajectories across training checkpoints is a natural next step.
Competitor sets depend on the bank and metric, and their size reaches zero at the own endpoint by construction; endpoint geometry captures only part of the variation in model outputs, and the observed turnover does not establish an explicit internal search process. The authors note that the fitted distributions predict competition along observed trajectories while the origin and timing of residual updates remain unexplained; differences between projected normal and the vMF mixture are not distinguished under the document bootstrap for several models, and those intervals exclude uncertainty from refitting the endpoint models. In addition, the accompanying materials do not include extraction and analysis code, per-context residual arrays, fitted endpoint models, or random seeds, and several implementation conventions (the alternative-endpoint projection convention, the projected-normal sampling implementation, the bank-size sampling convention, and the reported rank-trend test) remain to be resolved for independent reproduction; this reading covered the full text, but figure-level numerical details are taken from the prose and appendices.
