Skip to main content
Back to timeline
arXivSource publication:

Trajectory features fail to improve final-layer attention routing: six prespecified comparisons show no positive gain in SmolLM3 and Qwen3.5

Synopsis

Using paired enabled/bypassed executions of the final attention layer to produce signed next-token loss differences in frozen SmolLM3-3B-Base and Qwen3.5-4B-Base, the study tests whether hidden-state extrapolation error, curvature and error change improve attention routing beyond uncertainty, one-step displacement, position and state projections; on 100 held-out PG-19 books at an identical causal 20% invocation quota, none of six prespecified comparisons shows a positive gain after Holm correction, and in Qwen3.5 a parameter-matched fixed-projection control lowers NLL by 0.00356 nats/token relative to the trajectory router.

Source-provided article image: Evaluating Trajectory Features for Routing Final-Layer Attention
Figure 1 ·

Figure 1: Paired final-attention utility and controlled routing. Both branches receive the same preceding-trunk output and retain their corresponding final feed-forward computation. Primary features use the bypass state. Every compared router receives the same signed utility supervision. The prefix-only quota uses frozen calibration thresholds and allocates 99 calls among 495 scored positions. Incoming-state routing is a separate secondary experiment.

arXiv

Interpretation

In frozen SmolLM3-3B-Base and Qwen3.5-4B-Base, enabling final attention versus bypassing it reduces mean PG-19 NLL by 0.03813 and 0.16614 nats/token respectively, yet utility is positive on only 48.04% and 72.06% of scored predictions, showing that average attention benefit alone is insufficient to choose invocation positions. Prior adaptive-computation and memory-selection work largely compares end-to-end architectures or different retention targets; here the backbone and optional operation are held fixed and paired executions on the same prefix yield signed utility, separating whether attention helps from where to invoke it. Two public base checkpoints, two 512-prediction chunks per PG-19 book, 495 scored positions per chunk after excluding the first 17, 99,000 test predictions per checkpoint, and a fixed 20% quota.

None of the six prespecified comparisons shows a positive gain after Holm correction; in Qwen3.5 the parameter-matched fixed-projection control C_RP3 lowers NLL by 0.00356 nats/token relative to the trajectory router J32 (95% interval 0.00218–0.00487), and simultaneous upper bounds for all six contrasts fall below the prespecified 0.005-nats/token practical discussion threshold. This control matches head capacity and input dimension, isolating the incremental value of trajectory features from a simpler current-state projection rather than only from scalar baselines. Features, fits, budgets and six contrasts were locked before modern test utilities were collected; three fitting seeds, 10,000 whole-book bootstrap draws, seed 217, Holm correction.

Secondary results depend on the operation removed, feature location and scoring horizon: local 32-position attention recovers 54.8% and 93.8% of the corresponding full-versus-bypass mean gain in SmolLM3 and Qwen3.5; at a 4,096-prediction horizon, unforced J32 activation rises from 20.46% to 86.86% in SmolLM3 and from 18.71% to 35.55% in Qwen3.5. These decompositions show that removing final attention and withholding only its direct distant reads need separate evaluations, and that frozen thresholds drift substantially at longer horizons, making unforced NLL changes unsuitable as matched-budget evidence. Secondary analyses use midpoint sequences from the same test books and can overlap the core chunks; the authors label them a horizon-shift analysis rather than an independent book cohort.

Actual selected-query execution yields small long-sequence latency reductions with increased NLL: at 4,096 predictions I_S35 has native-relative speed ratios of 1.043 in SmolLM3 (95% paired-block interval 1.013–1.060) and 1.029 in Qwen3.5 (1.024–1.050), with executed NLL increases of 0.00925 and 0.11014 nats/token; in cached continuation all learned routers are slower and have higher NLL. The measurement separates allocation quality from measured inference benefit and discloses that bypass features repeat reusable feed-forward and vocabulary-head work, so measured overhead includes avoidable implementation cost. Of 100 planned policy–conditions, 99 passed preflight and supplied 2,970 timed executions; 30 randomized paired blocks with five warmups; SmolLM3 I_S35 at length 1,024 failed the unchanged 0.001 nats/token tolerance and was not timed.

Perspective

The work addresses researchers and engineers who must decide under a fixed budget whether to invoke final-layer attention, in settings with frozen base checkpoints, teacher-forced prefixes, a known scored horizon and a fixed 20% invocation quota. It supplies a reusable paired-intervention and causal-quota protocol and a measurement framework that separates allocation quality from measured latency; applicability to unknown-length serving, free-generation feedback, end-to-end efficiency with optimized fused kernels, and larger or differently pretrained model families requires separate evaluation.

A careful reader would still watch: the capability limits of rank-32 state projections and a fixed fitting recipe as learners, and whether the primary results rule out conditional independence of all trajectory information; one checkpoint per family and an overlapping horizon-shift cohort cannot separate parameter scale, architecture and pretraining; the deadline quota assumes a known horizon, and its forced final decisions plus observed threshold drift leave unknown-length serving unresolved; teacher-forced fixed prefixes exclude feedback from generated tokens; execution measurements ran under shared Windows desktop graphics with reference PyTorch fallbacks, so absolute production performance and performance with optimal fused kernels remain untested.

Sources