Skip to main content
Back to timeline
arXivSource publication:

With weights frozen, a single rank-8 LoRA at layer 14 lifts Qwen3-8B from 15.5% to 99% exact accuracy on 24-line reference chains

Synopsis

The work measures how far pretrained transformers follow in-context reference chains: thirteen base models reliably follow only 1.4–3.6 lines (median 2.2) and extra pretrained loops add little, yet with all model weights frozen a task-trained rank-8 LoRA at one early layer extends this computation, raising Qwen3-8B from 15.5% to 99% exact accuracy on 24-line chains, reaching 50 lines with a longer-trained LoRA and, in Ouro-1.4B, 60 lines after four loops and at least 160 after eight, via a relay in the middle layers where frozen heads read progressively further up the chain and removing parent-line attention stops it.

AI-generated editorial illustration: Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It

Interpretation

The default reference-following computation is short: thirteen base models reliably follow 1.4–3.6 lines (median 2.2), Qwen3-8B scores 83.5%, 54%, and 41.5% choice accuracy at three, four, and five lines, DeepSeek-V4-Flash (292B MoE) reaches 4.0 lines but scores 42% at six and 38% at eight against 33% chance, Ouro-1.4B's reach is 0, 1.6, 2.3, 2.2, and 2.2 lines after one through five loops, and Huginn reaches only 1.4–1.6 lines after 8, 16, or 32 recurrences. Prior work on multi-hop composition and recurrent depth asked whether models can compose facts or whether loops make inference computation adjustable; this work frames the issue as how little depth the default pass uses, and measures it with one reach and read-out protocol across standard and looped models. A panel of thirteen base models plus DeepSeek-V4-Flash, with 200 programs per cell in the standard survey and typically 150 in looped evaluations; reach is the first downward crossing of 80% accuracy, linearly interpolated.

With every model weight frozen, a single early-layer rank-8 LoRA extends chains substantially: Qwen3-8B's layer-14 LoRA trains 65,537 parameters (under 0.01% of the model) and raises exact accuracy on 16-, 20-, and 24-line chains from 15.5%, 14.5%, and 15.5% to 98.5%, 98.0%, and 99.0% (24 lines: 198 of 200, 95% Wilson interval 96.4–99.7%); a longer-trained LoRA answers 40- and 48-line chains at 98% and 88% versus 4% and 9% frozen, reaching 50 lines. Unlike training recurrent models for hop extrapolation or supervising one hop per loop, this leaves pretrained weights untouched and uses one low-rank representation edit to release computation already present in the frozen layers. Headline evaluations use 200 programs per cell, three training seeds score 94.7–97.5% at 24 lines, and WikiText-103 perplexity moves from 10.14 to 10.148 as a measured text-distribution control.

In looped models the LoRA makes extra loops useful: Ouro-1.4B's standard LoRA reaches 1.9, 7.1, and 17 lines after one, two, and three loops and answers every tested length through 24 after four; the longer-trained LoRA reaches 60 lines after four loops, about 146 after six, and at least 160 after eight (87% accuracy at 160 lines, four times the longest training chain); Huginn's longer-trained LoRA reaches 31 lines after eight recurrences and 58 after sixteen (95% interval 37–64). Pretrained looped models previously repeated a short computation with diminishing returns per loop; here the loop bodies can carry a longer relay once it is started at the right place. Reach estimates carry 95% parametric-bootstrap intervals, with typically 150 programs per cell and 100 or 60 for long chains and Huginn; a second longer-trained seed reaches 52 rather than 60 lines after four loops.

Mechanistically the LoRA starts a relay in the middle layers: Qwen3-8B's relay reaches lines 5, 6, 6, 8, 9, 13, and 16 at the outputs of layers 16–22 while the frozen model reaches line four at layer 17 and stops; attention reads further as the relay advances (two lines up at layers 16–18, then three, four, six, and seven at layers 19–22), and removing each pointer's attention to its parent in layers 14–22 returns six-, eight-, and twelve-line chains to chance (53%, 48%, 55%) whereas the same cut after the relay in layers 23–29 leaves 100%, 100%, and 98%. This moves the explanation from correlation to causal intervention: the relay runs in the frozen middle layers, and the LoRA itself exchanges no information across tokens (applying it only to program tokens retains the gain, while applying it only at the query adds just two to four lines). Combines causal tracing with clean-state restoration, attention knockout, and linear read-outs; in Ouro, ablating seven longer-reading heads lowers four-loop reach to 12.0 versus 12.0–17.4 for twenty layer-matched random-head draws.

Perspective

This work speaks to readers studying reference following and multi-hop composition: it offers a reproducible synthetic task, a reach and read-out measurement suite, and a recipe that starts a relay under frozen weights, suited to settings where in-context composition must be extended for a specific task format. For adaptation practitioners, the takeaway is that placement decides which useful computation can still follow a change: moving Qwen3-8B's LoRA from layer 20 to 21 drops reach from 20.5 to 5.2 lines, and in Ouro a last-loop-only LoRA reaches 21.6 lines at layer 6 but 2.6 at layer 20 (frozen 2.5). For inference, repeating the useful middle layers also helps: re-entering Qwen3-8B's layers 14–22 raises 64-line exact accuracy from 34% in one pass to 66% with one re-entry and 92% with two (frozen 10%). On MuSiQue, early-layer LoRAs add 11.4, 9.4, and 17.9 exact-match points in Qwen3-8B, OLMo-3-7B, and Llama-3.1-8B, and early projection LoRA preserves most of the all-layer gain.

The mechanistic tasks are synthetic and in-context, and MuSiQue provides evidence about placement rather than direct evidence of the same mechanism; detailed traces focus on Qwen3-8B and Ouro-1.4B, the looped models cover two families with Ouro-2.6B grown from Ouro-1.4B, LoRAs are trained per task format, the text penalty is a measured control, placement predictions have modest precision, first-loop-only sufficiency is established only through about 25 lines, and the longest-chain looped results use level order with a LoRA in every loop. The learned routing directions and minimum sufficient rank remain unidentified, so the causal tests constrain the computation without uniquely specifying its algorithm; small from-scratch models offer a further entry point, matching the relation between attention distance and relay progress exactly in 108 of 165 layer steps.

Sources