Skip to main content
Back to timeline
arXivSource publication:

Frozen-base logit correction: CRN v2 fixes 53.3% of errors with no MMLU/BoolQ degradation, while LoRA fixes 83.3% but loses 17–75 points

Synopsis

The study introduces CRN v2, a 34M-trainable-parameter (0.73% of the 4.65B text module) logit-level correction module sitting atop a fully frozen Gemma 4 E2B, trained by supervised fine-tuning plus reference-free DPO on 83,400 error-correction pairs, which corrects 53.3% of base-model errors on a 60-question CEHRI domain exam (43.3% on a reworded variant) with no degradation on MMLU, BoolQ, or car-wash, whereas a parameter-budget-matched LoRA baseline corrects 83.3% but loses 30–75% capability on the same benchmarks.

AI-generated editorial illustration: Safe Error Correction for Language Models: Frozen-Base Adjustment with Capability Preservation

Interpretation

It proposes and tests the design principle of frozen base plus logit correction plus KL anchoring: the correction module reads only the base's final hidden state and produces an additive logit correction, and the base receives no gradients. Unlike LoRA, adapters, and prefix tuning, which inject trainable capacity inside the base computation graph, CRN v2 is a side network that touches no base parameter; unlike model-editing methods such as ROME and MEMIT that rewrite weights directly, it avoids weight surgery and keeps the base intact. The paper describes itself as a study of a design principle rather than a claim of architectural novelty, calling the equation a standard low-rank bottleneck; training uses 83,400 pairs from 17,810 unique prompts, 2,000 SFT steps plus 500 DPO steps, about 25 minutes end-to-end on an Apple M4.

It provides measurable evidence for a correction–capability tradeoff: CRN v2 corrects 53.3% on the original exam (32/60, with the frozen base alone answering 7/60) and 43.3% on the reworded exam (52/120), while matching the frozen base exactly at MMLU 125/200, BoolQ 144/200, and car-wash 6/8. On the same benchmarks, a LoRA baseline matched to the CRN v1 budget (6.6M parameters, rank 19) corrects 83.3% (50/60) and 77.5% (93/120) but drops to MMLU 64/200, BoolQ 110/200, and car-wash 0/8, a 17–75 percentage-point degradation. The paper reports 95% Wilson intervals of [40.9%, 65.4%] for 53.3% and [72.0%, 90.7%] for 83.3%; the authors also note small per-task samples (200 each for MMLU/BoolQ, only 8 for car-wash) and that 'no degradation' means only these three tested benchmarks.

The KL preservation term is load-bearing: lowering the KL weight from 0.1 to 0.01 drops correction from 53.3% to 35.0% on the original exam and from 43.3% to 25.8% on the reworded one, while capability preservation is unaffected. The ablation turns 'is the anchoring term necessary' into a measured variable and yields the qualitative pattern that the correction module must stay anchored to base representations to be effective. The authors explicitly caution that this ablation confounds three variables—rank (128 vs 256), KL (0.1 vs 0.01), and DPO steps (500 vs 2000)—so the drop cannot be attributed to KL alone and a single-variable sweep is needed.

An exploratory injection-depth sweep found no configuration exceeding the rank-128 logit result: layer-7 hidden-state injection (1.6M parameters, SFT-only) reaches 50.0%/55.8%, layer 4 drops to 30.0%/28.3%, multi-depth logit correction (35M) reaches only 40%, and longer training (5,000 SFT plus 2,000 DPO) stays at 53.3%; DPO at depth 7 destroys capabilities (MMLU 13%, BoolQ degenerate). These results turn 'does deeper injection help' into an empirical question and give a pattern consistent with the frozen computation graph imposing a ceiling, rather than a proof of a theoretical one. The authors state these are single-run, session-observed results and that only the layer-7 SFT-only row is independently log-verified; depth-4 and DPO rows have removed logs and unretained checkpoints, and capability probes are small.

Perspective

The result targets deployment settings where the base is a frozen language model, the domain is fixed, and preserving general capability matters more than maximizing correction rate; it offers a reproducible engineering path—34M parameters, 0.73% of the base, one forward pass to produce corrected logits—with code, main-result weights, and evaluation scripts released and the deep variant released as code only. For researchers, it turns the correction–capability tradeoff into a comparable experimental setup: side network versus LoRA on the same data and DPO objective. For practitioners, it suggests accepting a lower correction rate where capability loss is unacceptable.

Readers should still watch that all 60 original exam prompts appear verbatim in the training data and the reworded exam is template-based near-duplicates (mean 71.6% character overlap), so neither tests out-of-distribution generalization; the KL ablation confounds rank, KL weight, and DPO steps, with a single-variable sweep still pending; capability benchmarks have small per-task samples (200 each for MMLU/BoolQ, only 8 for car-wash), so finer differences need larger samples; HellaSwag scores 0/200 for both models because of a scoring-parser mismatch, and the paper draws no conclusions from it; autoregressive inference needs one full base forward pass per generated token (no KV cache), about 1 s/token on M4; training data is small and scaling behavior is unknown; and the deep injection variant lacks independent verification except for the layer-7 SFT-only row. Whether 53% is a hard ceiling or a best-achieved point is left open by the paper itself.

Sources