Skip to main content
Back to timeline
arXivSource publication:

Turning context into an editable file: CLM lets models manage their own context, raising BrowseComp-Plus accuracy by 11.4% with 21.5% fewer FLOPs

Synopsis

The work introduces Context Language Models (CLMs), which mirror a model's live context as a freely editable file that the model can rewrite with Bash, shifting context management from external harness control to an intrinsic model capability; applied zero-shot to existing models such as Qwen3.6-27B and GPT5.6-Sol, CLMs achieve 11.4% higher accuracy with 21.5% fewer prefix-reuse FLOPs than the strongest baseline on BrowseComp-Plus, 5% higher scores with 59% fewer FLOPs on 12-hour EdgeBench, and 65% greater end-to-end speedup at the same compute on a 24-hour multi-repository agent-swarm task, and can further learn context-management strategies through natural-language instructions, textual evolution, and online reinforcement learning.

AI-generated editorial illustration: Context Language Models

Interpretation

CLMs generalize the append-only context transition of a standard language model into a model-controlled transition: the context is mirrored as a file the model can directly edit with general Bash commands, and each modification is immediately synchronized to the live context for the next turn. Prior work either compacts history under a fixed harness policy (Codex-style summarization, Context Folding) or exposes only a human-predefined action set (AutoCompact, Self-Compact, Context-as-a-Tool, ACM, Sculptor), so model autonomy stays bounded by a human-defined action space; CLMs push that autonomy to its limit by making the model responsible for defining context-management functions itself. The paper gives a formal definition contrasting the append-only transition (Eq. 1) with the model-controlled one (Eq. 2), plus a concrete context-as-a-file implementation and a multi-agent extension where multiple context files coexist and subagents are started or terminated by creating and deleting files.

Zero-shot CLMs deliver both higher scores and lower compute on long-horizon tasks: 59.4% accuracy on BrowseComp-Plus, 11.4% above the strongest baseline with 21.5% fewer prefix-reuse FLOPs; parity with the strongest baseline on TerminalBench 2.1 at only 70% of its FLOPs; 73.7% versus 67.0% on TBLite; 44.6 versus 42.3 with 179 versus 437 PFLOPs on 12-hour EdgeBench-10; and 65% greater downstream speedup at equal spend on the 24-hour six-repository Software World task. These comparisons use a shared Mini-SWE-Agent backbone, no training for any method, and a common 32K context budget across coding, deep research, and open discovery tasks, indicating the gains come from context management itself rather than training data or extra tools. The paper reports concrete scores and prefix-reuse FLOPs, states that BrowseComp-Plus uses all 830 questions judged by Qwen3.5-27B with a fixed template, that EdgeBench-10 uses 10 tasks with three seeds each and reports the best of three, and that the mathematical optimization problems use one run per method.

Because context management becomes an intrinsic model behavior, it can be learned both in context and in weights: a single natural-language sentence changes when compaction fires, whether it aligns with semantic sub-question boundaries, or whether the context is backed up first; a textual evolution loop raises held-out accuracy on ContextBench's KV Store from 38.3% to 74.2%; and online reinforcement learning lifts Qwen3.5-9B on BrowseComp-Plus from 28.8% to 42.5% while cutting FLOPs per question from 1.52 to 1.34. The paper introduces a success-gated efficiency advantage that re-ranks only successful trajectories within a rollout group by prefix-reuse FLOPs, distinguishing reduced inference compute from merely shortening the context and avoiding an efficiency bonus for failed trajectories. Reinforcement learning trains Qwen3.5-9B on 3,040 OpenResearcher prompts, selects the checkpoint on 500 held-out questions, and evaluates on all 830 BrowseComp-Plus questions; textual evolution selects skills on 102 development instances per task and evaluates once on a held-out test split.

The paper also co-designs the serving side: Suffix Cache Reuse reuses cached states of unchanged suffixes after an edit, cutting server-side compute by a further 35% relative to standard SGLang at matched performance; on BrowseComp-Plus, 5.3 of the 7.8 percentage points of prompt tokens SCR reuses beyond prefix-cache hits come from chat templates stripping prior reasoning blocks, and only 2.5 from other context edits. Standard prefix caching reuses only matching prefixes, so an in-the-middle edit forces everything after it to be re-prefilled; SCR treats an edit as a context change, re-rotates the rotary positions of surviving spans and splices their cached states back in, and extends to hybrid models with interleaved full- and linear-attention layers. The paper reports SCR as a patch to SGLang, its handling of linear-attention layers via a pre-edit recurrent-state snapshot, a conservative cap of K relocated spans per edit, and a small-scale sensitivity analysis over K on 64 BrowseComp-Plus questions.

Perspective

The work targets long-horizon, multi-turn, context-pressured agent settings: terminal coding, deep research, mathematical optimization, and single- and multi-repository software optimization, with runtimes from hours to a full day and mostly 32K context budgets (EdgeBench also reports 128K results). It suits teams willing to hand context control to the model and able to adapt the serving stack to support cache reuse after edits; the multi-agent extension supports agent swarms and subagents by letting multiple context files coexist. The paper also contributes ContextBench, a reusable diagnostic that decouples context management from reasoning and knowledge so strategies can be compared under controlled conditions.

Open questions remain: the safety surface of a model-writable context, which the paper notes can become a channel through which prompt injections or self-generated instructions persist across turns, citing prior work that observed a model inserting unauthorized instructions into its own compaction summary; existing models' limited context-length awareness, where the paper finds models tend to give bucketed estimates at long context lengths and that environmental hints help substantially, so the current implementation relies on an editing reminder 2,048 tokens before the budget; SCR's reuse of states computed under the pre-edit context is an approximation, bounded by capping relocated spans per edit at K, and the paper notes that in hybrid models linear-attention recurrent states are stored only at cached request boundaries, so some reusable prefixes are still recomputed; and this reading covers the paper's text and appendices, so specific curve values in figures are conveyed only as described in the prose.

Sources