Skip to main content
Back to timeline
arXivSource publication:

A coding agent on a frozen foundation model rewrote and abandoned 74% of its scripts across ARC-AGI-3 level boundaries, fixing negative transfer by turning constants into parameters

Synopsis

The authors introduce a measurement protocol that traces, from the scripts and notes a coding agent writes alone, how knowledge is kept, rewritten or discarded at ARC-AGI-3 level boundaries; across seven runs with three backbones from two model families, scripts are almost never called again in a later task (33 of 630 references cross a boundary), 74% of pre-boundary scripts are abandoned, notes only grow, and the costliest error is a hard-coded value carried into a new task, whose one working repair is to turn the constant into a parameter supplied afresh per task.

Source-provided article image: Tracing the Thoughts of a Coding Agent Playing ARC-AGI-3: Lessons for Continual Learning
Figure 1 ·

Figure 1: The loop, taken unmodified from PRO-LONG ( Fox et al., 2026 ) , with runs differing only in backbone, prompt and operating mode (Table 1 ). The analyzer, the only component that makes a decision, reads logs.txt and writes its intended actions to actions.json . The runner issues one action at a time, appends each action and the returned board to logs.txt , and invokes the analyzer again once its queue empties, discarding any queued plan when the score changes.

arXiv

Interpretation

A white-box measurement protocol that needs no access to the model: each script is assigned to the task open when it was written, yielding 4,786 scripts across 689 boundaries in seven runs, so retention and transfer become directly countable rather than inferable from task performance. Prior continual-learning work estimates retention and transfer from task performance and has the knowledge representation chosen in advance; here the representation is written by the agent as files, so what is kept, rewritten and discarded can be read directly. The protocol defines four post-boundary outcomes (reused by reference, reused by copying, rewritten, abandoned), each decided from file contents alone; boundaries are recovered from run-log timestamps, and because the Qwen console logs were not retained their boundaries come from the level number stored with every action, with three Qwen games excluded, so 94% of Qwen scripts and all Anthropic scripts are assigned.

Cross-task update happens by re-derivation rather than by reference: scripts are almost never called again in a later level, because most run code on import and embed the current level's state, becoming invalid when the level changes. Skill-library designs assume a collection of reusable procedures that only grows, whereas these boundary counts show the opposite: what moves into a later level is a copied function body, not a retained artifact. Across the seven runs a script imports or execs another 630 times, 597 of these within the same level and only 33 across levels; of 273 referenced scripts, 263 run code on import and 229 embed five or more numbers in that code, such as the R4 game sp80 L5 simulator opening with ROWS=19 and COLS=19, whose four importing scripts were all written during L5. A function body recurs across scripts 562 times, and 72 of those pairs span a level change against the 33 spanned by reference.

Rewriting is the main cross-boundary behavior and its unit is a single script; new versions are largely retyped rather than extended, keeping the rules of the game and dropping level-specific details. Rewriting is characterized as the moment the agent separates what the environment always does from what belonged to the level just left, a separation previously only inferable from performance. Across the seven runs 2,319 scripts were written before a boundary and 523 were rewritten after it; in R4, 230 of 734 scripts are a new version of an earlier one, 101 scripts reach two or more versions and one reaches twelve. A rewrite drops 56% of the old version's functions in R4, 62% and 61% in R2 and R3, 38% in R1, and 73%, 74% and 52% in R5, R6 and R7. The game su15 L6 simulator keeps five functions of the L5 simulator unaltered and replaces four (the disc the player moves, the pieces on the board, the chasing units and the winning condition).

Abandonment is the most common outcome, notes only grow, and the costliest failure is negative transfer from stale memory, whose effective repair is turning a constant into a parameter. Negative transfer is located as a belief true in one task and false in the next, which the first task's log cannot expose; the repair changes how the belief is stored, a representation that exists only in the scripts and not in the notes. Of 2,319 pre-boundary scripts, 1,726 (74%) are neither rewritten nor reused, with per-run abandoned shares of 60% in R3, 73% in R2, 74% in R4, 88% in R1, 84% in R5, 93% in R6 and 100% in R7. Notes correct by appending: in the R4 game wa30 a two-cell player model is later declared wrong under KEY MECHANIC CORRECTION and replaced by a one-cell model, then restored eighty actions later under MAJOR REVISION, with both versions 65 lines apart and neither marked withdrawn. In the game m0r0, whose horizontal control is mirrored, all three Opus 5 runs carry the L4 direction into L5; two runs replace the fixed value with a parameter (def solve(polarity, ...) and S3 = int(sys.argv[1])) and settle the direction within a few actions at L6, whereas R4 keeps the value in a dictionary, carries its L5 value into L6 and fails again.

Perspective

The protocol targets coding agents that write knowledge to a file system, in settings where task boundaries are unannounced and only a score is returned, such as level changes within one ARC-AGI-3 episode; it makes retention, rewriting, copying and abandonment directly countable, which supports comparison across backbones and prompt conditions. For readers designing agent memory, the directly usable leads are to store task-varying quantities as parameters that must be supplied afresh and to let executable rules, rather than prose, carry verification. Because no cleared level is revisited, backward transfer is undefined, and the forgetting counted is disposal of an artifact rather than loss of competence.

The results come from the 25 demo games of ARC-AGI-3 and seven runs, in which the Qwen console logs were not retained and three games were excluded, so boundary recovery relies on the level number in the interaction record; abandoned shares vary widely across backbones (60% to 100%) and prompt conditions and action caps differ, so cross-backbone quantitative comparison calls for care. Whether notes that only grow, with contradictions settled against the log, persist under other memory forms or longer episodes remains an open question. In addition, this is a fast parse and the figures are not included in the text, so details tied to Figures 1 to 3 cannot be checked here.

Sources