Confining self-evolution to the runtime harness behind one admission gate: 144 of 7,449 candidates admitted with no harmful change, while the same loops without the gate admitted 309 harmful changes
Synopsis
The work argues that self-evolution of LLM agents should be confined to the runtime harness—instruction text, tool-call logic and typed primitive composition—while weights stay fixed, and gives a dual-loop engine in which a local loop edits one primitive and a global loop replaces primitives from a typed library, both submitting to one admission gate; on a synthetic regulatory-shift benchmark (three families of re-interpretation, three severities, 10 seeds per cell) the gate admitted 144 of 7,449 candidates with none worsening error on held-out history, whereas replacing the gate with the check an unbounded system applies—fewer errors visible in recent traces—let the same loops admit 309 harmful changes and left missed flags above 10% in 49 of 90 runs.
Interpretation
Proposes and implements a harness-only architecture: the engine may rewrite instruction text, tool-call logic and primitive composition, may not fine-tune, and the served model is pinned by digest; a local loop changes one parameter or one clause of a primitive per candidate, a global loop replaces a primitive from a typed library, and both share one retrospective pool and one gate. The self-evolution literature treats the harness as one surface among several to optimise over; this work treats it as the only surface an institution may let evolve, turning “the agent changed” from an event requiring model re-validation into a text diff with a cause and a test attached. The architecture and admission rule are implemented in code, with the edit space, library and admission rule written into technical documentation; the experiments use synthetic data and neither the agent nor the proposer is a language model.
Measures the admission gate itself: across 90 runs the gate admitted 144 of 7,449 candidates with 0 harmful admissions, and restored the false-positive rate to the oracle level without raising missed flags in every low- and mid-severity cell. An earlier presentation specified the gate but did not measure it; this paper reports candidate counts, rejections by reason, a false-admission audit on held-out data, one worked rejection, and the gate’s own failure mode. Three families of re-interpretation at three severities, 10 seeds each, 90 runs, ten arms, a held-out window of 3,450 alerts per run, means with 95% intervals; the author also notes the held-out false-admission audit is weak because its halves are interleaved from the same stream.
Gives a design commitment on where labels come from: after a re-interpretation the retrospective pool must be relabelled by the new rule, or a correct adaptation is indistinguishable from a regression; on stale labels the gate rejected all 8,490 of its candidates, including the exactly correct harness 470 times. Positions the adaptation as supervised at the level of the rule and autonomous only at the level of the harness, with the relabelling rule authored by the compliance function, so the risk sits in a short, reviewable artefact. The gate-on-stale-labels arm left the pipeline at the degraded false-positive rate in every cell; all 8,490 of its candidates were rejected for lack of significant improvement and all also failed the non-regression clause.
Characterises the division of labour and its cost: parametric and clause shifts are repaired by the local loop alone, a structural shift only by primitive replacement; the dual loop recovers more slowly than the right single loop on the structural family because the local loop must stall before escalating, and it keeps proposing rejected candidates after recovery. The ablation shows escalation finds the right loop without being told the family, and that the gate rejects inexpressible local candidates rather than admitting the least-bad of them. The local-only arm proposed 1,890 candidates across the three structural cells and admitted 0, all rejected for regression; the compute-matched local-only arm did no better in the mid cell (27.4% against 27.4%).
Perspective
The result is aimed at teams in regulated credit pipelines that need to turn agentic adaptation into reviewable changes, especially compliance and model-risk functions: the edit space, typed library and admission rule can be written into technical documentation before deployment, making the set of reachable harnesses and the rule that admits them “predetermined changes” under EU AI Act Article 43(4) and Article 3(23); the gate writes a hash-chained record before deployment, corresponding to record-keeping under Article 12 and change management under Article 17(1)(a). The author states that for a deterministic pipeline auditability and behaviour reproducibility coincide, whereas for a hosted LLM behaviour reproducibility additionally requires pinned model versions and decoding parameters, recorded tool responses, and canary probes on a fixed prompt set. The author also states the bound has a real cost: a task the base model cannot do falls outside the loop and remains engineering work.
Open questions the author lists include: neither the agent nor the proposer is a language model, the proposer is a seeded search over a stated edit space, and the library contains the correct replacement primitive by construction, so the track measures the bound and the gate rather than proposal quality; an LLM proposer will produce candidates outside this space, including free-text edits whose regression surface is larger, and the interface accepts one, which is the next experiment. All data are synthetic, and the distributions, shift families and feedback model (all flags reviewed, 5% of clears audited) are the author’s; the no-gate failure depends on misses being less visible than false positives, and its size is an assumption. Execution noise is modelled as common random numbers; real LLM noise is not shared across versions, which would raise the noise floor of the regression clause and strengthen the case for a noise-scaled tolerance. The held-out false-admission audit is weak by construction because its halves are interleaved from the same pre-shift stream labelled by the same deterministic rule, so zero false admissions is expected for any gated arm and the informative contrasts are against the no-gate arm and on the post-shift held-out window. No static prompt-optimisation baseline or direct implementations of the cited self-evolution methods were run; one primitive and one shift at a time are tested, with interacting primitives and concurrent shifts untested; the relabelled pool is only as correct as the rule that relabels it; the benchmark carries no protected attributes, so there is no fairness claim, and bias monitoring under Article 10 has no mechanism in the design and no evidence in the experiments. In addition, the fixed tolerance rejected the exactly correct replacement in the high-severity structural cell (5 of 10 seeds recovered, 24 rejections of the correct harness), and the author proposes a tolerance scaled to the expected noise on the cases a candidate changes as the next revision, noting that no harmful change was admitted at either tolerance value, so the protective side of that trade-off is not measured by this track.
