Skip to main content
Back to timeline
arXivSource publication:

Memento 3 lets a frozen LLM agent clear every public ARC-AGI-3 game at the RHAE ceiling using 44% of human actions, via a natural-language rulebook compiled to executable code

Synopsis

Memento 3 introduces Code as Model: a frozen LLM agent writes its hypotheses about an environment into a natural-language rulebook, compiles that rulebook into executable code for prediction and planning, and revises both through a loop of observation, reflection, rule revision, compilation, and verification, accepting an update only when the LLM judges the code faithful to the rulebook and cell-exact replay reproduces every recorded transition. On ARC-AGI-3 the single-model agent clears every level of all public games, reaches the mean RHAE ceiling, and uses 44% of the human action count; in an Atari Pong case study a learned feedback controller wins all three evaluated episodes with different openings without further LLM calls.

Source-provided article image: Memento 3: Model-Based Recursive Self-Improvement through Reflective Rulebooks
Figure 1 ·

Figure 1: From a rulebook to an executable world model. Numbered rules map to corresponding branches of a simplified transition function. A planner uses the model to predict a route that changes the badge’s pattern and colour, collects a refill, and reaches the matching door within the move budget.

arXiv

Interpretation

Code as Model represents the agent's world model as a natural-language rulebook (world_model.md) paired with an executable realisation (world_model_engine.py): the rulebook carries semantic memory and rule-level revision, while the code carries prediction, replay verification, and planning. Earlier Memento work kept external memory on the policy side: Memento used episodic memory for continual adaptation, Memento 2 read memory writing as policy evaluation and memory reading as policy improvement, and Memento-Skills extended the learning state to reusable executable skills; this work moves reflective learning to the model side, the rules governing how the world behaves. Formalised as a deterministic goal-directed POMDP and a Bayes-adaptive POMDP, with an update operator and a control operator, a five-stage loop (observation, reflection, rule revision, compilation, verification), and an acceptance condition combining an LLM rulebook-fidelity judgement with exact replay that reproduces every recorded transition.

On all 25 public ARC-AGI-3 games the single-model agent clears all 183 levels, reaches the mean RHAE ceiling of 100.0, and uses 7,518 actions, or 0.44 of the 17,135-action human baseline. Compared systems are taken from official ARC Prize scorecards: baseline1 also clears every game but uses 8,347 actions (about 0.49 of the human budget), while NOOA reaches 85.1, OPINE-World 78.4, DreamTeam 38.1, and Continual Harness 20.5, with the latter four leaving games unfinished so their totals reflect incomplete runs. Every baseline number, including per-game RHAE, actions, and levels cleared, comes from each system's official ARC Prize scorecard, as does the human action baseline, so all columns are counted the same way; the agent observes only the action interface and cannot inspect the environment source.

An ablation shows the natural-language rulebook itself contributes: across ft09, ka59, and cn04, one game per difficulty band, the rulebook variant uses fewer actions and fewer agent turns in every band, with lower totals across the three games, while every run clears every level and reaches the RHAE ceiling. The ablation removes the rulebook from the update loop while keeping the harness, replay verifier, plan executor, backbone, and reasoning effort byte-identical, leaving a code world model, so the difference is attributable to the persistent rulebook rather than to other components. The three games represent easy, medium, and hard bands defined by human actions per level; actions count scored environment interactions and agent turns count iterations to completion, unaffected by wall-clock delays.

The population extension keeps multiple rulebook-executable pairs sharing one interaction history and samples a behavioural equivalence class per planning episode; on wa30 with two members, total actions fall from 899 to 597, per-level action efficiency matches or improves on eight of nine levels, and both arms clear all nine levels at the RHAE ceiling. The single-model point-belief approximation commits to one explanation, which can remain self-confirming if its actions never expose its errors; the population reduces the dependence of an entire run on one early stochastic initial model by letting different models guide exploration. Backbone, reasoning effort, and harness are held fixed at their single-model settings so only the number of retained pairs changes; during execution a communication mechanism keeps both members synchronised about which member supplied the active plan, which actions were sent, and what real transitions followed, so both update from the same evidence.

Perspective

The result targets agents that learn a deterministic goal-directed environment through black-box interaction under a limited real-action budget: ARC-AGI-3 levels are modelled as deterministic goal-directed POMDPs, the action interface semantics are initially unknown, and the agent sees neither the environment source nor a language task description. The rulebook and executable persist and are reused across attempts and levels, so later levels introducing new mechanics can still be accommodated; Git memory keeps every rulebook and module version addressable, diffable, and restorable, letting the agent revisit hypotheses it has already tested. The population extension targets cases where a single early model could remain self-confirming, using a shared history so different models guide exploration. The Atari Pong case indicates the same mechanism can support feedback control: the controller is fixed during evaluation and makes no LLM calls, with learning cost measured in emulator frames.

Rulebook fidelity is judged by the frozen LLM as a semantic condition, and the text gives no quantitative characterisation of how stable that judgement is; the population extension is examined only on wa30 with two members, so behaviour at larger member counts is not reported; the ablation covers one game per difficulty band, and the authors note that each Claude Code run incurs substantial time and monetary cost; the Atari Pong evaluation covers three episodes with different openings, and full-frame replay retains residual rendering and paddle-motion errors; the public set is near saturation with stronger frontier models, so public-set completion alone cannot isolate the contribution of a world-model architecture, and model choice and reasoning budget must be accounted for alongside it.

Sources