MirroS, Tsinghua and Peking University team present AgentGarten: code worlds become real-time interactive environments via a shared neural renderer, with agents learning shelter-building in 4 rounds and ramp use in 10
Synopsis
AgentGarten couples simulators or game engines with a shared real-time neural renderer: the engine maintains persistent state and executes program-defined interaction rules, while the renderer, distilled by Adversarial Forcing into a four-step block-causal model, turns engine-exported depth or normal conditions into visual observations in real time; in hide-and-seek, pretrained agents acting only on rendered frames built shelter with panels by round 4 and crossed walls with a ramp by round 10, and improved scores across rounds in four further worlds through written playbooks.
Interpretation
A code-world framework in which a scene program runs in a simulator or game engine that maintains persistent state and executes explicit interaction rules, while a shared neural renderer generates visual observations from structured conditions exported through a unified interface, so rendering errors cannot accumulate in the state and a recorded state trajectory can be re-rendered under a different appearance reference or camera. Unlike video world models whose state resides in generated history, the transition takes no rendered observation as input; unlike simulators that require 3D asset creation and complex rendering pipelines, new environments can be authored in code without bespoke visual assets per scene. The paper gives a formal formulation (Eqs. 1 and 2) and specifies colorized depth or surface normals as conditions, both exportable by an engine or estimable from real video (ViPE, Depth Anything 3, NormalCrafter).
Adversarial Forcing, a distillation method that converts a pretrained bidirectional video model into a few-step, block-causal real-time renderer: distribution matching on self-rollouts, exact replay so later losses update history encoding, and a real-data adversarial objective with exact R1/R2 regularization. Exact replay runs block by block with SDPA and is bitwise identical to the sampled trajectory, whereas SGF-style full-sequence FlexAttention replay shows 3.99% relative error; exact R1/R2 exploits the frozen backbone to avoid double-backward through fused attention kernels, without random-perturbation approximation. Table 1 reports zero (bitwise) replay error for block-by-block SDPA, forward time falling from 2.95 s to 2.79 s, and a 0.2% change in peak memory; Figure 7 contrasts Self Forcing rollouts that develop repetitive surface patterns and lose detail with Adversarial Forcing retaining natural textures.
Real-time inference on a single NVIDIA H100: hand-written Triton kernels fuse elementwise operations around attention, CUDA graph capture removes launch overhead, and an optional distilled tiny decoder reconstructs pixel frames in under 10 ms, giving over 35 frames per second overall. It avoids torch.compile to sidestep minutes-long cold-start compilation delays, and uses a bounded KV cache (five sink plus 44 recent latent frames) with top-aligned rotary position remapping to support rollouts beyond the training horizon. Table 2 breaks one served block of sixteen frames into condition encoding 10.6 ms, Transformer four denoising steps and cache publication 414.4 ms, and tiny decoder plus host transfer 8.8 ms, totaling 438.8 ms wall time at 36.5 frames/s.
Pretrained foundation-model agents practice in code worlds round by round: they perceive the world only through rendered observations and, after each round, write playbooks inherited by agents in fresh conversations in the next round, with outcomes improving across rounds in hide-and-seek and four further worlds. Unlike self-play reinforcement learning from random weights, these agents already hold abstract knowledge of objects, tools, and geometry, so the main challenge is grounding that knowledge into closed-loop sensorimotor execution; playbooks pass between independent agents, so each lesson must be stated well enough for another reader to check. In hide-and-seek, hiders built shelter with panels by round 4 and seekers crossed walls with a ramp by round 10, against roughly 25 million and 100 million training episodes in the 2019 self-play study; Table 4 shows companion-dog engagement rising 13 to 19, one-lane bridge arrival time falling from 71 to 41 seconds, all four sheep penned in herding, and the quarry loader scoring 90.9 in round 4.
Perspective
The framework targets embodied-agent training and evaluation settings that need inspectable state and programmable interaction rules while also wanting realistic visual observations. It lets new worlds be written as code and rendered through the same renderer, so environments can scale in number and difficulty alongside their agents; for researchers this means constructing navigation, manipulation, tool-use, and multi-agent tasks without authoring bespoke visual assets, and letting pretrained agents practice while seeing only rendered frames. The hide-and-seek results apply to a sequential one-hider, one-seeker setting where agents act by submitting short Python programs and receive no object coordinates or opponent-private observations; the four additional worlds each have their own actions, time limits, and scores, described to the agent only in its task file.
The hide-and-seek and four-world outcomes come from one recorded run with a limited number of rounds, so how results vary across more random seeds and layouts is not yet known. The paper explicitly notes that the 4 and 10 rounds in hide-and-seek versus roughly 25 million and 100 million training episodes in the 2019 study reflect two fundamentally different learning paradigms rather than a direct sample-efficiency ratio. The conditioning interface carries only geometry, so rotationally symmetric objects spinning, material, color, and object identity are not determined by geometry and must be carried by visual history, which may lose them after long occlusions or revisits beyond the memory window. Future work proposes more abstract and compact state representations, such as structured text or high-dimensional latent features, as conditioning interfaces, but aligning them with code-world state and adapting pretrained video models to use them remain open problems. This summary is based on the paper text and appendices and does not include frame-by-frame content of Figures 1 through 11, so descriptions of visual quality and emergent behavior follow the prose and tables.
