Skip to main content
Back to timeline
arXivSource publication:

JevSpawn lets agents derive finite action fields from natural-language tasks, topping five of eight benchmarks and cutting navigation latency to about 41 seconds

Synopsis

JevSpawn introduces a training-free compositional policy that turns natural-language task specifications into executable finite action fields, exploring them through parallel action spawning, feedback-driven branch selection, representation revision, and recovery from retained branches; across eight benchmark tasks against seven agent baselines and a TypeSafe Jev variant it achieves the highest scores on five of eight tasks and, on Maze and Grid, raises success rates to 0.96 and 0.95 while reducing end-to-end latency from the fastest baselines' 47.91 and 61.44 seconds to 40.91 and 40.58 seconds.

AI-generated editorial illustration: JevSpawn: Adaptive Agentic Inference through Compositional Action Spaces

Interpretation

The paper makes the action space itself part of the policy: the model infers action declarations from task context and interaction rules, combines field values into executable actions, and uses joint probabilities to guide parallel spawning. Earlier Jev-style finite-field prediction required the application to specify fields in advance, whereas JevSpawn infers and revises fields through interaction, extending finite prediction to tasks that require autonomous solving, including spatial navigation. The paper supports this with a formal compositional policy definition (conditional field domains, value-tree normalization, absorbing empty extension) and a worked construction in the appendix, and on Maze and Grid the domains are complete so execution requires no text generation.

Parallel spawning and feedback-driven state transitions are unified in one architecture: a beam retains assignments, each retained assignment defines a child branch and action, execution feedback guides branch selection and representation revision, and earlier branches can be resumed. Relative to methods that only search or only optimize workflows, this places branch retention, operation selection, and expansion in a single algorithm and allows action preferences to be revised after observations return without regenerating the preceding trajectory. Ablations show parallel spawning contributes strongly: reducing width from four to one lowers Maze success from 0.88 to 0.16, Grid success from 0.94 to 0.58, and LightsOut reward from 0.61 to 0.06; at a fixed budget of four actions, two parents with two actions each reach 1.00 on both Maze and Grid.

Shared action structure and model prefix reuse reduce repeated generation and context computation without additional training. The paper treats known declared values as known token sequences whose causal hidden states can be evaluated in parallel, projecting only onto vocabulary entries that distinguish continuations, reducing processed input tokens from per-branch repetition to a shared prefix plus suffixes. The paper derives conditions for equivalent shared attention and reports text decoding throughput of 544 and 580 tokens per second on PPNL and LightsOut, the highest values in its table, while Maze and Grid use finite evaluation throughout.

Across eight benchmark tasks, JevSpawn shows task-dependent trade-offs against seven agent baselines and a TypeSafe Jev variant: highest scores on five tasks, joint quality and latency gains on navigation, but longer execution on some tasks in exchange for higher scores. The results ground the question of whether finite prediction can support sustained task solving in a concrete task spectrum, and show that within the same architecture TypeSafe Jev scoring beats Qwen scoring on PPNL, LightsOut, and Sokoban while scoring lower on Grid, RushHour, 2048, and Nullify. Evaluation uses Qwen3.8-27B on four H100 GPUs, batch size 8, a 16,384-token context, up to 36 exploration rounds and a 300-second limit; PPNL has 1,136 instances, Maze 25 initial states, and the remaining tasks 100 instances each from seeds 42 through 141. On LightsOut and 2048 higher scores accompany longer execution, and Sokoban frontier membership is less stable.

Perspective

This work targets agent tasks where the action space must be derived from natural-language instructions and revised through interaction, and it applies where task rules supply or permit inference of finite fields, such as path planning, grid navigation, and sequential puzzles. For systems seeking to cut repeated generation and context computation without retraining a model, it offers a reusable compositional policy and a shared-prefix computation path. The paper also provides code and an interactive demo, supporting replication and extension on comparable benchmarks.

Open questions remain: how reliably finite fields can be inferred on tasks with open strings or no explicit domains; beam pruning removes probability mass and may discard a partial assignment with a higher-scoring completion, so beam selection only approximates exhaustive ranking; in long-horizon experiments 2048 and Sokoban show timeout rates of 82%, 85%, and 77% at 72 and 108 rounds, and without the deadline the 2048 score rises from 315.16 to 1102.24, indicating sensitivity to round limits and deadlines; Sokoban frontier membership is unstable, with only a 1.50-second latency margin over LLMCompiler and a lower score; and the TypeSafe Jev variant takes 1.4 to 2.1 times as long as Qwen scoring, including API communication. In addition, this reading covered the full text, but tables and figures appear as text, so some numeric cross-checks still require the original.

Sources