EVOKE ranks the same actions under alternative goals at fixed states, pushing pretrained world knowledge into decisions: ALFWorld unseen success rises from 60.4% to 91.8%
Synopsis
EVOKE is a post-training method that holds the environment state and interaction history fixed while introducing alternative goals for the same candidate actions and training on contrastive rankings, forcing the policy to elicit action-consequence knowledge already present from pretraining; across ALFWorld, WebShop, and search-based QA on three backbones it achieves the best average, raising ALFWorld unseen-game success from 60.4% for πboot to 91.8% while using fewer actions.
Interpretation
The paper reframes transferability for LLM agents as eliciting rather than acquiring world knowledge: in digital environments most action consequences are already internalized during pretraining, so no extra prediction objective or separate world model is needed. Against world-model lines such as WALL-E, ITP, MemWM, IWM, PaW, and EnvRL, which predict future observations or add auxiliary prediction signals, EVOKE trains no predictor and uses no inference-time planning; it changes only how the supervision signal is organized. Grounded in the paper's review of related work and in the observation from From Word to World that pretrained models already predict next states well in structured text-based environments; the paper's own probe experiments support that the knowledge is present.
The core mechanism is goal intervention at fixed states: the same state, history, and candidate action set are re-evaluated under different goals, and whenever preferences flip, a policy relying on contextual habits or single-goal correlations cannot rank them correctly. Existing policy-optimization methods (GRPO, self-distillation variants, ETO, DAgger) differ mainly in how the training signal is computed; EVOKE focuses on the goals under which each visited state is supervised. Goal-conditioned RL work such as universal value functions, successor features, HER, and Hindsight Supervised Learning shares experience across goals, whereas EVOKE applies goal diversity to a single decision. The theoretical motivation comes from Richens et al. (2025), who show an agent competent across a sufficiently rich set of goals must contain a world model recoverable from its goal-conditioned behavior; in practice an LLM annotator (Qwen2.5-72B-Instruct) labels executed outcomes, so labels rest on real transitions rather than predictions.
Training uses contrastive ranking rather than imitation: a listwise term raises the mass of the positive set as a whole, a pairwise term separates every positive from every competitor, and competitors are filled first with the policy's own high-probability non-positives as hard negatives. Compared with Positive SFT, which receives exactly the same goal contexts, positives, and updates but imitates instead of ranking, ranking leads by 7.5 points on unseen games (91.8% versus 84.3%), showing the gain is not merely more goal data. Ablations run on Qwen2.5-3B and ALFWorld: Source only (original goals only) reaches 82.8% unseen success, Source replay (repeating original-goal data to match updates) 85.9%, and EVOKE 91.8%; across nested data subsets from 20% to 100%, EVOKE beats SFT in all 15 paired runs and exceeds full-data SFT with 60% of the states.
Diagnostic experiments indicate the pretrained backbone already encodes action consequences, and EVOKE makes the policy decide by the goal rather than by habit. Linear probes decode consequences such as whether the target is held, lies in the named receptacle, or is hot, clean, or cool almost perfectly in Qwen2.5-3B before any ALFWorld training; on transitions where the same action text leads to different outcomes, an action-text probe is near chance while model representations stay above 97.0%, against 65.0–68.0% for the same architecture with random weights. On 812 goal pairs that share state, history, and available actions but have disjoint positive sets, πboot solves only 19.5%, and EVOKE makes habitual errors on 6.4% of decisions, 4.3 points fewer than Positive SFT (95% CI [3.0, 5.6]); on unseen games EVOKE solves 91.0% within 20 actions versus at most 80.6% for other trained variants.
Perspective
The results target LLM agents acting in digital environments such as websites, retrieval, and text-based household tasks; the method assumes action consequences largely fall within pretrained world knowledge, so the paper explicitly sets aside physical-environment dynamics such as contact and motion that are hard to capture without grounding. For teams aiming to cut training cost and inference-time planning overhead, EVOKE offers a path that replaces prediction objectives with decision-level supervision; at deployment the policy uses only the goal, history, and available actions, with no world-model module, no inference-time planning, and no annotator. Iterative aggregation collects new states and new mistakes round by round in the manner of DAgger, keeping hard negatives aligned with the current policy.
Note that all ablations and diagnostics concentrate on Qwen2.5-3B and ALFWorld, while other backbones and environments mainly supply main-table results; goal interventions rely on an LLM annotator to propose alternative goals and judge executed outcomes, and the effect of annotation quality on preference labels is not separately quantified in the main text. The probes show the backbone encodes action consequences, but how that encoding is brought into decisions leaves room for further characterization. Joint accuracy on goal pairs and unseen-game success measure different abilities: the former isolates a single decision, while the latter also depends on execution such as issuing valid actions and not revisiting locations, so the two are best read separately. The paper states that source code, trained checkpoints, and experiment scripts will be released, so independent reproduction and extension await those materials.
