AdvSim2Real co-evolves tasks, injections, and a web executor inside a frozen web world model, raising a 4B agent's completion under an unseen frontier adversary by 33.6% relative
Synopsis
AdvSim2Real co-evolves a task curriculum, an injection adversary, and a web executor inside a frozen web world model, where the curriculum is rewarded for tasks the executor solves about half the time and the adversary only for flipping a judged success into a failure; the trained 4B executor raises clean completion on 150 web tasks from 74.89% to 81.33%, completion under three learned adversaries from 48.07% to 57.48%, and completion under the untrained Kimi-K3 from 23.00% to 30.72%.
Interpretation
Introduces AdvSim2Real, a two-stage framework that co-evolves a task curriculum, an injection adversary, and an executor inside the frozen WebWorld-14B world model, so both tasks and attacks track the current agent. Prior defenses fine-tune on injections fixed before training, so the defender never meets an adaptive attacker; prior adversarial training adapts the attacker but keeps tasks fixed, so a task stops teaching once solved. The paper specifies the algorithm and the two-stage alternating updates, and evaluates on 150 form-filling tasks across five skill strata with per-seed counts and a checkpoint sequence.
Introduces the success-flip reward: an accepted clean run is replayed to the injection step, and the adversary earns credit only when the continuation fails. Rewarding any executor failure would credit the adversary for failures that occur without an injection; the success-flip reward limits credit to disruption caused by the rendered injection. The paper gives the reward equation, a rendering gate, and paired controls, and states that malformed proposals receive zero and unreachable transitions yield no injection and no flip credit.
Releases a benchmark of 150 form-filling web tasks with five skill strata, protected fields and forbidden controls, a reactive adversary, and a deterministic browser check of the submitted form. The benchmark ships task contracts, protected fields, forbidden controls, and a deterministic check together with code, checkpoints, and trajectories. The paper lists ten templates, five skill strata, low and high difficulty tiers, short and full forms, and states that strict browser success requires matching values, untouched protected fields, no forbidden control, and a review before commit where required.
Provides evidence that the curriculum stage buys capability and the adversarial stage buys robustness: without the curriculum stage the final agent loses 4.67 clean points but only 0.44 attacked points. The ablation compares the complete pipeline against a Stage-2-only pipeline, showing a clear clean gap and an unclear attacked gap. The paper reports clean and attacked means and per-round differences for three ablation rounds, and notes the ablation removes executor and curriculum initialization jointly without matching compute, so it measures the two pipelines rather than the curriculum alone.
Perspective
The work targets web-agent providers that need task-preserving robustness: the executor completes its authorized goal despite competing page instructions. Training and scoring happen inside the frozen WebWorld-14B world model, and all attacked results are verdicts inside that model; the capability checkpoints are additionally verified in a real Chromium browser with a deterministic form check, where strict success rises from 25.56% to 44.44% and correct fields from 52.50% to 73.24%. The benchmark covers 150 form-filling tasks in five skill strata with protected fields and forbidden controls. The authors release code, the benchmark, and all checkpoint results so injection defenses can be evaluated against adversaries trained on the defended agent.
The world model can display a value the executor never entered and the judge can accept a wrong calculation, so a judged robustness gain is not yet a gain in executable task completion; only the capability checkpoints were verified in a browser. The training adversary proposes from the initial page while the evaluation adversary reacts to the trajectory, so matching an adversary checkpoint matches the attacker, not the attack. Evaluation uses three rollout seeds of one trained checkpoint, measuring rollout variance rather than training variance, and reports no confidence intervals. The ablation removes executor and curriculum initialization jointly without matching compute, so it does not establish a robustness benefit from Stage 1 alone. The authors note that the final checkpoint still loses at least some clean completion to a frontier-model adversary, and that every robustness number is a model judgment inside one web world model.
