WebLoop trains an execution-free Critic with execution feedback, reaching 41.5 Overall on WebRise and 38.9% on WebGen-Bench with a 9B model
Synopsis
WebLoop jointly optimizes generation, critique, and refinement as roles of one shared policy for functional Web generation: Generator and Refiner outputs receive direct browser-execution rewards, while the Critic never sees execution results and is trained from requirement-level execution labels (discriminability) plus the execution gain induced by subsequent refinements (helpfulness), using a staged reward that first establishes diagnosis and then adds consequence-aware credit. With Qwen3.5-9B, WebLoop reaches 41.5 Overall on WebRise and 38.9% accuracy on WebGen-Bench, improving the base model by 11.3 and 15.4 points.
Figure 1: Conceptual comparison of execution-based learning paradigms for functional Web generation. (a) Generation-only RL optimizes executable outputs directly. (b) Execution-guided RL feeds external feedback back into the code policy. (c) WebLoop jointly learns Generator, Critic, and Refiner, assigning execution-derived credit to the intermediate Critic.
arXivInterpretation
The paper formulates functional Web self-refinement as a structured reinforcement learning problem and identifies intermediate Critic credit assignment as the central challenge: the Generator and Refiner produce executable artifacts with direct environment rewards, whereas the Critic influences downstream behavior without a directly executable outcome. Prior execution-based learning largely optimized final page quality or used execution feedback for iterative editing, without directly addressing how to optimize the intermediate Critic that guides subsequent refinement. The problem setting is supported by the paper's loop formulation and reward definitions, making it a framework-level argument rather than a single experimental result.
WebLoop trains an execution-free Critic with two complementary signals: requirement-level execution labels supervise discriminability, and the execution improvement induced by subsequent refinements measures helpfulness, combined through a staged reward that first establishes diagnosis and then introduces downstream utility. Relative to optimizing either signal alone or mixing both from the start, the staged combination performs better in controlled comparisons, and the two signals rank differently across benchmarks. Table 4 shows discriminability-only, helpfulness-only, and mixed-from-start variants all fall short of WebLoop; the paper also reports that discriminability is stronger on WebRise while helpfulness is stronger on WebGen-Bench.
All three roles are optimized jointly under a shared policy using role-specific comparison groups for group-relative advantage estimation and a clipped GRPO objective. Relative to treating refinement or Critic learning as isolated stages, joint optimization shows a gap already visible in first-pass generation. Table 3 reports Staged Training at 40.1 and 36.1% after refinement versus WebLoop at 41.5 and 38.9%; Table 1 shows WebLoop (Gen) raising the 9B base from 30.2 to 39.1 on WebRise and from 23.5% to 34.2% on WebGen-Bench.
The gains transfer to first-pass generation, persist at the 27B scale, and generalize from text-only training to Markdown, Sketch, Image, and Video inputs. This indicates that loop-level training strengthens the underlying generation policy itself, not only critique-conditioned revision. Table 2 shows average WebRise score rising from 31.9 for the base model to 37.5 for WebLoop (Gen) and 40.1 after the complete loop, with consistent gains across all three functional metrics; at 27B, WebLoop (Refine) reaches 52.2 and 42.8%.
Perspective
The result applies to the functional Web generation setting: tasks are defined by natural-language descriptions plus explicit and implicit requirements, and judged by browser-executed interactions and assertions specified in an Interaction Contract Graph. The training corpus contains 2,880 executable tasks across 21 domains and 1,442 scenarios; controlled training uses text-conditioned tasks, and the main WebRise results are evaluated under the Text modality. The approach fits generation settings that can supply executable verification, and the Critic stays execution-free at inference, so deployment does not require browser results or execution traces as input. For researchers and engineering teams wanting to use execution supervision for intermediate decisions rather than only final artifacts, the recipe of role-specific comparison groups plus joint GRPO is directly reusable.
A careful reader would still watch: the relative strength of discriminability and helpfulness differs across benchmarks, and whether that pattern holds on broader task distributions is open; how the coefficient controlling the second-stage discriminability weight affects outcomes is not fully mapped; the qualitative analysis shows failures can come from Critic misjudgment, insufficient repair guidance, or the Refiner not implementing guidance, but the share of each source in overall failures is not quantified; and the critique-length versus refinement-quality analysis is explicitly labeled descriptive rather than causal. In addition, formulas and some figures appear as placeholders in the main text, so reproducing exact objectives and prompt templates requires the full definitions in Appendices A, B, and D.
