Skip to main content
Back to timeline
arXivSource publication:

WebFovea placed 2nd at 57.0/100 in the WebRetriever Challenge 2026, raising its hidden-set score from 31.0 to 57.0 with the same model by reworking the harness between model and page

Related research and updates

Synopsis

The authors present WebFovea, a vision-based web agent that placed 2nd with a final score of 57.0 out of 100 in the WebRetriever Challenge 2026 on Protocol III, and report that because the same model was used in all four submissions, the rise of its official hidden-set score from 31.0 to 57.0 reflects changes to the harness, with many observed failures on real websites occurring at four stages—parsing, action effect, result reporting, and information shown to the model—rather than in the model's reasoning.

Source-provided article image: WebFovea: When the Model Is Right but the Click Is Wrong -- Reliable Round Trips for Vision-Based Web Agents on Live Websites
Figure 2 ·

Figure 2: A parsing failure: a self-generated chat-template token typed into a data portal’s filter box (CSO PxStat; task: CPI for June 2022). The site filters on the literal string and returns no results. The model eventually diagnosed the problem in its own reasoning but kept re-emitting the token and ran out of its 40 steps.

arXiv

Interpretation

The paper decomposes each step of a vision-based web agent into four things that must all go right: the model's reply is parsed into the intended action, the action takes effect on the page, the result is reported back accurately, and the model is shown the information it needs. Prior work often attributes web-agent success or failure to multimodal LLM reasoning; this work moves attention to the harness, the code between the model and the page. Evidence comes from failures observed on real websites: a coordinate-space mismatch placed every click at 3/4 of its intended coordinates, actions on native dropdowns, inside iframes, and in text boxes failed silently, and self-generated chat-template tokens contaminated 4.9% of task episodes.

WebFovea hardens each of the four stages and surrounds the loop with guardrails that keep the agent within the rules and its budget. The paper reports not only the design but also the evidence for each component (including negative results), a failure analysis, the limitations, and a roadmap that includes routing different steps to different models. Evidence comes from the official hidden-set evaluation of the WebRetriever Challenge 2026, with a final score of 57.0 out of 100 and a 2nd-place finish.

The same model was used in all four submissions, and the official hidden-set score rose from 31.0 to 57.0, indicating the gain came from harness changes rather than a model change. This provides same-model contrastive evidence for the 'model is right but the click is wrong' phenomenon, tying performance gains to harness changes rather than model capability. The authors state that the rise reflects changes to the harness, and note run-to-run variance on live sites.

The four-stage view does not depend on the model, although some individual fixes do. This view shifts part of the web-agent reliability problem from model capability to engineering interfaces, making it easier to transfer across models. The authors support this with score changes under harness changes with the same model, and explicitly distinguish the model-independent overall view from model-dependent individual fixes.

Perspective

The result targets vision-based web agents that complete retrieval tasks end to end on live websites, in settings where the agent must operate the site's own interface and return a verifiable answer, such as the evaluation defined by Protocol III of the WebRetriever benchmark. The authors describe the four-stage view as not depending on the model, so it can serve as an organizing approach when building harnesses for other vision-based web agents; the roadmap's routing of different steps to different models points to a multi-model division of labor as a next direction.

The summary level does not provide full numerical ablations for each component or the specific content of the negative results, so the individual contribution of each fix still needs the full text to confirm. The authors note run-to-run variance on live sites, so how stably the hidden-set score rise reproduces is an open question. Although the four-stage view is described as not depending on the model, some individual fixes are model-dependent, and their cross-model transfer remains to be verified.

Sources