Public articles linked to the same research event.
arXiv The authors present WebFovea, a vision-based web agent that placed 2nd with a final score of 57.0 out of 100 in the WebRetriever Challenge 2026 on Protocol III, and report that because the same model was used in all four submissions, the rise of its official hidden-set score from 31.0 to 57.0 reflects changes to the harness, with many observed failures on real websites occurring at four stages—parsing, action effect, result reporting, and information shown to the model—rather than in the model's reasoning.
The authors present WebFovea, a vision-based web agent that placed 2nd with a final score of 57.0 out of 100 in the WebRetriever Challenge 2026 on Protocol III, and report that because the same model was used in all four submissions, the rise of its official hidden-set score from 31.0 to 57.0 reflects changes to the harness, with many observed failures on real websites occurring at four stages—parsing, action effect, result reporting, and information shown to the model—rather than in the model's reasoning.
The authors present WebFovea, a vision-based web agent that placed 2nd with a final score of 57.0 out of 100 in the WebRetriever Challenge 2026 on Protocol III, and report that because the same model was used in all four submissions, the rise of its official hidden-set score from 31.0 to 57.0 reflects changes to the harness, with many observed failures on real websites occurring at four stages—parsing, action effect, result reporting, and information shown to the model—rather than in the model's reasoning.
The authors present WebFovea, a vision-based web agent that placed 2nd with a final score of 57.0 out of 100 in the WebRetriever Challenge 2026 on Protocol III, and report that because the same model was used in all four submissions, the rise of its official hidden-set score from 31.0 to 57.0 reflects changes to the harness, with many observed failures on real websites occurring at four stages—parsing, action effect, result reporting, and information shown to the model—rather than in the model's reasoning.