Skip to main content
Back to timeline
arXivSource publication:

CaptchaArena trains a single CaptchaAgent policy on 20,000 execution-verified CAPTCHA puzzles, lifting average Pass@1 from 11.4 to 71.7 against a human 94.1

Synopsis

The work builds CaptchaArena, a large-scale fine-grained computer-use training dataset of 20,000 interactive CAPTCHA puzzles across 20 types and five interaction modes, where every solution is replayed in a real browser and accepted by the page's own verifier, together with 20,000 screenshot-action trajectories (18,000 carrying judge-filtered step-by-step reasoning annotations) and pixel-mask supervision for irregular targets; training a single 9B policy, CaptchaAgent, on it reaches 70.5 average Pass@1 after supervised fine-tuning and 71.7 after reinforcement learning with the environment verifier as reward, versus 11.4 for the untrained backbone, 35.2 for the strongest open-weight GUI agent, 69.2 for the strongest closed-source model, and 94.1 for humans.

AI-generated editorial illustration: CaptchaArena: A Large-Scale, Fine-Grained Dataset for Training Computer-Use Agents on Interactive CAPTCHAs

Interpretation

The dataset represents each CAPTCHA puzzle as an executable browser task, admitting an answer only after its replay is accepted by the site's own verifier, upgrading static image-answer pairs into closed-loop interaction supervision. Earlier CAPTCHA resources traded off among type coverage, interaction fidelity, and trajectory supervision: ReCAP and CaptchaMind cover few types, MirrorCAPTCHA's synthetic training samples provide only answer and click-coordinate supervision, MCA-Bench predicts structured solutions rather than executing them step-by-step in a live browser, and CAPTCHA-X grades against manually defined geometric acceptance regions. The dataset contains 20,000 puzzles across 20 types and five interaction modes (single-click, multi-click, arrow-cycle, real-time, text-entry), split into 16,000 train, 2,000 validation, and 2,000 test puzzles; each puzzle has an executable reference solution run in a real browser and accepted by the page verifier, so each yields at least one screenshot-action trajectory.

Grading irregular targets against pixel masks rather than bounding boxes changes the measured grounding ability. The paper regrades the same recorded clicks twice, once against the stored mask and once against the tight bounding rectangle of that mask, and finds rectangle grading accepts clicks that miss the target but fall inside the rectangle. On Pick Area the mask-to-rectangle accepted-area ratio averages 0.62 (range 0.37-0.89), with the rectangle covering 65.1% of the image versus 38.8% for the mask; UI-TARS-1.5-7B rises from 33.0 under the mask to 77.5 under the rectangle on that type, while a click sampled uniformly inside the rectangle obtains 61.6.

A single 9B policy shares one set of parameters, tool schema, and observation format across all 20 types, with no per-task adapter or task-conditioned switching, and outperforms larger general models. MCA-Bench previously fine-tuned task-specific LoRA adapters, while general GUI agents and closed-source models perform unevenly on CAPTCHAs. CaptchaAgent reaches 71.7 average Pass@1 and 86.8 Pass@5 with a 97.9% submit rate; six open-weight GUI agents top out at 35.2 and three closed-source models span 49.4 to 69.2; scaling the Qwen3.5 family from 9B to 35B-A3B leaves the average at 11.4, indicating scale alone does not close the gap.

Using the environment verifier directly as the reinforcement learning reward, with no learned reward model or human reward labels, improves first-rollout reliability and transfers to two external benchmarks. Supervised fine-tuning leaves two failure modes: the policy does not always submit, and its actions are not always precise enough for the verifier to accept; the GRPO stage targets both and mines difficulty per puzzle rather than per type to preserve within-group reward variance. RL raises average Pass@1 from 70.5 to 71.7 (2.1 points excluding Click Order), improving 13 of 20 types; the gain at Pass@1 (1.2) exceeds that at Pass@5 (0.8), consistent with improving first-rollout reliability rather than expanding the solvable set; on external benchmarks Open CaptchaWorld rises from 47.2 to 51.0 and Halligan from 13.6 to 20.0.

Perspective

The dataset targets training and evaluating computer-use agents in a self-hosted environment, for CAPTCHA tasks that require closed-loop perception-action, pixel-precise localization, and multi-step interaction; because the environment deterministically verifies the final page state, reinforcement learning needs neither a learned reward model nor human reward labels, a setup reusable for other interactive tasks with executable verifiers. What a reader can take directly is the 20,000 puzzles, each with a reference solution and screenshot-action trajectory, the 18,000 reasoning annotations, and the pixel masks for irregular targets; the paper also notes that reasoning annotations depend on commercial models, the vision encoder stays frozen, and actions are discrete, so trajectories cannot reproduce the mouse-trajectory signals used by behavioral detectors.

The paper reports humans at 94.1 average on the same interface versus 71.7 for the policy, with the remaining gap concentrated in a few types: Place Dot, Dice Count, Patch Select, and Slide Puzzle together account for 19.5 of the 23.6-point gap. Failures are grouped into three modes: non-submission, weak grounding, and unstable execution; for Dice Count the supervised policy submits on only 54.4% of episodes, so nearly half its episodes are scored as failures with the answer never checked, meaning that type's Pass@1 reflects both task difficulty and non-submission. The Pass@1-to-Pass@5 gap exceeds 30 points on Slide Puzzle, Bingo, and Pick Area, indicating execution variance rather than missing capability. In addition, source images are reused between training and test: six types share source images (100% of Image Matching test images, 100% of Coordinates, 100% of Dart Count, 100% of Misleading Click, 100% of Patch Select, 100% of Rotation Match), which the paper audits with content hashes and notes does not imply puzzle duplication, but readers should still weigh it when interpreting generalization. External benchmark scores are not directly comparable with published numbers because the interaction and evaluation protocols differ.

Sources