OSWorld-Science tests 12 VLMs on 146 scientific-software tasks: the strongest agent averages 73.7% and only 40% on the hardest set
Synopsis
The work introduces OSWorld-Science, a benchmark and evaluation environment combining expert-defined scientific tasks, artifact-based execution graders, and a shared agent harness, and evaluates 12 vision-language models across 7 scientific domains, 17 software packages, and 3 languages, finding that the strongest model, Claude Fable 5.1, averages 73.7% and about 40% on the hardest task set, with failures concentrated in the interaction layer and result verification rather than in scientific reasoning itself.
Interpretation
The benchmark contains 146 tasks spanning chemistry (43), physics (31), medicine (26), statistics (20), biology (18), geographic information (6), and linguistics (2), across 17 software packages and 14 recorded configurations, generated through expert proposals and an AI-assisted co-worker design route and selected by scientific value and difficulty. Compared with general computer-use benchmarks such as WebArena, OSWorld, and Windows Agent Arena, and with scientific benchmarks such as ScienceBoard and Terminal-Bench-Science, this work anchors task collection in practicing scientists' judgments of importance and difficulty and covers hybrid GUI and CLI interaction. Task distribution, software shares (QuPath 23 tasks, ChemDraw and ASKCOS 21 each, SAS/R 20), and operation-primitive counts (click in 124 tasks, text entry in 101, drawing in 21, CLI supported in 122) are itemized in the main text and appendices.
Evaluation uses artifact-based execution graders that inspect application state and generated files such as molecular structures, segmentation masks, plots, and numerical results, awarding partial credit for incomplete outcomes; the accompanying harness integrates model adapters, interaction-loop control, and trajectory logging so models and interaction strategies can be compared. Unlike evaluations that match answers or read only final text, this design brings heterogeneous scientific artifacts into one scoring framework and treats the harness itself as an object of controlled analysis. Scores are normalized to a 0-1 pass rate; the appendix describes exact-match grading for single-value problems and error-band grading for simulation-parameter problems, plus an artifact-level labeling procedure over 1,500 released runs.
Across 12 vision-language models under the same harness, Claude Fable 5.1 leads at 73.7% mean score ($6.5 per task), ahead of Claude Opus 5 at 57.9% and GPT-6 Astra at 57.4%, while GPT-5.6 luna reaches 22.6% at $0.23 per task; rankings shift sharply by domain, with Astra at about 6% in chemistry but 100% on the two linguistics tasks. The results separate model capability from harness design and pair performance with cost, whereas most existing benchmarks report a single score. Scores are task-weighted means including partial credit with VOID runs assigned zero; the main text and Appendix B also report output tokens, total cost, and per-domain rankings.
Trajectory analysis shows only 402 of 1,500 released runs (26.3%) receive full credit, and of the 1,100 non-full-credit runs, 667 (60.6%) end without the graded artifact, 613 of them by exhausting the budget; narration and thinking are model signatures rather than predictors of success, and no interaction measure is significant after Bonferroni correction once model and task fixed effects are included. This provides quantitative evidence that failures concentrate at the delivery stage and indicates that raising reasoning effort or context length does not resolve them. Ablations run on 23 QuPath tasks: for Opus 5, moving from medium to max effort grows output tokens roughly tenfold and raises the score from 44.1% to 59.7%; increasing the history window from 1 to 10 screenshots cuts mean steps from 52 to 22 and raises the score from 35.2% to 57.0%.
Perspective
The benchmark targets evaluation settings where verifiable research workflows must be completed inside professional scientific software; it suits comparing vision-language models together with interaction harnesses and diagnosing whether failures arise in the interaction layer, the method layer, or delivery. Tasks, images, and code are public, with a contamination watermark added to detect training-data mixing; because some software requires licenses, this version does not plan to release all trajectories. For readers choosing a scientific-software agent or designing new tasks, the framework offers reusable task templates, graders, and trajectory-logging conventions.
Most model-task pairs are single runs, and the appendices state explicitly that these counts carry no variance estimate, so domain rankings and ablation differences should be read as directional observations. Some conclusions rest on manual attribution of trajectories, and coordinate frames are inferred from where clicks land, fitting most clicks but not all. Several tasks depend on third-party online services (such as a SynergyFinder server outage) or on licensed software, so reproducibility is subject to external conditions. In addition, some runs were made with Mnova not pre-installed or under a different evaluator version, making them not directly comparable with the released task; readers who see only the abstract may underestimate how much these conditions affect specific numbers.
