BIABench tests AI agents on 16 published studies: routine analyses complete, but 3D and time-lapse tasks fall to 0.00–0.10
Synopsis
The authors built BIABench, which reconstructs 16 published biological studies as end-to-end bioimage-analysis tasks giving the agent raw images plus a biologist's instruction, scores outcomes against the studies' own reported results with deterministic metrics and scores process with a VLM against expert-written rubrics, running six agents across several models and two instruction levels with three repeats each, and found routine 2D tasks reach up to 0.97 while 5D nuclear-pore quantification and 3D oncogenic puncta quantification stay at or below 0.25 in every configuration.
Interpretation
BIABench reduces published studies to verifiable tasks: each keeps only the raw images, the biologist's request for the study's key readout, and the reported outputs as ground truth, spanning eleven analysis subtasks across 2D, 3D and time-lapse data and modalities from H&E histology to single-molecule localization microscopy, with samples from six source organisms. Earlier bioimage benchmarks evaluate a single step (nucleus segmentation, cell tracking, in-silico labeling, quality control) or have agents train models on existing datasets and score held-out predictions, whereas this scores the whole cycle from raw images to the study's conclusion, with method choice including GUI software itself part of what is scored. Tasks are rebuilt from published studies, and a ground-truth replay reaches full score on every task with a ground-truth file, so no metric is capped below the maximum; each task carries a machine-readable YAML specification declaring modality, dimensionality, temporal mode, subtasks, output format, scoring metric and source accession.
Scoring is two-track: an outcome score uses field-standard metrics against ground truth (Dice for segmentation, the Cell Tracking Challenge SEG and TRA measures, F1 and localization error for spot detection, significance direction for colocalization, the Kolmogorov–Smirnov statistic or relative error of a derived quantity for kinetics), while a process score has a vision–language model rate every stage from loading raw data to the final report against severity-weighted expert rubrics. Unlike process-level benchmarks whose score is entirely judge-derived, rankings here use the deterministic outcome score alone and the judged process score is reported separately with its human agreement measured; unlike benchmarks that prescribe the method, the choice of analysis is left to the agent. Judge validation used twenty runs covering all six agents and all tasks, reviewed item by item by an expert cell biologist on blinded copies; on items both decided, the adopted judge (Claude Sonnet 5) agreed with the expert on 0.87 (Cohen's κ 0.67, CI 0.61–0.73), passing 0.27 of items the expert marked failed and failing 0.07 of those the expert marked passed.
Holding the model fixed at GPT-5.6 Sol across six agents, variation across tasks far exceeded variation across agents: the best agents reached 0.85–0.97 on NF-κB translocation quantification, HeLa nucleus/cytoplasm segmentation and cell counting, while 5D nuclear-pore assembly kinetics fell to 0.00–0.25 and 3D puncta quantification to 0.00–0.10; biological specialization conferred no overall advantage, with general-purpose agents achieving or tying the highest mean on most of the 16 tasks and leading by up to 0.97 versus 0.32 on bacterial tracking. This supplies a previously missing empirical fact: on end-to-end microscopy analysis with self-chosen methods, domain-specific design does not automatically translate into an advantage when run autonomously, and failures concentrate in tasks that add a third dimension or a time axis. Each agent–task pair was run three times independently, with outcome scores summarized as the mean of three runs and standard deviations across tasks; every agent ran through one shared interface, output files were recovered from the filesystem, and vendor coding command-line tools were treated as model-confounded harnesses.
Neither stronger models, cost, nor more detailed instructions reliably close the gap: on DeepSeek Harness a stronger model raised the mean from 0.52 to 0.65, while replacing the brief instruction with an expert-written detailed protocol left the mean almost unchanged (0.52 versus 0.52 on V4-Flash and 0.58 versus 0.58 on GPT-5.6 Sol), merely redistributing task-level performance; the cheapest configuration, DeepSeek-V4-Flash, reached about 80% of the strongest configuration's score at roughly $0.09 per run, while Codex cost five times as much for a gain of about 0.01. This rules out both 'stronger model' and 'more detailed instruction' as intuitive remedies and puts the cost–accuracy relation on the table: more expensive configurations were not necessarily more accurate. Model swaps were run on DeepSeek Harness (six models) and Claude Code (three models), and the instruction-detail comparison re-tested DeepSeek Harness on GPT-5.6 Sol and V4-Flash; cost comes from provider billing for the Claude Code and DeepSeek-V4-Pro configurations and from metered token usage priced at published rates elsewhere, with each agent's exposed usage signals stated and missing counts never read as zero.
Perspective
The benchmark targets autonomously running bioimage-analysis agents on tasks whose input is microscopy imagery and whose output is the physical quantity a study itself reports, covering 2D, 3D and time-lapse data and modalities from H&E histology to single-molecule localization microscopy; the task-construction recipe can be extended to new studies and the interface admits napari- or model-zoo-based agents. For a reader, it supplies a map of where current agents remain unreliable as data complexity grows, plus a reusable evaluation and task-generation method, rather than a specific biological conclusion. Because every task is scored at the study's own endpoint, the design assumes in theory that human experts would achieve nearly perfect scores on these tasks.
What a careful reader would still watch: process and outcome scores correlate only weakly (Pearson about 0.33 across all scored runs, about 0.13 excluding near-zero outcomes), expert-assigned process scores are no more predictive of outcome, and the judge compresses its ratings into a narrow range, assigning almost no run a score below a certain level; across three attempts on the same task, scores differed by more than 0.2 in a substantial fraction of agent–task pairs, and taking the best of three brought five of the six agents close to one another, indicating that differences between agents primarily reflect reliability rather than best-case capability. Judge agreement with the expert was lowest for tool choice and use (0.77, κ 0.47), where the judge tends to credit on the strength of the narration; the evidence a run preserves varies widely by harness, with expert skip rates from 0.10 for Agentic-J to 0.44 for CopilotJ, which affects how auditable the process score is. In addition, some configurations were excluded for technical reasons (GLM-5.1 and V4-Flash under Claude Code, where intermittent empty API responses were logged as finished turns, and wound-healing runs under V4-Flash with the detailed instruction that exceeded the provider's image-size limit and were rerun), and provider safety filters blocked some models on the SARS-CoV-2 task; the text states these exclusions item by item, so readers can judge the applicable scope.
