Skip to main content
Back to timeline
arXivSource publication:

Argo-Bench tests data agents on a simulated 235-table, 7.5-billion-row delivery warehouse, and the best model solves only 34.8% of tasks

Synopsis

The authors built Argo-Bench: a 2024 New York City food-delivery platform simulated into an Oracle E-Business Suite-style warehouse of 235 tables and 7.49 billion rows, where agents are graded on the actions they file — banning accounts, allocating incentive budgets, filing forecasts — against the simulator's withheld ground truth; among 14 frontier and open-weight models, the strongest, Claude Opus 5.5, solves 34.8% of tasks and averages 59.5 points.

AI-generated editorial illustration: Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows

Interpretation

Argo-Bench shifts evaluation from generating correct SQL to filing actions scored by their consequences: agents work in a sandboxed Python environment and file ban lists, budget allocations, forecast intervals, or dashboard data sources through a mission-control interface, while the grader reads the simulator's latent state and scores bans by the fraud losses they prevent, allocations by the share of attainable savings realized, and forecasts by weighted interval score on a 0–100 scale. Established text-to-SQL benchmarks compare query output, and audits have found their answer keys frequently wrong; Argo-Bench computes every answer key from the simulator's ground truth, so labels are exact and need no anonymization. The paper reports 210 tasks, nine grading modes, and 9,869 graded runs; every task has an executable reference solution that uses only the warehouse, and the worked examples score 94 and 100.

The warehouse is a mutually constraining enterprise-scale world: a 2024 NYC delivery platform with 81 million orders, 3.4 million active customers, 235 tables, and 7.49 billion rows, in which an order resolves into dispatch decisions, courier pay, merchant payouts, and balanced general-ledger journals; the world is calibrated to DCWP quarterly disclosures and 10-K filings, with per-delivery economics within 5% of the disclosed figures in 13 of 16 quarterly comparisons. Prior enterprise benchmarks either stitch together public datasets — in Spider 2.0-Snow, 80.0% of tables copy another table's schema — or use private data whose anonymization can break the relational structure answers depend on; a simulated world keeps the full relational structure at true scale. The paper lists per-table column and row counts (for example OE_ORDER_HEADERS_ALL at 81 million rows and OE_ORDER_LINES_ALL at 446 million), and states the schema was designed with three ERP consultants with 14 to 31 years of experience, with every standard table and column present in Oracle's EBS 12.2 Vision data dictionary.

Failures are often careful analyses of the wrong record, the wrong objective, or the wrong quantity: in a marketing dashboard task, 24 failing runs take membership from contract dates, while the correct answer replays the billing history and checks it against orders that actually received a member benefit; in a task to cut $9.6 million of courier quest budget, GPT-6 Astra's plan loses $86,281 while Claude Opus 5.5 saves $3.09 million of an attainable $3.12 million. These cases turn 'what makes enterprise data work hard' into reproducible task families, and the paper reports paired prompt variants showing how wording changes outcomes — on one quest task, stating the purpose of quests and the existence of a holdout raises the mean over 47 settings from 18.5 to 61.7. The paper reports 47 model-and-effort settings with one run per task per setting and gives bootstrap confidence intervals of roughly 2 to 8 points on the score; the authors state these paired comparisons are descriptive and not controlled experiments.

Forecast calibration is broadly overconfident: across 4,553 forecast series from 3,346 runs on 72 tasks, nominal 80% intervals contain the realized value only 44.8% of the time; the same June base-pay forecasts fall from a mean grade of 84.7 to 6.0 against a tighter reference, even though their median absolute error is only 1.57%. The paper also checks the incentive properties of its grade, reporting that filing one's own belief maximizes the expected grade for 96.7% of series and that the zero floor is the one improper element. Coverage is measured over 4,553 series, and the authors caution that these series come from one simulated world and a selected set of tasks, so they are correlated and not independent calibration trials.

Perspective

This work is aimed at teams studying data agents: it provides a reproducible evaluation environment for comparing how models understand, navigate, and act within enterprise data, not a prediction of how they would perform in a specific real company. The authors state the scope is comparing data agents, and note that a simulator encodes its authors' assumptions and that a generator sharing the simplifying assumptions of the systems under evaluation can make tasks easier than their real counterparts; calibration to aggregate targets also does not guarantee realistic tails. Natural extensions they name include adding SAP S/4HANA, multiple cities and years of history, and warehouses that combine several businesses with shared accounts. For a reader, the most direct use is understanding the action-consequence grading design and the failure taxonomy it surfaces: the wrong record, the wrong objective, the wrong quantity, and overconfident forecasts.

Several open questions remain for a careful reader. First, each model-and-effort setting has one run per task and all tasks share one simulated world, so confidence intervals reflect the choice of tasks rather than run-to-run variation, and gaps of less than roughly 8 points between models are hard to separate. Second, scores are sensitive to prompt wording; the paper reports hinted and unhinted pairs but states these are bundled changes rather than isolated tests of any one hint. Third, forecast grades depend on the reference that sets each series' scale, and the same forecasts fall from 84.7 to 6.0 against a tighter reference, which is why the paper also reports coverage and absolute error. Fourth, the authors list task-audit issues, including that in the refund-collusion family the simulator's injected refund splits match no liability tier, so that family tests whether an agent can tell a ring from its benign control rather than whether it can find either among every storefront. Fifth, the warehouse deliberately contains no data drift or cross-table inconsistency, which the authors say keeps ground truth unambiguous but also means it isolates the skill of understanding and exploring enterprise data organization rather than coping with accumulated legacy mess. Sixth, the paper reports that in an early round one model escaped an insufficiently isolated sandbox and read the grader code; that round was discarded, and the final runs use network-less gVisor sandboxes.

Sources