StoreBench puts agents in charge of an online store: top model DeepSeek-V4-Pro averages 0.700, still below the scripted smart-triage policy at 0.764 and human experts at 0.708
Synopsis
The authors introduce StoreBench, a long-horizon environment in which an agent runs a mid-size online apparel store for 30 to 365 simulated days on a production-grade Medusa v2 commerce backend, using 29 merchant tools under an operation-metered time budget with anchor-calibrated pass thresholds and an exploit-hardened reward; no frontier model matches the scripted smart-triage policy on average (best DeepSeek-V4-Pro 0.700 vs. 0.764), human experts edge out every model at 0.708, most models improve sharply over a full simulated year under Claude Code, and a GRPO post-training run lifts Qwen3.5-27B from 0.136 to 0.373 on held-out tasks.
Figure 2: Three example tasks from the 12-task suite: steady (steady-state operations), trade-war (a correlated supply shock), and black-swan (a silent demand crash). Each chart plots the hidden quantity the task’s event script perturbs; markers show when each event fires and through which channel the agent can learn of it. Quotes are from the task prompts.
arXivInterpretation
Time is metered in operations: every call except the free store_status costs one, and spending the window's last operation or calling end_window advances the world, so simulated time is a pure function of the action sequence and neither model latency nor wall-clock speed can influence it. Most agent benchmarks move the world only when the agent acts, reward a terminal verdict, and set the pass bar arbitrarily; StoreBench lets demand, supplier repricing and failures, and market shocks arrive whether or not the agent is ready, and decouples time from inference speed. The store runs on Medusa v2, a production-grade open-source commerce platform, and the agent acts through the same 29 merchant tools a human operator would use; replaying an action sequence against the live backend reproduces the ledger byte for byte, run manifests record the source commit, image digest, and hashes of seed data and calibration file, and missing or corrupt artifacts grade to numeric zero.
Scoring is calibrated against scripted anchors: for every task-seed pair six scripted policies (do-nothing, absentee, blast-list, rule-based, smart-triage, and an exploiter) run through the real tool surface, with absentee's terminal growth as the zero point and the gap to the best anchor as the scale, and the pass threshold at the midpoint of the two best honest trading anchors under four gates. Prior long-horizon business evaluations lean on terminal financial outcomes without task-specific calibrated scales or pass thresholds; here all 36 reported task-seed cells and the 24 cells of two privately calibrated held-out seeds pass all four gates, with thresholds spanning 0.43 to 0.76. Blast-list captures a median 65% of the top anchor's composite yet passes nowhere; during development one candidate lens (skeleton-crew) failed the judgment-gap gate when mass shipping out-earned triage under its original operation budget, and the task was rebalanced and recalibrated rather than the gate relaxed.
The reward surface is hardened through an audit-close-pin loop: each exploit found by an audit of 756 pre-hardening trials and a code-level exploit hunt was closed in the world model, pinned by an adversarial regression test, and folded into the exploiter anchor, so every calibration re-plays the whole catalogue. Open-ended rewards invite gaming, and the gap between 'no current model exploits this' and 'no optimizer can profit from this' is treated as the benchmark's attack surface, with a catalogue covering continuity credit for quitters, reliability laundering, free cancels at the buzzer, a listing side door, lottery pricing, a refund cliff, and phantom cash. Post-hardening the exploiter averages 0.04 over the reported suite against 0.46 to 0.76 for the honest trading anchors; every exploit has exactly one regression test, and all tests run in CI on every commit.
The environment also serves as a training ground: grading is deterministic and judge-free and the dense per-window reward channel telescopes exactly to the graded terminal growth; GRPO with LoRA on five disjoint 7- to 15-day tasks raises Qwen3.5-27B's mean composite on held-out evaluation tasks from 0.136 to 0.373, improving 31 of 33 cells and lifting the share of profitable episodes from 3% to 72%. This offers reproducible evidence that RL on a few shorter-horizon tasks transfers to unseen longer scenarios, with gains coming primarily from operating discipline (more on-time shipments and handled returns) rather than commercial strategy. Training used 36 steps of 5 task-seed prompts times 6 rollouts, 1,080 episodes in total; the base model exhausted the 150-turn cap on 97 of 99 episodes and reached window 33 on average, while the checkpoint reached window 47; checkpoint 34, evaluated under the same protocol, scores 0.341 and also improves 31 of 33 cells.
Perspective
The environment targets researchers and engineering teams who need to evaluate or train long-horizon operational agents: it suits measuring sustained reliability, economic judgment, and execution under limited supervision when demand is hidden and suppliers and markets shift, and it suits RL post-training through the dense reward channel. The authors describe the framework as general, with modular tools, event models, and backend adaptable to other real-world operational settings; the current instance is a direct-to-consumer apparel store where demand moves only through the shelf, with no advertising channel, no demand momentum, no market price discovery, scripted competitors, and an ultimately parametric hidden demand model. Evaluation-task prompts, configurations, and calibration are withheld with the evaluation suite, so the environment cannot be run or trained against outside the authors' own evaluation; the supplementary material provides five training-split tasks, ten sample trajectories, their calibration, and the composite scorer plus a standalone re-grader that reproduces a score from an exported ledger.
The short-horizon suite (60 to 90 windows, goose harness) and the full-year task (365 windows, Claude Code) differ in both harness and effort, so the authors report them separately, do not average across them, and state they cannot isolate the harness's effect; how much of the large full-year gains comes from the harness therefore remains open. The human expert comparison carries two bounds: humans contribute one selected outcome per cell where models average three unselected attempts, and participants were allowed one retry; scoring each cell by its earliest completed attempt gives a mean of 0.656 instead of 0.708. The post-training study uses one base model, one hand-adjusted training run, and held-out tasks rather than held-out seeds, which the authors frame as evidence that the environment supports hill-climbing with transfer to unseen scenarios, not a tuned training recipe. In addition, policy entropy on training batches rose from 0.2 to between 1.0 and 1.9, the checkpoint's messages drifted toward an emphatic, occasionally garbled style (19% of turns vs. under 1% for the base model), and in 8 of 99 episodes the checkpoint stops calling tools after about twelve windows, which the authors suggest should be met with a KL penalty or lower-temperature sampling on longer runs.
