SUSTech and CityU propose Box2-Bench: frontier models use good workflows but lose up to 43.3 points under bad ones
Synopsis
The authors introduce Box2-Bench, which holds the model and task fixed while varying workflow availability and reliability across five matched conditions (none, good, bad, partial, mixed), finding that frontier models such as Gemini 3.7 Flash, DeepSeek V4 Flash, and GLM 5.2 benefit from good workflows in 8 of 12 model-task pairs but are hurt by bad workflows in all 12 (drops of 3.3 to 43.3 points), and that mixed workflows underperform their matched partial workflows in 10 of 12 pairs; training two open-weight models on bad workflows only, counterfactual SFT cuts the bad-workflow penalty from 20.0 to 6.7 points on AIME and from 9.8 to 0.6 points on WebShop while weakening use of held-out good workflows, and outcome-based RL restores AIME utilization from -3.3 to +6.
Interpretation
The paper defines thinking outside the box as benefiting from useful guidance while overriding unreliable guidance, and measures it separately from task capability with Box2-Bench, which fixes the task, environment, evaluator, and model and varies only the workflow condition. Existing benchmarks evaluate autonomous task completion or compliance with prescribed procedures, but do not isolate whether models adjust their reliance as guidance reliability changes; Box2 uses five matched conditions (no workflow, Good, Bad, Partial, Mixed) to decompose utilization, robustness, recovery when guidance stops, and recovery when guidance becomes misleading into four paired effects. Workflows are generated by a fixed external model, checked by an independent verifier for usefulness, misleadingness, and answer leakage, then frozen and shared across target models; the frontier evaluation spans mathematics, software engineering, information search, and tool use.
The results show that using guidance and rejecting guidance are distinct capabilities: good workflows improve performance in 8 of 12 frontier model-task pairs (gains of 5.9, 13.7, and 16.8 points on AutomationBench), while bad workflows reduce performance in all 12 pairs, with drops from 3.3 to 43.3 points. This separation shows that task performance alone does not capture a model's ability to regulate reliance on guidance reliability, and open-weight models show the same pattern: good workflows improve all 6 pairs while bad workflows hurt 5 of 6. Paired comparisons hold task, environment, evaluator, and model fixed while varying only the workflow; AIME 2026 uses eight samples per problem with majority voting, and WebShop uses one complete episode per each of the 500 official test tasks.
Reliance shows inertia: mixed workflows, which share a useful prefix followed by a misleading suffix, underperform their matched partial workflows in 10 of 12 frontier model-task pairs, indicating that models often carry forward reliance established by initially useful workflows instead of revising it when reliability changes. Because Partial and Mixed share the same useful prefix and change point, their difference isolates the effect of continuing with misleading guidance rather than receiving no further guidance, a dimension raw task accuracy cannot reveal. The comparison runs under frozen workflows with identical task identifiers and reappears in open-weight models, motivating the training experiments.
Training on bad workflows alone can improve robustness, but two stages are needed to preserve utilization: counterfactual SFT shrinks the bad-workflow penalty from 20.0 to 6.7 points on AIME and from 9.8 to 0.6 points on WebShop while weakening reliance on held-out good workflows, and subsequent outcome-based RL restores AIME utilization from -3.3 to +6.7 while retaining a robustness improvement over the base model. The paper notes that the SFT loss cannot uniquely identify selective reliance, since selective rejection and broadly ignoring workflows both fit the training data, so preserving held-out utilization must arise through generalization; outcome-based RL does not privilege one verified trajectory and lets the model discover alternatives. Training uses only bad workflows with good workflows reserved for evaluation; AIME uses Qwen3-4B and WebShop uses Qwen3.5-9B, training and evaluation tasks are separated, and evaluation covers all 30 AIME problems and the 500 official WebShop test tasks.
Perspective
The results speak to model developers and evaluators whose agents receive external procedural guidance: Box2-Bench applies to settings where the task, environment, and evaluator are fixed and workflow reliability is controllable, and it can diagnose behavior when guidance turns bad or fails mid-execution, as well as serve as a training recipe that uses only bad workflows for robustness while reserving good workflows for evaluation. The authors note that workflows are generated by a fixed external model as a scalable proxy for human designers, and that future work should evaluate guidance written by people with more diverse intentions and error patterns; Box2 varies reliability through controlled and frozen workflows, leaving interactive, ambiguous, partially correct, and model-adaptive guidance to future study. The transfer evidence comes from multi-agent collaboration and memory-augmented reasoning, which the authors describe as preliminary.
Readers should still watch several things: workflows are generated by a fixed external model and audited by model-based semantic review, and the authors state that this audit reports neither an inter-annotator reliability coefficient nor a candidate-level rejection rate, nor does it estimate how often such errors occur in human-written workflows; the benchmark therefore evaluates a controlled collection of plausible procedural errors rather than their prevalence. The training conclusions rest on Qwen3-4B on AIME 2026 and Qwen3.5-9B on WebShop, and the SFT-versus-RL trade-off is not uniform across tasks, with the balance established by SFT on WebShop remaining largely stable after RL. In the transfer results, multi-agent episode success rises from 18.00% to 22.67% after SFT and reaches 18.67% with SFT+RL, and the authors note the gains are not monotonic across training stages; the memory-augmented scores are described by the authors as a transfer evaluation built on LongMemEval-V2-Small that is not directly comparable to the official leaderboard protocol. In addition, although this evidence bundle is full text, tables and figures appear as text, so some graphical detail of individual values cannot be cross-checked from the text.
