Skip to main content
Back to timeline
arXivSource publication:

Turning scientific rules into callable checks: R2T lifts SciCode repair from 26/30 to 29/30, with gains concentrated in a few tasks

Synopsis

Rules to Tools (R2T) packages public scientific requirements such as boundary conditions, equations, and required output files into callable executable checks; in matched SciCode repair comparisons where the text and tool groups share written rules, starting programs, model, and budgets, complete repair is 26/30 with text versus 29/30 with checks across two task-ID cohorts and 13/16 versus 15/16 in the eight-ID cohort, while the twelve-task shared-definition cohort ties at 13/24, indicating task-dependent gains alongside lower reported output and higher public CPU use in some cohorts.

AI-generated editorial illustration: Rules to Tools: Executable Checks for LLM Agents in Scientific Computing

Interpretation

R2T turns public scientific requirements into executable measurements of the current program: a diagnostic combines a public relation, probe input, tolerance, and reporting rules, and while the tool group can call a prepared implementation, the text group implements the same procedure in ordinary Python, with an independent evaluator grading the finished program. Prior work established feedback-driven revision (Self-Refine, Reflexion, CRITIC) and executable specifications (CodeMetaAgent, SecTDD, CodeSpec); R2T differs by turning public numerical relations into checks on an evolving program and, in matched repair, holding written criteria, starting programs, model, and budgets fixed while varying only access to a prepared implementation. The paper gives a formal definition of diagnostics, a specification-equivalence account, and Table 1 mapping four public requirements to measurements (SciCode 12 Numerov recurrence, SciCode 22 field reconstruction under rotation, SciCode 73 reciprocal-cell geometry, PDE flow boundary and equation residuals); checks were authored or reviewed by the study team, and the eight-ID cohort used static preflight and execution on each starting module.

In matched SciCode repair, prepared checks yield task-dependent complete-repair gains: across two task-ID cohorts, 26/30 with text versus 29/30 with checks; the eight-ID cohort is 13/16 versus 15/16 and the seven-ID extension is 13/14 versus 14/14. The gain is not uniform: in the eight-ID cohort two tasks favor tools, one favors text, and five tie, and removing task 77 leaves a zero difference across the other seven tasks; the twelve-task shared-definition cohort ties at 13/24. The eight-ID task-macro difference is 12.5 percentage points with a task-cluster bootstrap 95% interval of [-12.5, 43.75] points; the seven-ID extension differs by 7.14 points with a zero lower endpoint; exact two-sided sign tests on non-tied task effects give the reported p-values for the eight-ID (2:1), seven-ID (1:0), and combined 3:1 patterns. Each task has two continuations per group, and all 32 assigned continuations completed and received final grades.

Gains do not require an initial check to flag a violation: task 17 has an initial violation and favors tools, while tasks 77 and 11 report no initial violation yet still favor tools, and task 37 has no initial violation but favors text. This complicates the intuition that checks first surface a problem and then guide repair: task 77's starting program passes all five supplied public checks, and its two successful tool trajectories later use ordinary Python to revise periodic wrapping and pressure calculations; development-exposed task 22 begins with all four supplied rotation checks reporting no violation, yet tools repair 2/2 while text repairs 0/2. The paper reports that all sixteen held-out tool trajectories called the checker, making 34 calls; initial reports flagged tasks 17 and 48; task 77's initial checks reported no violation, as did task 11's two initial checks.

Access form and cost are measured separately: source delivered through ordinary Python and the dedicated command both reach 15/16 on the eight IDs but with different task-level outcomes, and in the matched PDE comparison detailed text scores 23/24 versus 24/24 with checks, with 31.2% lower reported output for checks. The paper reports complete repair, native-step progress, and resource cost as distinct endpoints, noting that in the eight-ID cohort both groups average 95.83% native-step completion, so the complete-repair difference occurs alongside unchanged average subtask completion. In the eight-ID cohort, text-to-tool totals are 7.25M versus 7.88M model tokens, 767,871 versus 819,981 output tokens, and 456.90 versus 633.15 public CPU seconds; across both cohorts reported output is at least 1,327,295 text versus 1,262,059 tool tokens, at least 4.9% lower for tools, while public CPU rises from 779.43 to 1,459.88 seconds; the development-exposed cohort saves 22.9% reported output and PDE uses 1,062,099 versus 730,723.

Perspective

The work targets revising an existing program in scientific computing: the agent already has the complete problem statement, written criteria, and ordinary code execution, and what is added is an executable measurement of the current program. It speaks to researchers and engineering teams building scientific coding-agent evaluations and toolchains, and applies to benchmarks such as SciCode-style numerical modules, PDE solvers, molecular simulation workflows, and scientific software repositories. The paper limits checks to measurable public requirements such as boundary values, physical relations, equations, and required output files, and stresses that final scores come from an independent evaluator while diagnostic observations receive no score credit. Natural next steps include repeating this matched design across more task families and more starting programs, folding checker authoring cost into the accounting, and running more repetitions across delivery forms (dedicated command, source through ordinary Python, initial report) to separate task-level differences.

A careful reader will still watch several things: the eight-ID 95% interval crosses zero with a zero lower endpoint, and the seven-ID extension also has a zero lower endpoint, so the robustness of task-level gains depends on more repetitions; the twelve-task shared-definition cohort ties at 13/24, showing gains do not appear in every cohort. The gains on tasks 77 and 11 occur where initial checks report no violation, leaving open how checker coverage relates to an agent's own investigation. On cost, some requests lack returned usage (a seven-ID text request, a PDE checker request, and two flow text requests), and the paper reports lower bounds using issued caps, so the exact size of output savings remains to be confirmed; public CPU rises in several cohorts, and the flow cohort's output-cost ordering is unresolved under the 25-response sensitivity. In addition, checker authoring time was not recorded, and the historical MDArena and AInstein panels include trigger conflicts and unparsed logs; these are scope and accounting boundaries rather than conclusions in themselves.

Sources