Skip to main content
Back to timeline
arXivSource publication:

An author–reviewer loop matched a multi-agent workflow on MCFM Fortran-to-C++ translation at about one-third the cost per file

Synopsis

The work proposes collaborative human–AI agentic workflows in which domain experts write the specification and plan and agents write code under a deterministic orchestration pattern, with each stage ending in a numerical comparison against reference code and human approval; across fourteen experiments on translating MCFM from Fortran to C++, a simple author–reviewer loop with enforced limits settled a comparable number of files to a multi-agent workflow at about one-third the cost per file.

Source-provided article image: Designing Collaborative AI-Driven Workflows for Scientific Software Engineering
Figure 1 ·

Figure 1: Authoring responsibilities and the flow from human-authored inputs to accepted code. See Section 2 for details.

arXiv · Page 3

Interpretation

It proposes and tests a collaborative workflow architecture centered on a verification boundary, with five design principles: numerical verification at every stage boundary, humans designing and accepting objectives while agents do repetitive work, reproducibility and provenance, collaboration between roles, and an open, user-controlled agent execution policy. Relative to using agents for full automation, the work explicitly separates supervision, implementation, and acceptance, and keeps the specification, plan, tools, and logs in a team-shared versioned Lab Notebook directory tree so the same task can be rerun under a different orchestrator for comparison. The architecture and principles are implemented on the MCFM translation case across fourteen runs, all starting from the same commit and the same 445 candidate files, with the gate and oracle being the same code in every run.

With the same orchestrator (Claude Code) and model (Opus 5), the author–reviewer loop (R12) and the multi-agent workflow (R2) settled the same files, with the loop costing about 40% as much at about the same time per file. This comparison varies only the orchestration pattern while holding the agent and model fixed, attributing the cost difference to the orchestration layer; the authors note that parallel agents each read the same background material while the number finished depends only on what survives the single integration step, so the time saved by parallelism does not compensate. R12 and R2 are single runs while most other configurations were replicated; costs are list prices and file counts come from each run's git branch rather than the agent's own checklist.

When behavioral rules were enforced architecturally by the orchestrator (CodeScribe, csloop), no phase exceeded the iteration bound or deviated from the execution policy; when rules were stated only through prompts (Claude Code, ccloop), author phases frequently ran past the bound, one reaching 67 iterations against a limit of 30, without corrective ReAct feedback. This comparison holds the author–reviewer pattern fixed and varies only how the execution policy is enforced, indicating that reproducible, auditable, and predictably terminating runs require enforced structure rather than prompted constraints. csloop covers eight runs (R4–R11) and ccloop three runs (R12–R14), all over the same agentic layer.

Model choice shaped tool-calling behavior and cost: the nominally cheapest tier, Sonnet 5, produced the two most expensive runs per file settled (R7, R8), using several times as many tool calls per file as every other csloop run and finishing fewer files, while GPT 5.6-Sol (R9–R11) performed similarly per file to the Opus 5 runs at similar token cost but with more tool calls. This comparison keeps the orchestrator fixed and changes only the model, attributing differences to the model and its API format; it also finds that all fourteen runs picked files in nearly the same order and settled on largely the same files, largely because the orchestrator's ranking tool determined which files were ready. Eight csloop runs span three model families, with R8 an incomplete run lacking an archived agent_log; cache-efficiency differences are attributed to the different API toggles Anthropic and OpenAI expose.

Perspective

The work targets scientific software teams that want to use agents without eroding domain understanding and correctness guarantees, especially teams maintaining large legacy Fortran codebases and migrating incrementally to C++ or device-portable kernels. The authors state that the Lab Notebook's inheritance structure, shared agentic layer design, and staged transformation format are generic, so another project can keep the same layout, change the project-specific parts, and add its own agentic layer with tools accessible via MCP. They also note the author–reviewer loop can be applied to deploy simulation campaigns, run analyses, and perform work in manageable portions, letting humans supervise AI work effectively. The cost conclusions are scoped by the authors to this task, where files do not depend on each other but each file is small next to the material an agent must read before touching it.

Agentic runs are stochastic, and the authors note that a single observation per condition cannot separate an orchestrator or model effect from run-to-run variance, with ccworkflow Opus 5 (R2) and ccloop Sonnet 5 (R14) being single runs. Claude Code keeps a user-specific memory directory outside the Lab Notebook; memory from R1 influenced R12 and led it to skip the Mods/ directory, and the authors state this influence left no trace in their records, while R13 and R14, run after purging that memory, processed the skipped files. A gap in the shared specification left the C++ representation of the Fortran kind parameters (sp, dp, ex, qp) unresolved, and the models produced two incompatible answers, a real extended-precision type and a placeholder integer, both of which passed the full benchmark suite because the converted code never referenced these constants directly. The authors also identify the model API format and agent execution policy as open problems, citing a report that multi-agent systems operating at scale circumvented well-enforced guardrails.

Sources