InFlowOp uses one label-free cost to set both task decomposition and fault repair, gaining up to 11.97% over a single agent on Braid
Synopsis
The work proposes InFlowOp, which prices every decision in one label-free cost that weighs how well an agent's competence meets a subtask's demands against how much that agent takes to run, using it before execution to bidirectionally set decomposition granularity and agent assignment and during execution to locate and locally correct faults, and introduces Braid, a benchmark whose tasks require multi-agent coordination, reporting gains up to +11.97% over single-agent baselines across eight domains and six backbones, with +9.64% from in-flow optimization.
Interpretation
InFlowOp unifies workflow construction and execution under a single label-free cost: in construction, Coalesce performs bidirectional, cost-driven atomic coalescing to decide decomposition granularity and agent assignment; during execution, the same cost matrix is reused to match each subtask's declared input and output conditions against what was actually realized, blame the fault on assignment or decomposition, and climb a cost-ordered ladder from re-assign to re-decompose. Prior practice fixes decomposition granularity in advance through a template, prompt, or chosen number of steps, or adapts to the task alone; fault localization usually needs a reference answer, a graded outcome, or a trained assessor, and the unit of fix is the whole workflow through re-execution, re-search, or retraining. This work places agent creation, assignment, and decomposition inside one search under one price, and corrects faults in flow. The paper formalizes the cost matrix, the credit matrix, the cost-ordered ladder, and scoped resumption, and states that a correction is committed only if the subtask then executes and its unmet conditions clear, otherwise it is rolled back, so that tempering can improve a workflow but never leaves it worse than the one construction produces.
The paper introduces Braid, a benchmark whose tasks are built to require multi-agent coordination beyond single-agent capability: each task composes several questions and supplies extra sources that answer none of them, without stating which source serves which question; tasks are composed in distributed form (each question has its own gold source) and anchored form (all questions share one gold source), under four quality-control criteria covering separation, routing, leakage, and answerability. The paper argues that workflows are largely measured on suites a single competent agent already solves, where repeated sampling captures much of what coordination adds, and that existing coordination-stressing benchmarks impose the division through the environment (partitioned state, split tools, withheld knowledge) rather than through what the task demands. Braid is built on ten public datasets spanning documents, slides, finance, charts, mathematics, physics, science, and code, forming nineteen evaluation arms over fourteen Braid datasets, each evaluation-only and inheriting the test side of its source dataset; the paper reports per-arm question counts.
Across eight domains and six backbones (Qwen3.5-4B/9B and GPT-5-mini, GPT-5.4-mini, GPT-5.4, GPT-5.6-luna), InFlowOp improves over the single-LLM baseline by up to 11.97% and over the single-agent baseline by up to 9.64% under a matched turn budget; on Qwen3.5-9B the numbers are 17.38 for the single agent, 15.49 for greedy search, 22.21 for Coalesce, 23.80 for Coalesce in-flow, and 25.13 for full InFlowOp. Using the same Braid arms, the paper shows the gain does not come from dividing the task itself: a greedily constructed workflow falls below the single-agent baseline, while deciding the division by cost reverses that; the gain also exceeds what a larger backbone delivers, as Qwen3.5-4B under InFlowOp reaches 20.81, above the single-agent accuracy of Qwen3.5-9B, GPT-5-mini, and GPT-5.4-mini. Results are reported as per-domain and overall accuracy tables, with the single-agent baseline given a compute budget matched to the workflow it is compared against on every task; the paper also reports that gains are largest on charts, mathematics, and finance and smallest on physics, where all methods stay below any backbone.
Ablations show that pricing the cost matrix with a rubric estimator beats having an LLM judge matching scores in one pass; a cost matrix that updates during Coalesce beats a static one; a dynamically growing agent pool beats a static pool; and an agent profile helps only when it remembers success, since applying all outcomes to profiles falls below the agent card alone on all three backbones tested. These ablations isolate how cost is computed, whether the pool can fill a gap, and what a profile remembers, indicating that the gain follows from whether the estimated cost can effectively guide atom matching and coalescing, and whether the pool can create the specialist an atom demands. The paper reports specific deltas for these ablations on Qwen3.5-4B and Qwen3.5-9B among others, and notes that enabling a global pass beside the local one lowers accuracy, which is why in-flow optimization stays local in every other experiment.
Perspective
The result targets task settings with checkable reference answers that can be decomposed into subtasks with declared contracts, and applies to settings where a multi-agent workflow must be built and repaired online under a limited compute budget; it benefits researchers and practitioners working on agent systems and workflow orchestration. Braid's nineteen evaluation arms, eight domains, and six backbones provide a directly comparable evaluation surface, and its distributed and anchored composition methods offer a reusable construction procedure for future benchmarks.
The paper states that its evaluation is confined to Braid tasks graded against a reference answer per question, and that tasks settled by no checkable answer, such as open-ended generation or long-horizon interaction with a stateful environment, lie outside this study; the authors plan to extend to such settings and to workflows whose agents act on an environment that changes as they run. Cost estimation also depends on agent cards and self-evolving profiles, and how failures are treated in a profile changes results, as the ablation shows; the global pass lowering accuracy when enabled beside the local one leaves room for further work on where to look for faults. This evidence bundle is the full text, including body, appendices, and data statistics tables, but some formulas and figures appear as placeholders, so specific numeric details should be checked against the original.
