FloWright treats the workflow as a training harness: co-evolving multiple roles lifts small open models by up to 7.41% across document, slide, chart, code, math, and finance tasks
Synopsis
The work proposes FloWright, which uses the workflow itself as a harness and a hierarchical, structure-aware reward to localize a workflow's single sparse outcome to individual roles, enabling one role to self-evolve and two or more roles to co-evolve, together with DataWright, which hardens existing single-agent datasets into workflow-level tasks; across document, slide, chart, code, math, and finance tasks, small open models trained with FloWright improve performance by up to 7.41%, with co-evolving (+5.03%) gaining more than optimizing one role alone (+2.83%).
Interpretation
FloWright lets more than the generator learn inside a workflow: it uses the execution trace to localize the workflow's single scalar outcome to specific roles, charging each role in proportion to the fraction of its own nodes at fault, so no additional models, labels, or executions are needed. Prior methods that train workflows from execution outcomes typically optimize only the workflow generator, while the other agents that build or execute each workflow remain fixed even though every outcome depends on all of them. The paper gives a role-level credit equation (Eq. 2) and a hierarchical reward (Eq. 3), and states that only genuine root causes are charged, that cascaded failures propagate along the workflow's edges, and that infrastructure faults are excluded; experiments compare four evolution modes on held-out test sets.
FloWright supports four evolution modes: single-role self-evolution, agent-skill co-evolution, upstream-downstream co-evolution, and multi-agent co-evolution; co-evolving more roles gains more. Existing work either trains each agent separately or co-trains agents only within a fixed pairing or a hand-designed system, whereas FloWright's co-evolution happens over the workflows constructed for each task. Among the four Qwen3.5-4B modes, multi-agent co-evolution gains most overall (+5.03%), ahead of agent-skill co-evolution (+2.83%), upstream-downstream co-evolution, and single-role self-evolution; every arm improves under at least one mode.
DataWright converts datasets a single agent can already handle into harder workflow-level tasks for both training and evaluation, through three hardening strategies (Shared, Paired, Decoy) at two hardening levels. The paper notes that workflows are commonly trained and evaluated on single-question math, function-level coding, and short question answering, which a single agent can already handle, so such data can neither reveal nor optimize what a workflow adds beyond a single agent. The paper reports a controlled before-and-after comparison: on the original MMLongBench-Doc (200-sample subset) the workflow and the single agent reach and ; on the original DocFinQA (200-sample subset) they reach and ; on MMLongBench-Doc hardened by DataWright the same workflow reaches against the single agent's , an improvement of points.
FloWright's gains compound across train time and test time and generalize across roles, backbones, RL algorithms, workflow topologies, and pools. One shared harness both provides the reinforcement learning signal at train time and, at test time, distills successful experiences into a reusable prior without changing any weights. Distilling the prior from the system's own experiences gains overall, and from a stronger teacher (GPT-5.4) as well; meta distillation gives the largest test-time gain at five shots, while one shot stays below the untrained baseline; train-time and test-time optimization compound to reach an overall level. Ablations show GRPO, CISPO, and DAPO all instantiate the objective, optimization improves every one of four topologies, and a pool that grows on demand gives the largest single design effect over a fixed pool.
Perspective
The result targets complex tasks that need multi-agent coordination, such as gathering evidence across heterogeneous inputs, reasoning over long contexts, and accumulating intermediate results over many steps; training and evaluation run on DataWright-hardened workflow-level tasks spanning document, slide, chart, code, math, and finance domains, with in-distribution, out-of-distribution, and out-of-domain regimes distinguished. The reinforcement learning experiments are mainly evaluated on two open models of a single family, Qwen3.5-4B and Qwen3.5-9B, while meta distillation is verified on four backbones including two closed-source models it never optimizes; evaluation tasks have inputs fixed once the task begins. For a reader, this means the method applies where a single agent can no longer solve a task in one pass and decomposition, routing, evidence localization, and aggregation are needed, rather than to single-question math or function-level coding that a single agent already handles.
The paper's stated limitations include that the reinforcement learning experiments are mainly evaluated on two open models of a single family, with the authors aiming to incorporate more model families, and that evaluation tasks have inputs fixed once the task begins, with future work planned for dynamic environments where the inputs a workflow reads and the components it needs change while the workflow runs. In addition, the homepage evidence bundle provides parsed full text in which tables and figures (such as Tab. 1 and Fig. 12-16) appear as placeholders, so some specific numbers can only be read from the prose; exact per-arm figures still require the original figures and tables. Strategy applicability is also constrained by prerequisites: Shared requires inputs that each carry several independent questions, and Decoy requires questions that identify their own input, so not every dataset supports every strategy.
