EMFA models error propagation to locate decisive errors in multi-agent failures, raising Who&When step-level attribution by 3.45 and 4.40 points
Synopsis
The work proposes EMFA, which reconstructs failed multi-agent trajectories into structured dependency representations, explicitly models cascading error propagation and persistent interaction loops, and combines dual-branch candidate screening with counterfactual verification to identify the decisive error—the agent–step pair whose correction would recover the failed execution; on the Who&When benchmark it achieves state-of-the-art step-level attribution accuracy, improving the previous best by 3.45 and 4.40 percentage points on the Hand-Crafted and Algorithm-Generated subsets, with the best agent-level accuracy on Hand-Crafted and competitive results on Algorithm-Generated.
Figure 1: Illustration of why error-propagation modeling is needed to separate decisive errors from misleading failure symptoms.
arXivInterpretation
It introduces an error-propagation-oriented perspective on failure attribution: a failed trajectory may contain multiple apparently erroneous steps, including deviations later corrected and downstream symptoms that merely inherit an earlier error, so the decisive error must be found by tracing how errors influence subsequent execution rather than judging steps in isolation. Existing direct-inference, replay-based, fine-tuning-based, and structure-guided approaches largely treat attribution as selecting suspicious steps from an observed failure, without explicitly modeling how errors propagate across the trajectory or persist in unresolved interaction loops before candidate selection. The perspective is grounded in the problem formulation: the decisive error is defined as the agent–step pair whose correction would recover the failed execution, and the earliest decisive error is taken as the attribution target following the benchmark annotation protocol.
EMFA converts flat execution logs into a two-level dependency representation (step-level producer–consumer dependencies and subtask-level dependencies) and extracts two complementary failure structures: an error-propagation graph along information dependencies and a persistent-loop graph capturing repeated pursuit of the same unresolved objective. Whereas CHIEF exploits hierarchical causal graphs to rank candidate errors and FALAT incorporates propagation analysis after candidate pruning, EMFA explicitly models complete error propagation before candidate selection and extracts loops independently of ordinary information-dependency paths. The method comprises five stages—final failure step identification, structured trajectory construction, error propagation modeling, dual-branch candidate screening, and counterfactual decisive error verification—supplemented by deterministic post-processing that adds instruction–response dependencies and deterministic checks for explicit or omitted required actions.
Dual-branch candidate screening is combined with counterfactual verification: the propagation branch prioritizes the earliest unrecovered deviation that materially affects subsequent reasoning or actions, the loop branch identifies the earliest step initiating or sustaining unresolved repetition, each branch contributes at most two candidates, and the verifier judges whether minimally repairing a candidate would prevent downstream failure. Verification uses model-based counterfactual assessment rather than environment replay, consistent with the expert-judgment annotation protocol of Who&When; when multiple candidate repairs appear capable of recovering the execution, predefined responsibility-boundary principles distinguish the step that introduces an error from downstream steps that merely execute, propagate, or expose it. Ablation shows full EMFA outperforms direct LLM screening and both single-branch variants across all metrics; removing the propagation branch causes substantial step-level degradation, while removing the loop branch causes a smaller drop on Hand-Crafted but reduces Algorithm-Generated step-level accuracy from 50.00% to 33.33%.
On the Who&When benchmark, EMFA achieves the best step-level attribution accuracy: 32.76% on Hand-Crafted (a 3.45-point gain over CHIEF) and 50.00% on Algorithm-Generated (a 4.40-point gain); agent-level accuracy is best on Hand-Crafted at 74.14% and competitive on Algorithm-Generated at 66.67%. Across all methods, step-level accuracy is markedly lower than agent-level accuracy, highlighting the greater difficulty of precisely locating the decisive step; EMFA shows consistent gains on this stricter metric across both subsets. The test set contains 58 Magentic-One logs in Hand-Crafted and 126 CaptainAgent logs in Algorithm-Generated, with tasks from GAIA and AssistantBench and ground truth obtained through multi-round expert consensus; EMFA results are averaged over three independent runs under a strict Top-1 criterion and without using task ground truth.
Perspective
The results target turn-based LLM multi-agent systems in which exactly one agent acts at each step, and are evaluated under the Who&When failure-attribution setting, where the attribution target is the earliest decisive agent–step pair per the annotation protocol. Intended uses include debugging failed trajectories, targeted repair, and learning from failed executions; the method can be instantiated with different backbones such as GPT-4o, DeepSeek-V4-Flash, Qwen3.5-397B-A17B, and GLM-4.7, with GPT-4o performing best on both subsets. Evaluation covers the Hand-Crafted subset (58 Magentic-One logs) and the Algorithm-Generated subset (126 CaptainAgent logs), with tasks from GAIA and AssistantBench and ground truth obtained through multi-round expert consensus, and all experiments are conducted without using task ground truth.
The trajectory-length stratification shows that EMFA and CHIEF have lower agent-level accuracy on some longer trajectories; the text offers the possible explanation that structure-guided methods introduce stronger inductive biases and may miss failure patterns not represented by their structures, which remains to be further verified. The method introduces additional inference-time overhead, and more efficient propagation modeling is listed as future work. In addition, several formulas, thresholds, and implementation details (such as temperature and maximum generation length) are not fully rendered in the text, and detailed prompts, configurations, and cost analyses are provided in the supplementary materials, so reproducing those details requires consulting the supplement.
