ReCast turns frozen-LLM hidden states into attribution-oriented step representations, achieving the best Hit@1 on four failure-attribution benchmarks
Synopsis
After quantifying that text embeddings and hidden states separate root-cause steps only weakly, the work proposes ReCast: an attribution-aware probe selects layers, a sparse Johnson–Lindenstrauss projection builds complementary pattern and deviation features, and a bidirectional encoder trained with contrastive and ranking objectives converts frozen-LLM hidden states into attribution-oriented step representations; it also releases the ReCast-2K training set and attains the best Hit@1 on Who&When, TraceElephant, AFTraj-2K, and Who&When Pro.
Figure 1: Comparison of text embeddings and last-layer hidden-state representations on failed trajectories from ReCast-2K and the Handcrafted and Algorithm subsets of Who&When. Lower ARS/GRS and higher SI indicate better separation. See Appendix A for metric definitions.
arXivInterpretation
The paper first quantifies how well Qwen3 text embeddings and Qwen3.5-27B last-layer hidden states distinguish root-cause steps, using Adjacent Root Similarity, Global Root Similarity, and a cosine Silhouette Index, and finds limited separation: root causes remain similar to their predecessor and other steps within a trace, and cross-trace separation is weak. Prior work largely reused text embeddings or hidden states as step representations without systematically testing their attribution discriminability; this work frames “which step representation suits failure attribution” as a measurable question and supplies within-trace and cross-trace separation metrics. Three basic representations are compared on failed training trajectories and on the Handcrafted and Algorithm subsets of Who&When, with metric and representation details in Appendix A; the finding is stated as limited separation rather than a single significance test.
ReCast has three parts: an attribution-aware probe selects layers by their ability to separate the root cause while covering model depth; a fixed sparse Johnson–Lindenstrauss projection yields a pattern branch and a deviation branch referenced to the mean of successful trajectories; and a bidirectional encoder, trained jointly with root-cause contrastive, non-root-cause contrastive, and trajectory-level ranking losses, outputs contextualized step representations. Unlike methods that score internal signals directly (MASPrism’s NLL and attention) or model successful dynamics (OAT), ReCast learns attribution-oriented step representations rather than scoring raw signals, and unlike external-trajectory methods it exploits frozen-LLM internal signals. Ablations show the full model attains the highest mean Hit@1; attribution-aware layer selection exceeds five alternatives by 2.75–4.40 percentage points, the joint loss improves over ranking-only training by 3.85 points, and removing augmentation lowers Hit@1 by 4.22 points; ablations of the projection branches, encoder, and non-root anchor selection also fall below the full model.
ReCast achieves the best Hit@1 on all four benchmarks: it surpasses the second-best method CHIEF by 5.65 and 9.19 percentage points on Who&When Algorithm and Handcrafted; it exceeds ASCon by 4.71 and 1.11 points on TraceElephant Captain-Agent and Magentic-One; and it exceeds ASCon by 1.02 and 8.71 points on AFTraj-2K and Who&When Pro. These gains appear under different training-data settings (ReCast-2K for Who&When and TraceElephant, the official AFTraj-2K training split, and the Who&When Pro training partition), indicating the benefit is not confined to the authors’ own dataset. All methods use the same three seeds with mean and standard deviation reported; AgenTracer-8B is not publicly available, so only its reported values are cited without multi-seed results.
The learned step representations attain lower ARS and GRS and higher SI under the same metrics, indicating improved within-trace and cross-trace separation; with precomputed features, ReCast’s attribution stage averages 0.11–8.28 ms per trajectory, whereas the measured LLM baselines require 2.88–30.50 seconds per trajectory. This both answers the limited-separation observation that motivated the work and shows the attribution stage can run efficiently once step features are precomputed, in contrast to methods relying on autoregressive decoding. The representation-separation comparison covers the training set, Handcrafted, and Algorithm; the efficiency comparison reports mean per-trajectory time on Who&When and TraceElephant, excluding feature extraction and embedding-cache preparation from attribution time.
Perspective
The results target multi-agent execution traces whose candidate steps are model-generated actions: environment observations and tool-result messages are excluded, so traces whose annotated decisive error is a tool output fall outside the evaluation (two such AG trajectories in Who&When were removed). The method suits settings where frozen-LLM hidden states are available, and its attribution stage runs in milliseconds once features are precomputed, fitting debugging and online-auditing workflows that need fast root-cause localization; training requires failed trajectories with root-cause labels, supplied here by ReCast-2K and by each benchmark’s own training split. Backbone benefits vary by benchmark: 27B gives the highest Hit@1 on Who&When, AFTraj-2K, and Who&When Pro, while 0.8B and 9B lead on Captain-Agent and Magentic-One respectively, so backbone choice can follow the target benchmark.
How root-cause labels are constructed affects transferability: ReCast-2K labels come from step-level voting by three LLM judges (766 unanimous, 346 two-of-three) plus manual review of 384 trajectories, whereas Who&When Pro injects controlled errors to obtain causally grounded labels, so the two label semantics are not identical. The relationship between the separation metrics (ARS, GRS, SI) and Hit@1 is presented consistently but without a quantitative mapping from metric improvement to accuracy gain. Backbone size and accuracy are not monotonic, with smaller backbones leading on some subsets, so benefits depend on benchmark and metric. In addition, several numeric values in the main text (for example, specific figures in some tables and in Figures 3 and 4) are absent from the parsed text, so exact experimental numbers should be checked against the original tables and appendices.
