Skip to main content
Back to timeline
arXivSource publication:

TraceDance auto-builds 107 behavior benchmarks from 252,557 real deployment traces, and nine frontier models pass only 26.7% on average

Synopsis

TraceDance is an agent system that constructs targeted benchmarks from real deployment traces for user-specified undesirable behaviors, using Anchor-and-Confirm (programmable retrieval plus candidate-level confirmation by a Flash LLM) and an Anchor Synthesis Loop to generate behavior specifications, and using decision-point continuation so an evaluated LLM produces its next turn at a recorded decision point and is graded by a behavior-specific rubric; experiments on 252,557 sessions from Claude Code and OpenClaw produce 107 benchmarks with 4,125 instances, fulfilling 95.3% of build-target requests, with both human annotators confirming the requested behavior in 84% of sampled instances, while nine frontier LLMs achieve a mean pass rate of only 26.7%.

AI-generated editorial illustration: TraceDance: An Automated System for Building Agent Behavior Benchmarks from Real-World Agent Deployment Traces

Interpretation

TraceDance turns deployment traces into on-demand behavior benchmarks: a user describes an undesirable behavior in natural language, and the system returns a benchmark containing context inputs preceding confirmed occurrences of that behavior plus a behavior-specific rubric, or a rejection reason when it cannot comply. Prior task-completion benchmarks such as Terminal-Bench check the final container state, and safety and process-level suites also use fixed test cases and target behaviors, while auditing tools identify undesirable behaviors in traces without converting them into reusable tests; the paper states that TraceDance is the first system to turn this idea into on-demand benchmarks for user-specified undesirable behaviors. The input/output contract, the three decision frames (action, failure, claim) with their cut rules, and the three instance-validity conditions are given in Section 3; experiments cover 139 test queries, of which 107 are build targets and 32 are rejection targets.

Anchor-and-Confirm scans broadly with programmable anchors on CPUs and uses a Flash LLM only to confirm candidates, keeping construction cost tractable, while the Anchor Synthesis Loop synthesizes, validates, and revises custom specifications when no predefined one matches. Reviewing every session with an LLM or human is too costly at deployment scale, so the paper separates broad scanning from semantic confirmation and requires generated anchor code to pass programmatic safety checks and Reviewer Model code review before execution, then uses retrieval-quality statistics such as confirmation rate and candidate counts to drive revision. Successful queries scan an average of 128,297 sessions, pass 706 candidates to the Fast Model, and ultimately produce an average of 40.4 instances, requiring a mean of 552 LLM calls across all construction stages versus 1.3 for rejected adversarial queries; Appendix A.2 gives round budgets and retrieval-quality thresholds, plus an anchor revision example that raises matched sessions from 3 to 32.

Decision-point continuation removes the need for a reference answer or environment replay: an instance is the evaluated LLM's next turn on the recorded context, graded by a behavior-specific rubric on a 0-5 scale, with a three-LLM judge panel averaging scores and a mean of at least 4 counting as a pass. This makes traces from private repositories, internal tools, MCP servers, and other non-reproducible environments usable for evaluation, and every evaluated LLM receives exactly the same recorded context; the paper notes that OpenAI's production evaluations regenerate responses on deployment conversations, while TraceDance turns this into on-demand benchmarks for user-specified undesirable behaviors. Rubric quality was rated at least 4/5 by both annotators in 90% of 100 sampled instances, with a mean of 4.75/5; on the 84 instances both annotators confirmed and scored, judge-panel pass/fail agreement with humans is 81.0%, comparable to agreement between the two annotators, with Cohen's kappa of 0.40 for the panel versus 0.31 between annotators.

The benchmarks expose behavioral weaknesses that task-completion evaluation overlooks: nine frontier LLMs average a 26.7% pass rate (range 22.9%-33.5%), and grouping by rubric requirement gives 67.9% for Valid call versus only 8.1% for Check first. The paper reports that GPT-5.6-Sol scores higher than DeepSeek-V4-Pro and GLM-5.2 on task-oriented benchmarks such as SWE-bench Pro and Terminal-Bench 2.1 but does not outperform either on these behavior-focused benchmarks, and that overall rankings hide behavior-specific strengths, for example Claude Opus 4.8 ranks first overall yet trails Kimi-K3 on Error-guided correction, 27.8% versus 57.8%. Table 2 reports pass rates for the nine LLMs overall, by frame, and by setting; the Figure 4a grouping uses 102 single-behavior benchmarks and 3,966 instances, and the paper reports that the ordering of the four groups holds under alternative passing rules and individual judges; at the behavior level, Secret protection averages 6.9%, Commit hygiene 0.9%, and Failure-log inspection 0.6%.

Perspective

The method is aimed at agent developers: when developers have real deployment traces and can describe the undesirable behavior they want to test in natural language, TraceDance can produce targeted benchmarks on demand for comparing models, guiding improvement, and detecting regressions. It applies to the two deployment settings studied, coding and general tool use, and to environments that cannot be replayed, such as private repositories, internal tools, MCP servers, or accounts on external services, because evaluation uses only the recorded pre-decision context. The paper positions TraceDance as an evaluation component of the recursive self-improvement loop and explicitly leaves evaluating TraceDance within such a loop to future work.

Evaluation covers only the evaluated LLM's next turn at a recorded decision point, without executing actions or continuing the trace, so what happens after that turn is unobserved: a response that begins by planning or gathering information may act appropriately later, and a passing response may still fail when executed. Instances are cut from contexts in which the source LLM already exhibited the requested behavior, so pass rates describe performance at decision points where the behavior has occurred in practice, not its frequency in deployment. Rubrics credit only the most appropriate responses, so a response that avoids the undesirable behavior without meeting the full rubric can still fail, and absolute pass rates therefore depend on the passing rule; the paper reports that relative comparisons remain stable across alternative passing rules and individual judges. Construction and grading are themselves imperfect: both annotators confirm the requested behavior in 84% of sampled instances, the judge panel grades more leniently than humans, and GPT-5.6-Sol as a judge shows a relative preference of about 0.22 points on the 0-5 scale for its own responses. In addition, the loaded text includes the main body and appendices, but some table cells are not fully rendered in the text, so the specific numbers in those cells cannot be restated here.

Sources