Skip to main content
Back to timeline
arXivSource publication:

ToolRACER generates 5.6K conversation trajectories with 66% failure-prone scenarios via three-agent emulation and four-stage validation, lifting fine-tuned Qwen3-4B to 39.20% macro Pass@1 on τ-bench

Synopsis

The authors present ToolRACER, a three-stage synthetic data pipeline coordinating user, assistant, and tool emulation agents that injects user-originated and environment-originated adversarial behaviors, validates traces through four evaluators plus a refinement loop, and yields ToolRACERBench: six domains, 55 personas, 5.6K trajectories with about 66% failure-prone scenarios; models fine-tuned on it improve end-to-end agentic accuracy on τ-bench, BFCLv3, and ACEBench over baseline training sets.

Source-provided article image: ToolRACER: A Robust Agentic Conversation Emulation Resource for Agent Training and Evaluation
Figure 1 ·

Figure 1: Examples of ToolRACERBench conversation trajectories: (a) an Unhappy path and (b) an Impossible path.

arXiv

Interpretation

ToolRACER encodes robustness directly into data generation: three LLM agents (assistant, user, tool) each maintain a distinct view of conversation history, orchestrated by a central runner, and generate traces from skeleton scenarios classified as Happy (non-adversarial users, achievable tasks), Unhappy (adversarial users, achievable tasks), and Impossible (non-achievable tasks). Prior multi-agent simulation training corpora such as APIGen-MT emphasize successful, cooperative interactions, leaving user-side adversity and tool-side failure underrepresented; this work models both user-originated and environment-originated failures as explicit, classifiable trajectory types. The paper defines all three path types and lists injected behaviors (partial information, incremental constraint revelation, goal revision, adversarial attitudes; backend failure, empty results, timeout, policy constraint), with Table 1 giving per-domain and per-path counts of 3,762 training and 1,881 test trajectories.

Four automatic validation checks (syntax, faithfulness, role confusion, task success) plus a refinement loop substantially raise usable data yield: without feedback only 26.3% of simulations pass validation, and the refinement loop recovers 65.4% of initially failed conversations. Failure reasons from LLM-as-a-judge evaluators are converted into refinement hints and fed back into simulation, rather than discarding failed traces or regenerating from scratch, improving validated yield at equal generation budget. Appendix A Table 6 reports per-path success rates (Happy 32.2%, Unhappy 23.9%, Impossible 23.6%) and recovery rates (72.6%, 62.3%, 75.0%), with Unhappy paths hardest to repair.

Fine-tuning on ToolRACERBench improves small-model end-to-end agentic accuracy and cross-domain generalization: fine-tuned Qwen3-4B-Instruct reaches 39.20% macro Pass@1 on τ-bench, and mixing with APIGen-MT reaches 52.5% on retail and 41.25% macro. Against APIGen-MT alone (36.95% macro) and open-source function-calling models (ToolACE-MT 20.60%, xLAM-2-3B-FC-R 38.20%), this work achieves higher or comparable results with a smaller corpus (3.7K vs 5K), indicating that failure-prone trajectories themselves carry transferable training signal. Table 3 reports BFCLv3 single-turn and hallucination subcategories plus τ-bench airline/retail Pass@1; Table 5 reports ACEBench end-to-end accuracy rising from 8.4% to 10.0% for Qwen3-4B and from 19.2% to 25.0% for Qwen2.5-32B.

Data-quality comparison shows ToolRACERBench leads on diversity while using shorter dialogues: 11.8 tools and 12.6 turns per dialogue versus 15.1 and 18.5 for APIGen-MT-5k, but higher entropy 9.29, Distinct-3 0.348, dialogue Vendi 23.8, and tool-trajectory Vendi 18.8. This indicates the robustness gains come not from longer conversations or more tool calls but from the trajectory-distribution diversity introduced by personas and failure injection. Table 2 compares both corpora using metrics from Wang et al. (2025c); ToolRACERBench scores higher on five of six automatic evaluations, with the one lower metric, entailment rate (0.042 vs 0.071), explained by the authors as reflecting that real user conversations are more variable.

Perspective

This resource targets researchers and engineering teams training or evaluating tool-calling conversational agents, in settings where multi-turn interactions involve function calls and where user behavior and tool environments may deviate from ideal paths. The three-path taxonomy and injected-behavior list can be used directly to construct in-domain skeleton scenarios; the refinement loop's failure-reason feedback mechanism can raise validated yield under a limited generation budget. Cross-domain results suggest that mixing failure-prone trajectories with in-domain data is a viable configuration for balancing robustness and task performance in small models.

Readers should still watch: how the gap between synthetic trajectories and real user distributions affects downstream deployment; the rise in end-to-end accuracy alongside a drop in process accuracy on ACEBench leaves the improvement path for long-horizon multi-turn tasks unclear; the lower entailment rate relative to the comparison corpus is given an explanatory account rather than a controlled experiment; and whether gains hold across other base models and larger scales is not concluded in the text. If later versions add figure details, the stability of each ablation combination can be checked further.

Sources