Skip to main content
Back to timeline
arXivSource publication:

EdgeGen extracts compliance rules from agent policy documents and generates database-grounded rule-violating tasks, lifting finetuning mean progress by 2% to 42% on the tau2-bench airline domain and improving Gemma-4-e4b harness optimization by 30% over the base harness

Synopsis

EdgeGen automatically extracts compliance rules from an agent's policy document, enumerates combinations of rule violations, and grounds each scenario in executable database states via a scenario-aware SQL agent, admitting tasks only after data-grounding and scenario-feasibility checks; combined with a synthetic database generation method it forms a fully automated closed loop requiring no human annotation, yielding a 2% to 42% mean progress improvement from finetuning on the tau2-bench airline domain while some baselines degrade for some models, and harness optimization improving 10% and 30% over human-curated and base harnesses for Gemma-4-e4b.

AI-generated editorial illustration: EDGEGEN: Improving Tool-Calling Agents Beyond Happy Paths with Synthetic Edge Case Generation

Interpretation

It proposes EdgeGen, a violation-driven synthetic task generation framework that extracts compliance rules from an agent specification, enumerates the power set of those rules to construct violation scenarios, uses a scenario-aware SQL agent to sample database values that make the task feasible, and admits a task only after data-grounding and scenario-feasibility verification. Prior synthetic task generation methods such as TaskBench and FuncBenchGen largely produce generic function-calling tasks, do not systematically enumerate a structured space of behavioral rule violations, and are not grounded in a database state; EdgeGen explicitly brings the question of when an agent should refuse into task construction. The paper describes a five-stage pipeline (tool graph construction and path sampling, violation scenario generation, scenario-aware database sampling, test case generation, verification) and states that verification is performed by a SQL agent executing deterministic SQL queries; rule extraction quality was audited against manual annotation, with airline precision 90.0%, recall 87.5%, F1 88.7%, retail 100.0%/95.2%/97.6%, ToolSandbox 100.0%/100.0%/100.0%, and overall 94.6%/92.1%/93.3%.

It chains synthetic database generation, EdgeGen task construction, and synthetic-data-based agent optimization into a fully automated closed loop that needs no production logs, human-annotated tasks, or real data, yet still improves performance on human-curated test sets. Existing agent benchmarks each provide a curated set of test cases but no synthetic-data layer for automatically expanding evaluation to support agent improvement; this loop connects environment state, task generation, and optimization. The paper reports supervised finetuning and harness optimization on the tau2-bench airline domain and ToolSandbox; finetuning data came from gpt-5.4 as the agent model over 8 trials per sample, giving 120 attempted rollouts per benchmark, with only successful traces (progress rate 1) retained for supervised finetuning.

On the tau2-bench airline domain, finetuning on EdgeGen-generated data improves every evaluated model, whereas human-curated data and TaskBench degrade performance for some models. The paper draws a coverage hypothesis from this: tasks with less coverage and no violation scenarios are insufficient and can be actively harmful, while EdgeGen explicitly maximizes coverage over combinations of compliance rules. The paper reports a 2% to 42% mean progress improvement from finetuning; human-curated data regressed on 3 of 6 models (up to about 8 points on gemma-4-e2b) and TaskBench regressed on 4 of 6 models; the aggregate win-rate table reports EdgeGen vs. Base 100.00%, vs. Naive 81.25%, vs. Human-Curated 86.67%, vs. TaskBench 94.12%, vs. FuncBenchGen 100.00%, which the paper describes as descriptive comparisons rather than statistical significance tests.

The same EdgeGen samples can optimize the harness for black-box models: on Gemma-4-e4b mean progress rises from 0.43 to 0.56, a 30% relative improvement over the base harness and 10% over the harness optimized on human-curated data, while using fewer tool calls and tokens. A case study shows the human-curated training split concentrates on approve-branch workflows, so the optimized system prompt becomes a long checklist hard-coding business rules case by case; the EdgeGen split allocates half its budget to denial and verification cases, and the resulting prompt collapses into a short discover-verify-confirm-write protocol that references the policy document and tool descriptions instead of enumerating them inline. Table 2 reports for Gemma-4-e4b base harness mean progress 0.4301, human-curated optimized 0.5083, EdgeGen optimized 0.5608, with tool calls 5.8/8.8/6.9 and tokens 81.0k/121.0k/85.0k; for gpt-5.4 the EdgeGen harness is within a 2% to 4% relative difference of the human-curated counterpart on mean and maximum progress while using fewer tool calls and tokens; on ToolSandbox with gpt-4o-mini the EdgeGen-optimized harness reaches mean progress 0.923 versus 0.931 for human-curated, with higher token and tool usage.

Perspective

The result targets enterprise tool-calling agent development settings that have a structured policy document and a relational database, and it supports two downstream uses: supervised finetuning on the generated trajectories, and harness optimization for black-box models that cannot be finetuned. The paper states EdgeGen is agnostic to the database source and can use any compatible relational database, so synthetic database generation and task generation can be swapped independently. The main beneficiaries are small and mid-sized models: the paper observes larger gains for Qwen 2.5-3b and the Gemma models, with smaller headroom for larger models already near saturation on the test set.

The paper states its evaluation is limited in scale, with main experiments using 10 to 15 evaluation scenarios per benchmark and supplementary scaling experiments using 21 airline scenarios, so larger evaluation sets and broader benchmark suites remain open. The complexity ablation shows gains are not monotonic in the number of violations: for Qwen2.5-3b maximum progress is 0.67 at one violation but drops to 0.24 and 0.38 at two and three violations, with the two-violation setting below the 0.35 base. Data scaling shows human-curated test mean progress peaks at 30 tasks and declines at 60, and improvements on synthetic validation do not consistently transfer to the human-curated test set. In the human study, among runs marked successful, 6 of 10 airline runs and 4 of 10 ToolSandbox runs were verified correct, with the rest involving loosely defined assertions, judge errors, or missing task details; most sampled failures reflect genuine agent limitations, with a minority from task or data generation artifacts. The paper also notes that verification does not guarantee every extracted rule or assertion is semantically correct, and that rule extraction depends on a structured policy document. These are scope and open questions that bound how far the conclusions apply.

Sources