Skip to main content
Back to timeline
Hugging Face - BlogSource publication:

ServiceNow CoreAI's AutoSynthData turns failure diagnoses into enterprise-agent training tasks, lifting Gemma's Pass@1 by 7.2 points on EnterpriseOps Gym

Synopsis

ServiceNow CoreAI introduces AutoSynthData, which uses a target model's failures and a stronger teacher's successes to locate capability gaps and then generates and validates new executable tasks (system specification, user prompt, verifier); on the Hybrid domain of EnterpriseOps Gym, with Gemma-4-26B-A4B-it as target and Qwen3.8-27B as teacher, it generated 2,000 samples in about 18 hours, and after fine-tuning the best checkpoint (epoch 5) raised mean Pass@1 by 7.2 percentage points (a 35% relative improvement), lifted verifier success from 63.01% to 68.55%, and closed 59% of the original Pass@1 gap between Gemma and the reference model; on ITSM, with DeepSeek-V4.1-Flash as teacher, it generated 1,994 samples in 66 hours and raised mean Pass@1 from 18.77% to 27.18%.

AI-generated editorial illustration: AutoSynthData: Generating Training Data for Enterprise Agents

Interpretation

It converts a model's failures in a specific enterprise environment into trainable data: AutoSynthData first evaluates the target model on diagnostic tasks in the environment while a stronger teacher runs the same tasks, and from those runs it identifies the capability being tested, the tools and workflow structure involved, where the target fails and how the teacher succeeds, the properties a correct final state must satisfy, and the dimensions that can vary while preserving the capability. Rather than reusing evaluation tasks or hand-authoring tasks, it distills the evaluation runs into sanitized capability specification cards; the generator receives only the cards and not the original prompts, entities, trajectories, or verifier details, so new tasks differ in prompts, states, and solution paths. The pipeline is illustrated with experiments in the Hybrid and ITSM domains of EnterpriseOps Gym, and the text states that training tasks were newly generated from capability specifications and that the generator did not receive the original evaluation tasks; the number of specification cards or the scale of human review is not reported.

It states quality criteria for tasks and verifiers: tasks should satisfy feasibility (at least one trajectory in the current environment satisfies the prompt while respecting the system specification), realism (the prompt resembles something a user would plausibly ask), and difficulty (the task exposes a weakness of the current agent); verifiers should satisfy consistency, soundness, and completeness, agreeing with prompt, specification, and task state, rejecting failing or policy-violating trajectories, and accepting valid solutions rather than encoding one reference trajectory. It separates generating a plausible request from generating a useful training sample, and notes that a lax verifier can reward incorrect behavior while an overly restrictive verifier can penalize valid solutions, tying verifier quality directly to training-signal quality. These properties are given as definitions and design principles, paired with a positive gate (does the intended solution pass the verifier) and a negative gate (do mutated expected outcomes fail); per-gate pass rates or error rates are not reported.

Two-phase dataset construction with two levels of quality control: the target phase generates core samples in parallel, with each candidate going through validation, execution, solver evaluation, and repair before acceptance; the multiply phase creates novel variants of accepted samples, each with its own user request, environment state, entity configuration, reference trajectory, and verifier, and a multiplied sample cannot seed another multiplied sample, which limits drift across generations. For difficulty, the configuration favors tasks the target model solves on no more than one of three trials and a stronger solver solves on at least two of three. Sample-level checks repair individual candidates (a critic diagnoses the failure and guides targeted repair, with a fixed retry limit), while batch-level meta-review checks whether task families are overrepresented, capability dimensions are missing, the same kinds of examples repeat, or particular targets keep failing generation; the controller then reduces generation in overrepresented regions and directs work toward gaps. The process description is fairly complete, including the fixed difficulty thresholds and retry limit; the rejection rate, repair success rate, and gains from batch-level adjustment are not reported.

It demonstrates effectiveness in two enterprise domains: in Hybrid, with Gemma-4-26B-A4B-it as target and Qwen3.8-27B as teacher, about 18 hours produced 2,000 samples, the best checkpoint was epoch 5, mean Pass@1 rose 7.2 percentage points (35% relative), verifier success went from 63.01% to 68.55%, and 59% of the original Pass@1 gap between Gemma and the reference model was closed; in ITSM, with DeepSeek-V4.1-Flash as teacher, 66 hours produced 1,994 samples and mean Pass@1 rose from 18.77% to 27.18%. The same failure-driven generation mechanism improves performance in two different domains, indicating it is not tied to a single environment; ITSM generation was slower, which the text attributes to a larger teacher model and to that run preceding pipeline optimizations that improved throughput. Concrete sample counts, generation times, best epoch, and before/after metrics are reported; the experiments focus on SFT and evaluation is within the same environment used for generation, with no cross-environment transfer or comparison against other data-synthesis methods reported.

Perspective

The work targets enterprise agents that must operate within their own systems, rules, and data states, and it applies where an executable environment, replayable reference trajectories, and deterministic verification exist, such as stateful enterprise environments like the Hybrid and ITSM domains of EnterpriseOps Gym. For a reader, it offers a reusable process: diagnose failures, distill capability specifications, generate and validate tasks, post-train on accepted samples, and re-locate remaining gaps after the model updates. The text states the experiments focus on SFT and that the same mechanism is planned for RL, generating tasks that challenge the current policy, training, and then moving the generation target with the updated policy.

The text does not report quality-gate pass rates, the share of rejected candidates, or repair success rates, nor does it compare against other data-synthesis methods, so it is hard to tell how much of the gain comes from this pipeline itself. The 66-hour ITSM generation is attributed to a larger teacher model and to preceding throughput optimizations, indicating that generation cost is sensitive to teacher size and pipeline implementation. In addition, evaluation is within the same environment used for generation, so cross-environment transfer and the moving difficulty frontier under RL remain open questions. The loaded text is incomplete and lacks figures and appendices, so details such as the concrete form of the specification cards and the prompt design of the critic and meta-review cannot be confirmed from the available content.

Sources