ServiceNow CoreAI published a description of AutoSynthData, a system designed to convert a target model’s failures into validated training tasks for agents in enterprise environments. The pipeline creates executable tasks that reflect the tools, policies and state of a specific environment, and selects tasks that expose weaknesses the target model still exhibits.
How AutoSynthData works
AutoSynthData represents each training item as a tuple: system specification, user prompt and verifier. The system specification captures environment constraints, policies and any task initialization. Generated tasks must be feasible (solvable in the environment), realistic (plausible user requests) and difficult enough to provide new training signal.
The verifier is required to be consistent with the prompt and environment state, sound in rejecting incorrect outcomes, and complete enough to accept valid but different solution trajectories. ServiceNow CoreAI emphasizes that lax or overly strict verifiers reduce the dataset’s usefulness.
The pipeline evaluates the target model and a stronger teacher on diagnostic tasks, distills capability gaps into sanitized specification cards, and uses those cards to generate many new tasks. Generation happens in two phases: a target phase that produces core validated samples, and a multiply phase that creates vetted variants anchored to the core set.
Validation, repair and dataset control
Candidates pass a quality loop including solver evaluation, positive and negative verification, and a critic-driven repair process with limited retries. Positive verification runs a reference solution in the environment; negative verification checks that mutated incorrect outcomes fail. A batch-level meta-review then adjusts coverage and generation guidance to avoid overrepresentation and to close capability gaps.
ServiceNow CoreAI reports experiments on EnterpriseOps Gym. In the Hybrid domain, AutoSynthData generated 2,000 samples in about 18 hours and, after supervised fine-tuning of Gemma-4-26B-A4B-it using Qwen3.8-27B as teacher, improved mean Pass@1 by 7.2 percentage points (a 35% relative gain) and raised verifier success from 63.01% to 68.55%, closing 59% of the original Pass@1 gap with the reference model. In the ITSM domain, 1,994 samples were generated in 66 hours using DeepSeek-V4.1-Flash as teacher; synthetic SFT raised mean Pass@1 from 18.77% to 27.18%.
ServiceNow CoreAI notes that the approach treated synthetic data generation as a moving frontier: after post-training, evaluation guides the next round of task generation. The published account (October 2, 2026) focuses on supervised fine-tuning but states the same mechanism could be applied to reinforcement learning.
Original source: Hugging Face Blog