Invent-A-Dataset generates post-training datasets from zero data, reporting 17% relative quality gains and 19% relative diversity gains
Synopsis
This technical report introduces Invent-A-Dataset, a prompt-based system that goes from a dataset description to realistic and large-scale post-training datasets; evaluated in a zero-data regime against five frontier model APIs (Anthropic, Google, OpenAI, DeepSeek, Zai) across eight task types and dataset sizes up to 20K samples, it reports the highest quality (17% relative gains) and the most diverse samples (19% relative gains), with its diversity advantage widening with scale (from parity at 200 samples to 37% relative gains at 20K samples), translating into downstream training gains and higher rankings for its fine-tune across post-trained model architectures.
Figure 1 : Diversity vs. quality trade-off. Quality (x-axis) versus diversity (y-axis), measured by DCScore with τ = 0.05 \tau=0.05 , for four external models and the Invent API. Both scores are averaged over eight 5K-sample datasets in unconstrained setting. Higher is better on both axes; the upper right is optimal. The Invent API attains high quality and high diversity simultaneously.
arXivInterpretation
Introduces Invent-A-Dataset, a prompt-based system that goes from a dataset description to realistic and large-scale post-training datasets, targeting the zero-data regime as the most extreme yet most prevalent setting. Rather than treating dataset building as a manual and brittle process, the work treats description-to-dataset generation itself as a measurable capability. Presented as a technical report; the visible text gives the system positioning and evaluation setting without implementation details.
Across eight task types and dataset sizes up to 20K samples, Invent-A-Dataset achieves both the highest quality (17% relative gains) and the most diverse samples (19% relative gains). Compared against five frontier model APIs (Anthropic, Google, OpenAI, DeepSeek, Zai), the reported result is simultaneous advantage in quality and diversity rather than a single-metric trade-off. The abstract reports relative gain figures and comparison targets; evaluation protocol and statistical details are not expanded in the visible text.
The diversity advantage widens with training dataset size: parity at 200 samples and 37% relative gains at 20K samples. Scale is treated as a variable, indicating the advantage is not fixed but grows with data volume. The abstract provides comparison values at two scale points; the trend between them is not described in the visible text.
These quality and diversity advantages translate into downstream training gains, with the Invent-A-Dataset fine-tune ranking higher than other generator fine-tunes across post-trained model architectures. Links dataset-level metrics to the actual performance of post-trained models rather than stopping at dataset-level measurement. The abstract reports downstream training gains and a cross-architecture ranking conclusion; specific models, tasks, and ranking basis are not listed in the visible text.
Perspective
The work targets the zero-data regime: practitioners have no data for the capability they want to learn, only a dataset description. Under that premise, Invent-A-Dataset is positioned as a prompt-based system that goes from description to realistic and large-scale post-training datasets, suited to development settings where post-training data must be obtained quickly; the evaluation covers eight task types and dataset sizes up to 20K samples, and reports how its fine-tune ranks across post-trained model architectures. For teams seeking to reduce manual data construction and start post-training without seed data, this setting directly matches their working conditions.
The visible text is an abstract and does not include figures, task lists, prompt design, or the specific measures of quality and diversity, nor does it describe the API configurations or fairness settings for the five frontier model APIs. The evaluation basis, sample construction, and statistical significance behind the 17%, 19%, and 37% relative gains, as well as the specific models and tasks behind the downstream training gains and the higher cross-architecture ranking, therefore remain to be confirmed in the full text. The scale trend is given only at the 200 and 20K endpoints, so the shape of the middle range is also an open question.
