Skip to main content
Back to timeline
arXivSource publication:

Prompt2Skill builds skills automatically from a single natural-language task description, averaging a 36.1% relative gain across four benchmarks and five models

Synopsis

Prompt2Skill is a multi-agent framework that, from a natural-language task description alone, has a Data Agent derive a task specification and build a proxy data pool through retrieval, adaptation, and synthesis, and a Skill Agent refine the skill in a closed loop of execution feedback, reflective editing, and paired statistical screening, improving over direct prompting by 36.1% and over off-the-shelf skills by 83.0% on average across question answering, reading comprehension, spreadsheet manipulation, and mathematical reasoning with five open-source and commercial target models, without significant regressions.

AI-generated editorial illustration: Prompt2Skill: Unsupervised Skill Optimization From Natural Language Instructions

Interpretation

The work moves skill construction from a setting that requires a curated, in-distribution labeled training set to one that needs only a task description: the system receives a description and optional few examples, never observes the target distribution or its metric, and exports a SKILL.md consumed by a frozen model at inference time with no weight updates. Prior work such as Trace2Skill consolidates execution trajectories and SkillOpt performs validation-gated text-space edits, both presupposing labeled training tasks; the paper notes that shrinking the training set to a single example collapses gains to the level of a data-blind draft, an assumption Prompt2Skill removes. The paper formalizes the prompt-to-skill problem, where neither the target distribution nor the metric is observed, and specifies the full two-agent procedure and algorithm.

The Data Agent constructs task-matched proxy data by retrieving candidate sources from the Hugging Face dataset catalog and Wikipedia, adapting records into input-reference pairs via column selection and deterministic row mapping, generating new items from retrieved material when needed, and filtering through structural checks, task-compatibility comparison, a seed-skill answerability screen, and a consistency check for synthesized items. Unlike prompt-optimization lines such as APE, ProTeGi, and GEPA, or skill-document editing such as SkillOpt, this work chains data acquisition and skill optimization into one end-to-end pipeline. The paper gives formula-level descriptions of retrieval, adaptation, generation, and validation, and states that these checks provide practical evidence of usability, since solver agreement alone does not establish reference correctness.

The Skill Agent refines the skill through reflective edits and a paired statistical acceptance rule: each round compares a candidate against the incumbent on the same fresh acceptance batch, accepting only when the direction of improvement is positive and the statistic exceeds a threshold, after which a frozen selection set determines the exported skill in a single final stage. An appendix result bounds the selection loss under proxy distribution and evaluator mismatch, showing that retaining the initial skill and previously accepted skills and comparing them on an independent selection sample keeps the exported skill within a stated error allowance of the best retained skill. The acceptance rule is a per-candidate screening threshold rather than a conventional significance criterion, and the implementation does not adjust for comparisons across candidates or rounds; the paper states this explicitly and relies on the separate frozen selection set as a common basis.

Across SearchQA, SQuAD, AIME, and SpreadsheetBench with Qwen3-8B, Qwen3-32B, Llama-3.2-1B-Instruct, Claude Haiku 4.5, and GPT-5.5, Prompt2Skill achieves the highest average relative improvement for every target model, ranging from 22.0% to 29.1% for the Qwen and commercial models and 93.5% for Llama-3.2-1B over its three reported benchmarks, and attains the highest mean score in 10 pairs, including all five solvers on SQuAD and all four benchmarks for Qwen3-32B. Fixed skills can regress substantially, for example the off-the-shelf skill reduces Llama-3.2-1B's SQuAD score from 0.207 to 0.048, whereas Prompt2Skill raises it to 0.492, illustrating that a skill's usefulness depends on the target model and task. The main table reports means and standard deviations per method and benchmark; the paper notes these are averages of relative changes rather than percentage-point accuracy gains, and that a pair with a small baseline score can contribute a large relative change.

Perspective

The result targets deployers who have a task description but no in-distribution labeled data: the skill is a text artifact injected into a frozen model at inference time, requires no weight updates, and applies to tasks measurable by an executable proxy metric, such as question answering, reading comprehension, spreadsheet manipulation, and mathematical reasoning. The paper also explores NLP AutoML, where on Japanese-to-Python code generation and temporal expression normalization Prompt2Skill scores comparably to a modified Prompt2Model baseline that generates up to 1,000 training examples and fine-tunes a fixed Qwen-32B model, suggesting skill-based task adaptation as an alternative to fine-tuning. On auditing, the paper applies text and workbook-hash checks to retrieved sources and reports that no audited input was flagged.

The paper states that the proxy metric implements the system's interpretation of the request while the target metric defines true task performance, so distribution and evaluator mismatch from the data agent remain a separate source of error from selection error; the acceptance threshold is a per-candidate screen rather than a significance criterion, and no adjustment is made for comparisons across candidates or rounds. On the audit, the paper notes that the listed sources are configured or discovered source banks rather than verified item-level attribution, that source cursors indicate attempted fetching rather than how many items were retained, that HotpotQA was recorded for SQuAD but had no bulk-fetch progress in most runs, and that three manifest items have no declared workbook files and receive text checks only. The conclusion's 36.1% and 83.0% are equal-weighted averages of relative changes over 19 model-benchmark pairs, not percentage-point accuracy gains, and a pair with a small baseline can contribute a large relative change. This reading covered the full text, but some table and appendix values appear as text, so exact reproduction should still consult the original tables and appendices.

Sources