ProgramDistill: Turning Interactive Web Apps into Verifiable Reference-Guided SWE Tasks
Synopsis
This work introduces the ProgramDistill benchmark and a fully automated mine–craft–patch pipeline that factorizes 26 interactive web applications into 1,975 replay-verified behaviors and constructs 4,063 tasks, letting coding agents infer and restore missing functionality by interacting with a reference application whose source is hidden; across nine frontier agents, the best full-application reconstruction success is 49.2%, and partial-application reconstruction success falls from 100% to 64.0% and from 96% to 32% as restoration depth increases from 1 to 8.
Figure 1: ProgramDistill evaluates reference-guided software engineering for interactive web applications. A coding agent is given a working reference application whose source implementation is hidden and an editable current application with missing functionality. The agent interacts with the reference to infer the intended behavior, restores that behavior by modifying the current implementation, and validates the result against the reference.
arXivInterpretation
Introduces ProgramDistill, a benchmark for reference-guided software engineering comprising 4,063 replay-verified SWE tasks derived from 1,975 behaviors across 26 interactive web applications, spanning atomic repair, cumulative repair, and full-application reconstruction. Prior coding-agent evaluations typically specify desired behavior through issues, instructions, or tests, whereas this benchmark requires agents to infer behavior from a working reference and judges success by whether the replayed interaction executes identically. Tasks are generated by a fully automated pipeline, and each behavior is admitted only when the replay verifier returns V(A,L_d)=1, with gold patches providing an end-to-end positive control.
Formulates program distillation along two axes: application-to-task factorization converts working applications into structured SWE tasks, while reference-to-current distillation requires agents to recover behavior from an executable reference. Unlike ProgramBench, which treats the program as a whole-program reconstruction target, this work exploits the prerequisite dependencies formed by UI actions and application state in interactive web applications, organizing behaviors into prerequisite lineages. Both axes are operationalized by the mine–craft–patch pipeline across 26 applications without human intervention, with 629 of 1,201 cumulative tasks composed deterministically and 572 requiring a merge agent.
Introduces restoration depth as a controlled difficulty axis, where prerequisite lineages provide a progression from atomic repair to full reconstruction and a natural basis for future curriculum-based training. This axis lets a single benchmark systematically expose failure patterns that emerge as restoration depth increases, rather than reporting only a single aggregate success rate. Reports results across nine frontier coding agents: in full-application reconstruction, GPT-6 Astra and Claude Opus 5 recover 49.2% and 28.8% of evaluated workflows; in partial-application reconstruction, success falls from 100% to 64.0% and from 96% to 32% as restoration depth increases from 1 to 8.
Trajectory analysis reveals a growing mismatch between reconstruction burden and agent effort, with observation effort per required behavior declining especially sharply as tasks deepen, and Astra achieving the strongest repair performance while exhibiting the highest observation activity and the fewest edit/write steps. Identifies effort allocation across observation, validation, and editing as an important dimension of reference-guided software engineering alongside implementation capability. Based on analysis of trajectories from nine frontier coding agents, constituting an observational summary of behavioral patterns during evaluation.
Perspective
The benchmark targets the setting of interactive web applications: tasks are generated by the mine–craft–patch pipeline from 26 applications, including ones adapted from the OSWorld web-application suite and real-world open-source projects and SaaS clones from public repositories, and evaluation occurs in an isolated environment where the reference source is hidden and public-network egress is disabled, with success defined by whether replayed interactions match the reference rather than by source-level matching. It is suited to evaluating and diagnosing coding agents on atomic repair, cumulative repair, and full-application reconstruction, and to supplying tasks for future curriculum-based training; applicability to non-web software or systems whose behavior cannot be replayed through browser interaction would need separate examination.
A careful reader might still watch: replay verification relies on deterministic execution and stable observable attributes, with time-dependent behavior reduced by a shared deterministic clock, so how broadly this setting remains stable across behaviors is an open question; the mask-depth critic rejects superficial changes, and its judgment boundary is worth understanding further; cumulative tasks increasingly rely on a merge agent as restoration depth grows, so the influence of merge quality on task validity is a direction for further observation; additionally, the loaded text is the paper body, and appendix details on the corpus, prompts, and implementation are not included, so those appendices are the open part to consult when assessing pipeline reproducibility.
