Skip to main content
Back to timeline
arXivSource publication:

D-OPCD distills agent-harness prompt experience into diffusion weights, lifting direct-generation average from 60.52 to 65.09

Synopsis

The authors propose Diffusion On-Policy Context Distillation (D-OPCD), which treats the agent-harness-produced prompt as privileged teacher context and supervises a student conditioned only on the original query at states along the student's own denoising trajectory, thereby writing prompt-mediated harness gains into the diffusion model's weights; paired with the Auto Skill Evolver (ASE), direct generation improves from 60.52 to 65.09 averaged over four benchmarks, and the updated generator under a skill-free harness slightly exceeds the original skill-equipped agent on average.

Source-provided article image: Internalizing Agent Experience into Diffusion Model Weights via On-Policy Context Distillation
Figure 1 ·

Figure 1: Method overview. As tasks accumulate, a T2I harness learns reusable skills that improve its prompts while the generator remains fixed. D-OPCD conditions an EMA teacher on the original query q i q_{i} and the harness-selected prompt p i p_{i} , and trains a student conditioned only on q i q_{i} to match the teacher’s predictions along the student’s own denoising trajectory. The updated generator can generate from q i q_{i} alone, while the harness resets its skills to continue adapting around it.

arXiv

Interpretation

D-OPCD uses the harness-produced prompt as privileged context: a teacher conditioned on both the original query and the harness prompt supervises a student that receives only the original query, comparing predicted velocities at states visited by the student's own rollout. Prior on-policy diffusion self-distillation draws privileged information from target or reference images, whereas this work substitutes the prompt constructed by an evolving harness, making prompt-mediated harness gains available in the generator's weights. The paper formalizes asymmetric conditioning, the on-policy rollout, and the objective, and compares against Vanilla SFT, Diffusion-DPO, and D-OPSD on matched records, the same Z-Image-Turbo weights, and the same LoRA optimization budget, with D-OPCD reaching the highest average of 65.09.

ASE organizes task execution trajectories into episodes, insights, and skills, and a skill manager commits mature insights into versioned skill libraries, each committed update defining a new harness version. Harness experience becomes a versionable, replayable data source for distillation rather than external guidance that must be retrieved and reapplied at every request, giving the harness-to-generator interface a concrete implementation. The paper specifies six insight actions and nine skill actions, maturity thresholds, capacity and length bounds, and benchmark-specific trigger and review-interval settings, and reports which skill-update snapshot was selected in each evolution round and how many tasks contributed experience.

After internalization, the generator beats the base generator on all four benchmarks in direct generation, raising the average from 60.52 to 65.09; with the skill-free harness reattached it averages 82.33, slightly above the original skill-equipped agent's 81.83. This indicates that part of the skill benefit can reside in the weights instead of being retrieved and applied as external guidance on every request. Results use each benchmark's official metric, with task sets split into disjoint evolution, selection, and reporting portions, and the signed internalization gap is also positive on R2I-Bench.

After clearing the internalized skills and rerunning ASE, the renewed skills improve three benchmarks and raise the average to 84.16, above both the updated skill-free agent and the original skill-equipped agent. This demonstrates an iterative harness-model co-evolution loop in which the generator absorbs old skills, the harness sheds saturated ones, and new guidance is evolved around the updated generator. The paper reports one subsequent round of skill evolution across four benchmarks and notes that R2I-Bench declines after skill evolution, which the authors attribute to the already high skill-free score leaving less room for additional gains.

Perspective

The result targets text-to-image agents whose gains flow mainly through prompt construction, verification, and iterative refinement, in settings where the harness and generator are separate and harness experience can be reduced to query-prompt pairs. For practitioners, this means part of the repeatedly retrieved and reapplied harness guidance can be written into the generator's weights once, so inference needs only the original query; the paper also demonstrates clearing the skill library after the weights absorb old skills and letting ASE evolve again around the updated generator. The authors note that extending this to image-generation agents that search the web, retrieve reference images, or use other tools requires first determining which reusable capabilities suit transfer into weights and which information should remain available through tools at inference.

The gains internalized into the weights are modest: direct generation averages 65.09, still well below the 81.83 of the original skill-equipped harness, and the authors suggest one possible explanation is that prompts produced by evolved skills, while effective when used by the harness, may be poorly aligned with the diffusion model's latent space for learning the same behavior from the original query alone. The study evaluates one base generator, one D-OPCD update, and one subsequent round of skill evolution, so retention of earlier gains over many update-and-reset cycles, transfer to new tasks, and the tradeoff between offline training cost and recurring harness inference cost remain open questions. In addition, the improvement filter and diversity adapter context designs affect the four-benchmark mean inconsistently, leaving open which harness-produced contexts best support model internalization; the decline on R2I-Bench after skill evolution also suggests that room for improvement differs across benchmarks.

Sources