APPL uses each skill's structural prior as both a training bias and a runtime interface, raising few-demo OOD success on MetaWorld from 37.9% to 89.6% and solving 8 of 16 unseen ManiSkill skill compositions
Synopsis
The work introduces Agent Priors-guided Policy Learning (APPL): a construction agent segments complete demonstrations into reusable skills, proposes several structural priors per skill and trains one Diffusion Policy per prior, then writes the same prior into a runtime interface so a runtime agent can select and compose the frozen policies; on six MetaWorld tasks it raises OOD success with two demonstrations from 37.9% for a fixed relational prior to 89.6%, and on five long-horizon ManiSkill tasks it reaches 50.0% success under shifted objects (versus 10.0% for a full-task Diffusion Policy), 92.5% on task-level variants, and solves 8 of 16 previously unseen skill compositions, while hiding the interface information substantially reduces performance.
Interpretation
The paper uses each skill policy's structural prior in two roles: during training it shapes where the policy generalizes through representation, action parameterization, architecture, or an auxiliary loss, and at runtime the same prior is exposed in the interface together with handoff conditions, measured training support, and verification evidence so the runtime agent can judge where that policy applies. Prior work either used structural priors only for training (for example LGA and KALM each automate one family of priors and produce one design per task) or described skills at the composition level only by name, instruction, or symbolic operator, losing the policy's structure; APPL carries the same prior across both levels. The paper formalizes faithfulness and informativeness conditions for the interface and compares, in two experiment suites, ablations that hold the trained policy library fixed while varying the information exposed at runtime (hiding interface information, hiding prior information, removing the high-level agent, removing verification).
On six MetaWorld tasks with 20 demonstrations each in a few-shot setting, agent-designed priors substantially improve out-of-distribution skill generalization: with two demonstrations APPL reaches 89.58% OOD success versus 28.96% for vanilla Diffusion Policy and 37.92% for a fixed relational prior. Relative to the fixed relational prior, APPL is ahead by 51.67, 50.83, 41.04, and 36.88 percentage points at 1, 5, 10, and 20 demonstrations; IID success is similar across methods at larger demonstration counts, indicating the gains come from generalization rather than fitting. Six tasks, four demonstration counts, and six trained systems per task-condition, 144 models in total, tested on 20 IID, 40 C, and 40 E states with 95% stratified bootstrap intervals; the authors state the intervals are conditional on the trained models and the six tasks and do not measure variation across training seeds or independent agent designs.
On five long-horizon ManiSkill tasks with twelve demonstrations each, APPL succeeds on 50.0% of the eight frozen motion-OOD layouts, versus 10.0% for both the full-task Diffusion Policy and the single-prior baseline, and reaches 92.5% on task-level variants that resume mid-task or request early termination. This is a system-level comparison that combines runtime agent scheduling with a larger library of skill policies trained with diverse agent-designed priors; the paper explicitly notes the comparison does not isolate scheduling from library size. Eight motion-OOD layouts, eight task-level cases, and sixteen composition cases per task, with paired McNemar tests; a scripted planner completes all task-level and composition cases and 37 of 40 motion-OOD layouts.
Keeping the same verified policy library but hiding the interface information drops motion-OOD success from 50.0% to 20.0%, task-level success from 92.5% to 65.0%, and composition from 8/16 to 2/16; replacing the runtime agent with a fixed rule gives 22.5%, 52.5%, and 1/16. These ablations hold the trained policy library fixed and vary only the information exposed at runtime, attributing the gains to interface information and online policy selection rather than to the policies themselves. The paper reports paired win-loss counts and p-values (for example 17 wins to 5 losses on motion OOD against the no-interface-information ablation, p=0.017, and 6 wins to 0 losses on composition, p=0.031), and notes that restoring verification scores and exit criteria recovers most of the motion-level gain while the prior, handoff, and support descriptions contribute more on task-level variants.
Perspective
The results target simulated manipulation settings where reusable skills are learned from a few complete demonstrations and then composed under shifted object positions or task variants, assuming structured state observations, segmentable demonstrations, and a shared inverse-kinematics interface. They enable follow-up work on improving verification (for example, aligning it with deployment states and handoff rules), expanding skill support through targeted data collection or augmentation, and integrating complementary tools such as motion planners; for a reader, the implication is that information about why a policy generalizes can remain available to downstream agents after learning instead of being discarded when training ends.
The paper itself notes that a structural prior describes intended generalization rather than guaranteeing competence, that a policy may still fail outside its demonstrated support, and that the current library may not suffice for recovery; verification uses only demonstrated entry states and can misrank policies under shift (for example, success in covered peg assembly fell from 8/8 to 1/8). Common failures follow a dropped object, a drawer pushed closed, or a skill invoked outside its trained entry conditions, and the demonstrations contain no recovery trajectories. On evaluation, task-level success uses an executor that interrupts at first goal attainment, so it measures goal attainment rather than unaided stopping or sustained stability, and the motion-OOD split was chosen after results on a harder split were seen by halving the out-of-range displacement. In addition, re-running the language-model agents may produce different policy designs and runtime decisions, and exact numerical agreement also depends on the software and hardware environment.
