Skip to main content
Back to timeline
arXivSource publication:

PAM encodes long-horizon intent as proprioceptive sketches, raising success from 47.5% to 75.0% on four real-world bimanual tasks

Related research and updates

Synopsis

The work proposes Proprioceptive Action Models (PAM), which jointly generate a compact, timing-free sketch of the robot's remaining joint-space path and a dense executable action chunk within a single transformer denoiser, with the sketch parameterized by arc length rather than time to capture geometric intent invariant to execution timing and with block-causal attention and a staggered denoising schedule maintaining directed sketch-to-action dependence; in simulation PAM improves over its action-only counterparts on Push-T and LIBERO-Long, and on four real-world bimanual tasks it raises success from 47.5% to 75.0%.

Source-provided article image: Proprioceptive Sketches as Long-Horizon Intent for Generative Action Policies
Figure 1 ·

Figure 1: PAM replans during a real-world dual-arm long-horizon task. At each step, PAM jointly generates a proprioceptive sketch 𝐙 t \mathbf{Z}_{t} of the future configuration-space path and a short action chunk 𝐀 t \mathbf{A}_{t} conditioned on the sketch, while only 𝐀 t \mathbf{A}_{t} is executed. For visualization, the sketch and the action chunk are mapped through forward kinematics and projected into the camera view.

arXiv

Interpretation

PAM jointly generates two outputs within a single transformer denoiser: a compact, timing-free sketch of the robot's remaining joint-space path and a dense executable action chunk. Prior long-horizon structure was often supplied by language plans, subgoal images, or video forecasts, which are costly to generate and still need to be translated into robot motion; directly predicting future motion avoids that translation, but a dense, time-indexed trajectory requires numerous parameters to cover the full remaining task and over a short horizon largely repeats the action chunk. PAM compresses long-horizon intent into a sketch and generates it jointly with the action chunk. Method description at the abstract level, stating that joint generation occurs within a single transformer denoiser; model scale, training data, and ablation details are not given.

The sketch is parameterized by arc length rather than time, expressing geometric intent that is invariant to execution timing. Relative to time-indexed dense trajectory prediction, arc-length parameterization decouples path geometry from execution speed, so long-horizon intent need not be carried as a time series. Design description at the abstract level; no quantitative validation of timing invariance is provided.

Block-causal attention and a staggered denoising schedule maintain directed sketch-to-action dependence, so action tokens condition on a progressively cleaner sketch throughout sampling. This supplies an explicit dependence direction and sampling mechanism for joint sketch-and-action generation, rather than generating the two independently or without ordering. Mechanism description at the abstract level; attention patterns and schedule parameters are not detailed.

In simulation PAM improves over its action-only counterparts on Push-T and LIBERO-Long; on four real-world bimanual tasks it raises success from 47.5% to 75.0%. Relative to generative policies that predict only action chunks, adding the sketch as a long-horizon intent representation yields gains on simulation and real-world bimanual tasks. The abstract reports simulation benchmark names and a real-world success-rate comparison (47.5% versus 75.0%); task list, trial counts, statistical uncertainty, and baseline configuration are not given.

Perspective

The work targets generative robot policies that need long-horizon intent and applies to joint-space action generation settings, with reported effects on the simulation benchmarks Push-T and LIBERO-Long and on four real-world bimanual tasks. For robot-learning researchers and practitioners who want to avoid costly intermediate representations such as language plans, subgoal images, or video forecasts, PAM offers a route that carries long-horizon intent in a proprioceptive sketch and directly produces executable action chunks within the same denoiser.

The abstract does not specify the content of the real-world bimanual tasks, trial counts, or statistical uncertainty, nor does it give quantitative magnitudes on the simulation benchmarks or ablation results, so the source of the sketch's gain over action chunks remains to be unfolded. How arc-length parameterization behaves on tasks with complex path geometry or precise timing requirements, and how the specific configuration of block-causal attention and the staggered denoising schedule affects results, are directions a reader can continue to watch. This document is an abstract-level parse and does not include figures or experimental details from the body, so the judgments above are limited to what the abstract states.

Sources