Skip to main content
Back to timeline
灵初智能 PsiBotSource publication:

PsiBot releases Psi-R2.5: reverse generation mass-produces strong Pair Data, and HIL+RL post-training lifts on-site success rate to 99%

Synopsis

PsiBot released the embodied-intelligence model Psi-R2.5, which uses a two-layer architecture of a high-level planner (QwenVL3.5-4B) and a low-level controller (Wan2.2-IT2V-5B), proposes and implements a "strong Pair Data" standard, reverse-invokes the Psi-W0 world model to generate human-hand demonstration videos from real-robot execution trajectories, distills an end-to-end human-to-robot data conversion model, and pairs it with a dexterous-hand HIL-in-the-loop plus RL post-training framework that the article says raises success rate to 99% in only 1-2 working days after a few iterations, while measuring compositional generalization with simulation evaluation embedded in pre-training and a 50-task real-robot multi-task benchmark.

AI-generated editorial illustration: 灵初智能Psi‑R2.5正式发布!模型能力持续强化,破解人机数据对齐难题,定义高质量人类数据标准

Interpretation

It proposes and implements a "strong Pair Data" standard: on the basis of identical interaction objects, scene details are aligned frame by frame and temporal trajectories matched one to one, so the converted robot action sequences can be replayed on real hardware and fully reproduce the task. The article says the industry mostly relies on "weak Pair Data" that only satisfies task-semantic similarity, with scene details and temporal frame sequences that cannot be precisely matched; strong Pair Data aligns the human manipulation dynamics domain to the robot execution dynamics domain. Presented through two self-defined validation criteria: converted trajectories can be replayed and reproduced on real hardware, and post-training on converted data generalizes to new tasks of the same kind; the text says "the vast majority" of converted human demonstration data can directly drive the robot to complete the task, but gives no sample size or statistical basis.

It proposes a reverse-generation idea: instead of fitting robot data from human data, it reverse-invokes the Psi-W0 world model to batch-generate human-hand demonstration videos whose scene, timing, and actions highly match massive high-precision robot execution trajectories, and distills an end-to-end human-to-robot data conversion model from this paired data. Compared with the R2 version, which replaced traditional simulators with Psi-W0 and used reinforcement learning to optimize human-hand trajectories but was cumbersome and relied on repeated iteration between policy and world models, reverse generation is described as a more efficient path to mass-producing strong Pair Data. The article makes a qualitative comparison of reverse generation against prior approaches such as Real2Sim reconstruction transfer and image segmentation plus inpainting to replace the robotic hand, and says any human-hand video shot with an ordinary phone or camera can be converted in one click into robot visual frames and executable action sequences; no quantitative comparison is provided.

It pairs the model with a dexterous-hand HIL-in-the-loop plus RL post-training framework: on top of the base model, only a small amount of demonstration data is needed to fine-tune for complex business scenarios, and failure cases are collected in real time during deployment and automatically fed back for iteration. It treats post-training and on-site data feedback as a core part of commercialization rather than relying only on a purely pre-trained base model. The article says that after a few iterations the model can raise success rate to 99% and takes only 1-2 working days, and that the approach has been deployed at actual customer sites; it does not disclose task types, number of trials, or statistical method.

On evaluation, it embeds simulation evaluation deeply into the entire pre-training process and builds a proprietary real-robot multi-task benchmark containing 50 high-difficulty complex tasks, with task initial states randomly reset throughout evaluation. Randomly resetting initial states is used to avoid overfitting and to measure compositional generalization and adaptation to complex scenes. The article gives the scale figure of 50 tasks and the random-reset design, but reports no specific evaluation scores or comparisons against baselines.

Perspective

The article targets embodied-intelligence R&D and industrial deployment scenarios that need to convert human demonstration data into robot-executable data, especially sites with diverse SKUs and highly customized work cycles. Its reusable assets include: the criteria for strong Pair Data (identical interaction objects, frame-by-frame scene alignment, one-to-one temporal trajectory matching, converted trajectories replayable on real hardware), the reverse-generation plus distillation end-to-end conversion model, and the HIL plus RL post-training and failure-case feedback loop. For readers, this framework can be used to assess whether their own data pipeline stops at "weak Pair Data," and to judge whether evaluation should be embedded in pre-training and whether a proprietary real-robot multi-task benchmark is needed.

The article does not disclose the generation scale of strong Pair Data, the training data volume and structure of the conversion model, the task types and number of trials behind the 99% success rate, or specific scores on the 50-task real-robot benchmark or comparisons against baselines. The ICL section says that "without updating model parameters and without secondary training" the robot can zero-shot generalize to new complex tasks, but the technical details are explicitly deferred to a later release, with demonstration demos only on the project homepage. Readers will therefore still watch: how broadly the reverse-generated human-hand demonstration videos maintain frame-by-frame matching with real-robot trajectories, and whether the HIL plus RL success-rate improvement reproduces stably across different tasks and hardware.

Sources