Skip to main content
Back to timeline
arXivSource publication:

WiSPER combines pose-supervised pretraining with residual flow refinement to cut WiFi-CSI multi-person 3D pose error to 63.72 mm MPJPE

Synopsis

WiSPER is a two-stage framework: PAMEL couples JEPA-style masked latent prediction with auxiliary pose-set supervision on the same visible CSI context to guide the encoder toward joint localization from partial observations, and ReFT refines pose candidates through a conditional residual flow guided by coarse coordinates and per-joint decoder features. On PiW3D it reaches 63.72 mm overall MPJPE, a 40.0% reduction relative to WiFi-JEPA, with 42.1% and 38.1% reductions for two- and three-person observations.

Source-provided article image: WiSPER: Pose-Supervised Predictive and Residual Flow Refinement For Multi-Person 3D Pose Estimation With WiFi CSI
Figure 1 ·

Figure 1 : Our WiSPER consists of 2 stages: Stage 1, PAMEL, combines masked latent prediction with auxiliary pose supervision on visible CSI context. Stage 2, ReFT, uses the pretrained encoder and a pose Transformer to generate coarse pose candidates, then refines them through conditional residual flow guided by their coordinates and joint features. Residuals from H H sampled trajectories are averaged, scaled by σ \sigma , and added to each coarse pose.

arXiv

Interpretation

PAMEL jointly optimizes pose-set supervision and masked latent prediction on the same visible CSI context, so the encoder supports both cross-link prediction and joint localization from partial observations. Prior WiFi-JEPA masked latent pretraining used no pose annotations; this work introduces pose supervision into predictive pretraining when paired annotations are available. Ablation shows pose-supervised configurations obtain lower errors than CSI-only JEPA: PAMEL with refinement reaches 63.72 mm overall MPJPE, versus 92.24 mm for CSI-only JEPA with refinement and 65.15 mm for pose-only pretraining with refinement.

ReFT conditions an endpoint-parameterized residual flow on a candidate's coarse coordinates and per-joint Transformer features to correct pose candidates. Remaining coordinate errors of a structured hierarchical pose Transformer are modeled explicitly as conditional residuals rather than left to direct regression output. Enabling the trained residual refiner reduces overall MPJPE by 15.6%, 14.6%, and 13.8% across the three pretraining configurations; the authors state this supports applying the refiner within the evaluated models but does not establish superiority over direct residual regression.

On PiW3D, WiSPER attains lower error than WiFi-JEPA for single-, two-, and three-person observations, with error rising as the number of people grows. The work extends single-person improvement to multi-person CSI pose estimation and reports multi-person comparisons under the same local protocol. Single-person 49.42 mm versus 81.20 mm, with PCK@0.2 rising from 74.26% to 89.07%; two-person 63.76 mm versus 110.10 mm; three-person 81.65 mm versus 131.93 mm.

Accuracy does not improve monotonically with more flow integration steps. Trajectory count and Euler-step count are varied at the same checkpoint to characterize the refiner's sampling behavior. One trajectory and one step gives 64.71 mm, the main configuration (8 trajectories, 5 steps) gives 63.72 mm, 32 trajectories gives 63.33 mm, and 8 trajectories with 10 steps gives 63.98 mm.

Perspective

The result targets indoor multi-person pose estimation from WiFi CSI where paired CSI and 14-joint pose annotations are available during training, while inference requires only CSI. It enables multi-person skeleton recovery in settings where visual observation is unavailable or undesirable, such as rehabilitation and exercise monitoring. Evaluation uses the PiW3D split of 89,946 training frames and 7,824 test frames, covering seven volunteers, eight actions, and three indoor environments; the authors note that training and test frames originate from different temporal portions of shared recordings, so this evaluation does not establish generalization to unseen participants or environments. Future work is directed at generalization evaluation across environments and unseen participants, and at reducing paired supervision through independently pretrained CSI representations and human-pose priors.

Results come from individual training runs, and variation across training seeds is not evaluated; the reported matched-pose errors do not assess person-count accuracy; evaluation uses temporal portions of shared recordings, leaving generalization to unseen participants and environments, as well as inference latency, for future investigation. In the ablation, PAMEL and pose-only pretraining also differ in loss weighting, and the authors state this comparison does not isolate the contribution of masked latent prediction; PAMEL performs better on two- and three-person observations while pose-only pretraining is slightly better on single-person observations (49.36 versus 49.42 mm). In addition, GraphPose-Fi's 622.38 mm is the authors' local adaptation to nine CSI links and a 14-joint graph, which the authors state characterizes that adaptation rather than its published performance.

Sources