Skip to main content
Back to timeline
University of Twente Research InformationSource publication:

WiFi CSI cannot substitute for wearable IMU: 6-11% weighted accuracy on 27 activities versus 73-80% for IMU

Synopsis

Using the Multimodal Activity Sensing Dataset (MASD; 27 activities, 20 participants), the study tests whether ambient WiFi Channel State Information (CSI) can substitute for a body-worn inertial measurement unit (IMU) in human activity recognition; the authors find that the dataset's released machine-learning-ready files expose a degraded signed CSI representation rather than the documented amplitude, so they rebuild amplitude and antenna-ratio Doppler from the raw complex CSI, recover omitted participant identifiers, and re-evaluate subject-independently under a controlled multi-seed protocol, finding WiFi systematically weak: every backbone they try reaches only 6-11% weighted accuracy on the 27-class set against 73-80% for IMU, with usable accuracy on 0 of 27 activities; the corrected r

Source-provided article image: Where WiFi Sensing Fails:A Reproducibility Study of CSI-IMU Substitution for Human Activity Recognition

Interpretation

The study provides the first at-scale account of where and why WiFi CSI fails to substitute for IMU in activity recognition. Prior work often reports feasibility of WiFi sensing, whereas this study systematically characterises the distribution and causes of substitution failure under subject-independent, multi-seed controlled evaluation. Based on MASD with 27 activities and 20 participants, using multi-seed controlled evaluation and subject-independent splits; WiFi backbones reach 6-11% weighted accuracy versus 73-80% for IMU, with 0 of 27 activities usable.

The study reveals that the dataset's released machine-learning-ready files expose a degraded signed CSI representation rather than the documented amplitude, invalidating conclusions drawn from the files as shipped. This is a reproducibility finding about the released data artifact itself, not merely a re-evaluation of a model or algorithm. The authors rebuild proper amplitude and antenna-ratio Doppler from the raw complex CSI and recover omitted participant identifiers; the corrected representation yields a 14 percentage point gain on the easy subset.

Even with the corrected CSI representation, WiFi still cannot substitute for IMU, and fusing it with IMU does not beat IMU alone. The conclusion directly addresses the CSI-IMU substitutability hypothesis with negative evidence, rather than comparing performance of a single model. WiFi remains systematically weak after representation correction; fusion does not exceed the IMU-only modality; a well-regularised gated fusion model closes its gate and becomes invariant to WiFi.

The authors release a corrected, subject-labelled CSI pipeline for MASD to support subsequent reproduction and re-evaluation. The pipeline corrects both the representation and the participant identifiers, enabling subject-independent evaluation to be redone consistently on this dataset. The pipeline is public at https://github.com/Joost080/masd-wifi-imu-substitutability and underpins the paper's multi-seed, subject-independent evaluation.

Perspective

The result targets activity recognition research that uses the MASD dataset and seeks to substitute or supplement body-worn IMU with WiFi CSI, under subject-independent, multi-seed controlled evaluation settings; the authors' released corrected, subject-labelled CSI pipeline supports subsequent reproduction and re-evaluation, and can be used to test whether other representation-rebuilding or fusion strategies change the substitution conclusion.

Readers should still watch: whether the corrected CSI representation changes conclusions on other datasets or sensing configurations; whether the gated fusion model's gate-closing behaviour is stable across different regularisation strengths; the joint failure of the 27K-parameter feature model and the deep network on the same activities WiFi cannot recover suggests recoverability of those activities may be limited by the sensing modality itself rather than model capacity; and, as the paper is presented here in abstract form, specific activity categories, failure-mode breakdowns, and statistical details require consulting the original text and the public pipeline.

Sources