Skip to main content
Back to timeline
arXivSource publication:

Tactile-JEPA pretrains on the taxel connectivity graph and cuts force-estimation error by 6.3% and in-hand pose error by 20.8%

Synopsis

The work introduces Tactile-JEPA, a self-supervised pretraining method for distributed tactile sensors (e-skins) that samples masks at local and global scales over the taxel connectivity graph and predicts the embeddings of masked taxels; across magnetic and piezoresistive sensors and three datasets (Sparsh-skin, Tactile socks, DECO-50) it reduces force-estimation RMSE by 6.3% and in-hand pose RMSE by 20.8% over the strongest prior baseline, with consistent gains in action classification, object classification, and tactile-conditioned policy learning.

AI-generated editorial illustration: Tactile-JEPA: Topology-Aware Self-Supervised Representation Learning for Distributed Tactile Sensors

Interpretation

Tactile-JEPA moves self-supervised masked prediction from the visual pixel grid onto the taxel connectivity graph, using local masks (connected subgraphs grown by Dijkstra expansion) to capture localized contact detail and global masks (uniformly sampled across the graph) to reflect the overall contact state of the hand, mixed in equal proportion. Prior distributed-tactile representation methods either reshape the signal into an image and reuse visual SSL (e.g., T-DEX with BYOL), or treat taxel location only as a per-taxel feature (Sparsh-skin, STAT), or require contact-surface geometry available only in simulation (HyperTaxel); this work is the first to put sensor topology directly into self-supervised mask sampling without needing coordinates, labels, or privileged simulation information. The ablation shows that with the same compact connected masks and the same budget, sampling on the taxel graph (4L) beats graph-agnostic I-JEPA block masking (4I-JEPA) on both force estimation and policy learning; the default 2L+2G mix achieves the best force RMSE and in-hand pose accuracy and ranks second on DECO-50 behind pure global masking (4G).

Under a protocol where the pretrained encoder is frozen and only a lightweight task head is trained, Tactile-JEPA outperforms BYOL, MAE, DINO, and end-to-end baselines on multiple downstream tasks. Against the strongest baseline on each metric, force-estimation RMSE drops by 6.3% and in-hand pose RMSE by 20.8%; on Tactile socks, action-classification accuracy improves by 1.28% and full-body pose RMSE drops by 2.6%; on DECO-50 visuo-tactile policy learning it reaches second-best (MAE is lowest) without any per-dataset tuning. Three datasets span magnetic (Xela uSkin) and piezoresistive (knitted fabric, Inspire FTP) sensors, unimanual and bimanual setups, and human and robotic hands; each SSL method is pretrained with three seeds and three to four heads per encoder, giving 9 or 12 runs per task, with means, sample standard deviations, and a one-sided Welch t-test between the top two methods.

The method trains stably on heterogeneous tactile signals, whereas several baselines collapse on some datasets. BYOL collapses on Tactile socks and DECO-50, MAE collapses on Tactile socks, and DINO collapses on DECO-50 in its standard configuration, becoming viable only after disabling the iBOT loss; Tactile-JEPA pretrains stably on all datasets under the same masking configuration. The corresponding cells in Tables II and III are marked collapsed; the authors also report that Tactile-JEPA pretrains 65-75% faster than DINO and give per-dataset GPU hours (e.g., 6.1 h for Sparsh-skin and 5.5 h for DECO-50 on four NVIDIA A100 80 GB GPUs).

Mask scale steers what the representation favors: local masks help fine contact tasks, global masks help tasks tied to whole-hand contact state, and mixing scales is the robust choice for general-purpose pretraining. On DECO-50, pure global masking (4G) gives the lowest policy error, consistent with action prediction depending on the whole-hand contact state; the default 2L+2G is best for force estimation and in-hand pose accuracy. Replacing the default global context mask with a local one improves force estimation but substantially degrades pose and policy performance. The conclusion comes from the masking-strategy ablation in Table IV, comparing 4I-JEPA, 4L, 4G, 2L+2G, and local-context variants (4L†, 4G†, 2L+2G†), with the local-context variants excluded from the best and second-best markings.

Perspective

The result targets distributed tactile sensors covering a robot hand or the human body (magnetic e-skins, piezoresistive textiles), and pretraining needs only the sensor's own signal plus its known layout, not taxel coordinates, task labels, or privileged simulation information; the encoder stays frozen downstream with only a lightweight task head trained, which suits robot-learning settings with limited labels and a need to adapt quickly to new tasks. The pretraining objective is defined within short time windows and does not model temporal structure across windows, leaving temporal reasoning to the downstream decoder, so it fits deployments that invoke the encoder at control-loop rate and compose successive embeddings. For hardware with many taxels such as DECO-50, the authors group adjacent taxels to bring sequence length to a tractable size, indicating the route scales to high-coverage skin.

The pretraining objective is defined only within short windows, with cross-window temporal reasoning delegated to the downstream decoder, so how long-horizon contact dynamics are represented remains an open question. Evaluation centers on three public datasets and their tasks, leaving other transduction principles, other embodiments, and long-term stability in real deployment to be observed. On DECO-50, MAE achieves the lowest policy error, indicating that different self-supervised objectives have different strengths across tasks and that which objective combination is most general deserves further comparison. In addition, this reading is of the full paper text, where table values appear as text; some layout details of individual cells (such as significance markings) are hard to reconstruct exactly from plain text, so precise numbers are best checked against the original tables.

Sources