Skip to main content
Back to timeline
arXivSource publication:

TacEx Uses Tactile Curiosity to Drive Robot Exploration, Learning Grasping Without Task Rewards and Improving VLA Downstream Performance

Related research and updates

Synopsis

The work introduces TacEx, a framework that decomposes model uncertainty across sensory modalities and directs curiosity toward the tactile channel, enabling a robot to learn manipulation and grasping during exploration without task rewards or expert demonstrations; the resulting interaction-dense dataset supports offline learning of downstream pick-and-place policies, and tactile-driven exploration is further used to post-train vision-language-action (VLA) models, substantially improving downstream performance while remaining highly sample-efficient.

Source-provided article image: Tactile Curiosity Drives Robot Interaction
Figure 1 ·

Figure 1: Overview of TacEx . The state s s concatenates the fused latent z z , a visual embedding v v , and a tactile embedding ι \iota , where v v and ι \iota are produced by separate CNN encoders from RGB and tactile force maps. At each step, the policy π ⁡ ( a | s ) \pi(a|s) acts in the environment, and the resulting transitions train an ensemble environment model that predicts the next state s ′ s^{\prime} . The model’s epistemic uncertainty decomposes per modality into σ n z \sigma_{n}^{z} , σ n v \sigma_{n}^{v} , and σ n ι \sigma_{n}^{\iota} ; unlike standard MaxInfoRL [ 31 ] , which rewards the agent for total predictive uncertainty, TacEx explicitly weights the tactile component σ n ι ​ ( s , a ) \sigma_{n}^{\iota}(s,a) in the exploration bonus. This anchors curiosity to the sense of touch, driving the agent toward contact-rich interactions and away from uncertainty in functionally irrelevant regions and transitions.

arXiv

Interpretation

It proposes TacEx, which treats tactile feedback as a natural exploration signal by decomposing model uncertainty across sensory modalities and directing epistemic-uncertainty-driven curiosity toward the tactile channel. Relative to existing intrinsic motivation methods based on model disagreement or epistemic uncertainty, which can reward uncertainty in functionally irrelevant transitions such as erratic motions in free space, TacEx uses modality decomposition to concentrate exploration on contact-related transitions. At the abstract level the text presents the framework design and its motivation, indicating the intended improvement over isotropic noise and prior intrinsic motivation methods; implementation details, comparison baselines, and quantitative metrics are not given in the provided text.

Under tactile-driven curiosity, the robot discovers complex contact dynamics and learns to manipulate and grasp objects without task rewards or expert demonstrations during exploration. It extends reward-free, demonstration-free exploration from free-space motion toward contact-rich manipulation and grasping, indicating that the tactile channel can serve as a source of exploration signal. This is an experimental result stated by the authors in the abstract; no task counts, success rates, or sample sizes are provided.

The interaction-dense dataset collected through tactile-driven curiosity supports offline learning of downstream pick-and-place policies without additional environment interaction. It turns exploration-phase data into a reusable offline learning resource, so downstream policy learning no longer requires new online interaction. The abstract states that the dataset supports offline learning of downstream policies; dataset size, offline algorithms, and performance comparisons are not given.

Tactile-driven exploration is used to post-train vision-language-action (VLA) models: although the VLAs are initially pre-trained without tactile feedback, post-training with TacEx substantially improves downstream performance while remaining highly sample-efficient. It shows tactile exploration can serve as a post-training procedure for VLAs, adding contact information even when the tactile modality was absent during pre-training. The abstract gives a directional conclusion (substantial improvement, high sample efficiency) without specific improvement magnitudes, benchmark tasks, or sample sizes.

Perspective

The work targets robot manipulation and grasping, especially settings where contact is dense and task rewards or expert demonstrations are hard to obtain; its value lies in using the tactile channel to guide exploration, using the resulting interaction-dense data for offline learning of downstream pick-and-place policies, and post-training VLA models that initially lack tactile feedback. It is relevant to research and engineering teams that need low-cost contact data collection or want to add tactile capability to an existing VLA.

The provided text is abstract-level and lacks figures and experimental detail, so it cannot be determined under which task distributions or tactile hardware conditions tactile curiosity holds, nor how large its gains are relative to existing intrinsic motivation methods; the magnitude of VLA post-training improvement, how sample efficiency is measured, the size of the offline dataset, and the generalization range of downstream policies all need confirmation in the full text. In addition, how modality-decomposed uncertainty estimation behaves under noisy tactile signals or sparse contact is an open question worth watching.

Sources