acceptodds
Under review as a conference paper at ICLR 2027

Cortically Inspired Visuo-Tactile Alignment for Cross-Embodiment Tactile Inference

Abstract

Observing touch on another person's body can recruit somatosensory sensations in observers, providing a striking example of how the brain translates visual events into bodily states. Neuroscientific evidence suggests that this visuo-tactile resonance relies on structured alignment between visual and somatosensory neural representations, yet the computational instantiation and functional role of such alignment remain unclear. Here we introduce MTNet, a cortically inspired dual-stream framework that learns visual-to-tactile prediction under semantic, distributional, and geometric alignment constraints. Given monocular RGB observations, MTNet predicts force fields across 1,140 taxels on a robotic hand at 1-mm spatial resolution. Compared with reconstruction-driven models, the alignment constraints substantially improve spatiotemporal prediction and reshape visual representations to match the tactile manifold geometry, thereby simplifying the mapping from vision to touch. By aligning paired visual observations of human and robotic hands within the learned visuo-tactile space, the physical robot generates corresponding tactile states from observed human touch, achieving cross-domain transfer without human tactile labels. These results offer a computational account of how the cross-modal alignment supports visuo-tactile resonance across embodiments, translating neuroscientific findings into a design principle for tactile inference in embodied artificial systems. Code and videos are available in the supplementary materials and will be publicly released upon acceptance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.