TouchAgent: Benchmarking and Advancing Real-World Tactile Intelligence
Abstract
Recent advances in multimodal large language models (MLLMs) have expanded multimodal intelligence. However, progress in the tactile domain remains constrained by several distinctive challenges, including heterogeneous sensor modalities, diverse domain knowledge, and a fragmented task landscape. To address this gap, we introduce TouchBench, a comprehensive benchmark for real-world tactile interaction with **10K** questions and approximately **383K** observation frames, covering **6** real-world tactile task categories ranging from perception to reasoning and decision-making, thereby enabling broader and more systematic evaluation than prior benchmarks. Using TouchBench, we evaluate 14 open-source, proprietary, and tactile-language models, revealing systematic deficiencies in perceptual grounding, domain knowledge, and reasoning, all of which are essential for real-world tactile intelligence. Beyond evaluation, we propose TouchAgent, an agent framework for tactile intelligence that strategically integrates knowledge retrieval, perception, and reasoning through domain-specific tactile models and tools. To further improve tool planning and evidence integration, we construct TouchAgent-Instruct with **40K** multi-round interaction trajectories and optimize TouchAgent through supervised fine-tuning and reinforcement learning. Extensive experiments show that TouchAgent consistently outperforms existing MLLMs and tactile-specific models across tactile benchmarks, highlighting the effectiveness of tool-augmented agents in advancing real-world tactile intelligence. Our code and benchmark will be released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.