EgoFact: Estimating Dense Full-Hand Tactile Fields from Egocentric Human Videos
Abstract
Understanding human dexterity from egocentric videos requires the ability to recover not only the kinematics, but also the physical tactile signals underlying the hand-object interface. However, estimating such tactile signals from egocentric videos remains a challenging task due to severe occlusions and heterogeneous tactile representations across sensing devices and datasets. To bridge this gap, we introduce EgoFact, a unified framework for estimating dense, full-hand tactile fields of both hands directly from egocentric videos. Specifically, 1) we map heterogeneous tactile measurements onto a canonical 3D hand surface with spatial decay and propagation around sensor centers, yielding a unified full-hand tactile representation for cross-device training and evaluation; 2) we develop a tracker-free vision-to-tactile model with force-guided learning objective to effectively infer tactile fields directly from RGB inputs; and 3) we introduce EgoFact-Data, a cross-device egocentric-tactile dataset with 100 hours of recordings, 800 clips, three glove types, six scenes, and 200 manipulation tasks, together with scene- and task-held-out protocols for evaluating generalization. Extensive experiments across five egocentric–tactile benchmarks demonstrate consistent improvements over existing baselines, achieving up to around 25% higher contact accuracy, 30% higher field IoU, and 40% lower force error. Notably, EgoFact further improves 3D hand–object interaction reconstruction and supports dexterous skill imitation, presenting the broader value of vision-to-tactile estimation for physically grounded egocentric intelligence.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.