Elephant: Recognize and Reconstruct the World Blindly
Abstract
Understanding an object's geometry and semantic identity through physical contact is an important capability for robots operating with limited visual information. Yet each tactile observation reveals only a small region of the surface, leaving open how much object-level understanding can be achieved from just a few contacts. We introduce Elephant, a framework for joint object reconstruction and recognition from sparse tactile observations, without visual input at inference time. Elephant uses sensor poses to express local observations in a common spatial frame and aggregates evidence across grasps to infer global geometry and semantic category. We also introduce a real-world tactile dataset collected using UMI-based grippers equipped with optical Daimon and magnetic eFlesh sensors, pairing tactile measurements with sensor poses and object geometry. The curated Daimon subset comprises 52 objects from 3 families, while the eFlesh study covers 6 semantic categories composed of 18 instances. These real-world observations are complemented by simulated tactile data from 279 object instances spanning 43 semantic categories. On real Daimon data, Elephant achieves a CD-L1 of 6.19 mm with ten grasps. In a separate six-category eFlesh study with six held-out object instances, it achieves 94.4% recognition accuracy over 18 sampled inputs at 20 tactile episodes, alongside a reconstruction F-score of 56.2%, compared with 51.8% for the strongest evaluated reconstruction baseline. These results highlight the potential of spatially localized tactile observations for recovering global object shape and semantic identity across optical and magnetic sensing technologies.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.