Where Is Here? Spatial Frame-of-Reference Mismatch in Egocentric Text-to-Point-Cloud Localization
Abstract
Textual queries specify target locations through surrounding scene cues when exact coordinates are unavailable. Text-to-point-cloud (T2P) localization estimates the target position in a 3D point-cloud map from such queries. Existing T2P query protocols typically use map-frame directional relations, which can be interpreted as map-space constraints once the referenced objects are grounded. However, localization requests can instead describe local observations relative to an implicit observer-centered reference frame. An observer may know that a tree is on the left while the corresponding map-frame direction remains unknown. Even with correct object grounding, the same egocentric relation can imply different map-space constraints under different reference directions. We formalize this problem as Spatial Frame-of-Reference Mismatch. To study this setting, we introduce EgoT2P, a controlled benchmark with paired map-frame and egocentric queries over matched local maps and target positions. To test whether explicit reference-state modeling helps in this setting, we propose FoRLoc, which predicts a target-centered reference state consisting of a facing anchor and a coarse target-to-anchor bearing, together with auxiliary object correspondences, before estimating the target position. In the paired EgoT2P evaluation, all five evaluated T2P methods yield lower recall with egocentric queries. FoRLoc achieves a 20.6% relative improvement in 5 m localization recall over the strongest compared method.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.