Anchoring the Observer: Observer-Centeric Positional Encoding for Situated 3D Reasoning
Abstract
We present an observer-centric framework for situated question answering (SQA), coined OC-SQA, that reasons about spatial relations from an observer’s position and orientation. Existing SQA approaches are largely insensitive to situation descriptions, and observer-specific situation information is insufficiently exploited. To address this limitation, we explicitly incorporate observer poses for the positional encoding of visual tokens via a novel observer-centric rotary position embedding (OC-RoPE). OC-RoPE projects 3D locations of visual tokens onto a shared bird’s-eye-view (BEV) plane and then transforms projected positions into the observer’s reference frame for observer-centric positional encoding. We further introduce a learnable observer anchor at the head of the visual tokens of each frame to explicitly indicate the observer’s origin. Together, these two designs provide direct geometric cues for situated reasoning. Extensive experiments demonstrate that our method achieves state-of-the-art performance on SQA3D and Real3DQA, while also demonstrating strong performance on MSQA. Further in-depth analyses show that our design increases the reliance on visual tokens relevent to the situated questions when predicting answers.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.