Spatial Relations Are Not Geometry Alone
Abstract
3D spatial reasoning requires effective representations of inter-object spatial relations. Existing methods typically construct spatial relations from predefined geometric schemes, while incorporating query semantics and reference frames at stages outside relation formation. This separation leads to two fundamental limitations: geometric measurements are fixed before the model knows what spatial evidence the query requires, while directional semantics are not defined directly under the relevant reference frame. Consequently, query-critical geometric cues can be irreversibly omitted, while the mapping from reference frames to directional semantics is poorly constrained and must be learned implicitly. Downstream reasoning is forced to operate on incomplete and ambiguous spatial representations, directly degrading reasoning performance. We therefore propose FROG, a unified relation-centric framework for 3D spatial reasoning, built on the principle that spatial relations should be formed jointly from reFerence fRames, geOmetry, and lanGuage. An order-aware geometry controller captures the relational semantics of the query and drives language-adaptive spatial measurement, adapting measurement directions and relative weights across spatial scales according to the query. In parallel, a relation-centric reference-frame mechanism uses observer orientation to define frame-dependent directional semantics directly during object-object relation formation, without modifying the underlying object or scene representations. By jointly integrating geometry, language, and reference frames into relation formation, FROG effectively mitigates the omission of query-critical geometric cues and the ambiguity of directional semantics, providing downstream reasoning with more complete and context-consistent spatial relation representations. Extensive experiments across multiple 3D LLM backbones, object settings, and spatial reasoning benchmarks demonstrate performance gains and validate the complementary roles of language-adaptive relation modeling and lightweight reference-frame integration within a unified relation-formation framework.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.