ArcX: A Unified Relation Encoder for Vector Spatial Representation Learning
Abstract
Vector spatial representation learning (VSRL) has primarily focused on representing individual spatial entities, such as their locations, geometries, and attributes, while paying limited attention to pairwise spatial entity relations. Current methods face two limitations in spatial relation representation. First, they often rely on centroid distance, which can discard geometric details. Second, relations are typically inferred only after the two entities are encoded independently, resulting in limited spatial interaction. We introduce pair-level VSRL as a complementary representation perspective and propose ArcX, a unified VSRL encoder that generates dedicated spatial relation embeddings for entity pairs. ArcX introduces an arc-wise cross-entity modulation mechanism, allowing each entity to respond to the other before embedding readout. We evaluate ArcX across five datasets spanning pair-, entity-, and scene-level VSRL tasks. The results show that ArcX can capture multiple dimensions of spatial relations, outperforming the second-best method by an average of 10.8% on pairwise tasks. ArcX can also improve entity-level neighborhood aggregation by an average of 4.4% over the second-best method and scene-level information propagation by an average of 16.5%, demonstrating consistent utility across diverse task settings. ArcX introduces a relation-centric perspective to VSRL, complementing existing entity- and scene-level perspectives. The code and data are available at https://anonymous.4open.science/r/ArcX-6BB3.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.