acceptodds
Under review as a conference paper at ICLR 2027

Learning Spatially-aware Representations from Scene Graphs Using Optimal Transport

Abstract

Visual similarity is typically considered from the semantic content of images, while the structure and the spatial organization of objects are often discarded. Spatial information curated from objects (such as bounding box coordinates and relative position relations between objects) is usually absent from the neural model similarity assessment. In the meantime, scene graphs computed from images provide an intermediate representation for explicitly modeling the structure of objects in images, yet their spatial configurations remain underexplored, limiting their usability for spatially-aware computer vision tasks like retrieval based on object arrangement. To address this gap and in this effort to better account for spatial information in images, we propose and formalize the Spatial Scene Graph (SSG), which encodes absolute and relative positions between objects as node and edge attributes, respectively. Building on this structured representation, we propose SPOT (SPatial scene graph Optimal Transport), a novel criterion for spatial similarity between SSGs, and indirectly between images. SPOT allows both the ability to evaluate neural models based on the spatial similarity found in the images and the training of neural backbones to learn more spatially-aware representations. We experimentally demonstrate these claims across two spatially-aware downstream tasks: content-based image retrieval and visual question answering.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.