acceptodds
Under review as a conference paper at ICLR 2027

SOMo: Unified Spatial and Observation-Grounded Scene Memory

Abstract

Questions about a robot's past observations often identify an object through its relationships to other objects and the context of a recorded encounter. An agent answering such questions through separate searches must assemble compatible objects and observation evidence from intermediate results; an early selection can exclude the intended answer and misdirect subsequent tool calls. We introduce SOMo, a scene memory that lets spatial relationships and observation context guide object selection during retrieval. Through a shared interface, the agent specifies object descriptions, required relationships, and spatial and temporal restrictions. SOMo jointly selects compatible objects and returns their mapped locations with the requested observation evidence, supporting proximity, co-observation, and robot-relative bearing queries. A sparse spatial hierarchy connects object groups to recorded robot visits and summarizes their spatial extent and observation intervals, allowing constraints to eliminate incompatible groups before individual verification. Adaptive planning selects the grouping resolution and constraint order using execution-cost estimates and counts of the object pairs remaining after grouped filtering, before expanding those pairs. Joint semantic ranking then selects among the verified combinations. With Qwen3.5-9B for both methods, SOMo reduces positional and temporal error relative to DAAAM by 45.3% and 37.6% on NaVQA, and by 68.09% and 64.32% on CODaAG's category averages. On SG3D, SOMo with Qwen3.5-9B achieves 31.82% sub-task grounding accuracy, compared with 22.16% for DAAAM with GPT-5-mini. On CODaAG, average tool use falls from 5.62 to 1.46 calls and answering time from 31.26 to 6.29 seconds. Separate candidate-generation benchmarks show 7–25 speedups over blocked NumPy scans on CODa and 16–193 on KITTI-360 with identical candidate results.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.