acceptodds
Under review as a conference paper at ICLR 2027

CaRAMeL: Compressed and Retrieval-Augmented Spatio-Temporal Memory for Long-Horizon Robot Navigation

Abstract

Caption-based spatio-temporal memory lets a robot answer where and when something happened over hours of operation, but two costs keep it off the robot: the captioner's visual-token budget and a reasoning LLM that such systems assume to be a frontier API model. Consecutive frames from a moving camera repeat the same scene, consecutive captions repeat whenever the robot is stationary or revisits a place, and the quantitative answers the memory is queried for are anchored to the pose and timestamp stored with each entry rather than to its caption; a local reasoner declines to transcribe those fields on close to half of the quantitative questions, so its dependence on an API model is one of coverage rather than accuracy. CaRAMeL acts on each observation without training. Evidence-grounded decoding grounds quantitative fields in the retrieved metadata. Chimera token compression transplants the group-of-pictures structure of video codecs into the vision encoder, keeping one reference frame per group intact and carrying only the changed regions in a second frame, so that no token pools across observations. Caption compression merges consecutive captions only while the robot's odometry keeps it within a bounded radius, and every segment's pose and timestamp remain in the store. On NaVQA, an 8B local reasoner on two-thirds of the visual tokens and a 40% smaller store reaches the positional and temporal error of the same memory driven by a frontier API model (33.7 against 34.4 m), with the lowest hallucination rate among compression methods and a 1.2x faster captioner on a Jetson AGX Orin.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.