acceptodds
Under review as a conference paper at ICLR 2027

UMMI: Symbolic Time Binding for Bounded Multimodal Memory with a Frozen Language Model

Abstract

A bounded memory can give a frozen language model access to multimodal sensor recordings: a small network compresses the streams into a fixed number of embeddings that are spliced into the model's input. We find that such memories support questions about what happened far better than questions about when. The limitation is not the reader alone: given the same facts as timestamped text, the same frozen Qwen3-Omni model places events in the correct interval with 77-78% accuracy without any training. We introduce symbolic time binding in UMMI, a learned memory interface. Memory embeddings are grouped by recording interval, and each group is prefixed, in ordinary text, with its interval index and step range. The backbone is not updated and the continuous memory is not enlarged; the prompt grows by sixteen short phrases. On a synthetic dense-event benchmark, interval localization rises from 60.7% to 98.0% over three training seeds, whereas learned interval embeddings added to the same tokens help much less. On question subsets from held-out OPPORTUNITY subjects, overall accuracy rises from 41.6% to 53.8% for the same shared-memory configuration, with gains that vary across subjects. A flat tokenizer given the same labels does not benefit. Time binding also fails to help in the tested Aria Everyday Activities configuration, where several transcript words share each interval. Explicit textual time labels are thus a simple and effective interface for reading time from a bounded memory, while preserving several items within one interval remains open.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.