UMMI: Symbolic Time Binding for Bounded Multimodal Memory with a Frozen Language Model
Abstract
A bounded memory can give a frozen language model access to multimodal sensor recordings: a small network compresses the streams into a fixed number of embeddings that are spliced into the model's input. We find that such memories support questions about what happened far better than questions about when. The limitation is not the reader alone: given the same facts as timestamped text, the same frozen Qwen3-Omni model places events in the correct interval with 77-78% accuracy without any training. We introduce symbolic time binding in UMMI, a learned memory interface. Memory embeddings are grouped by recording interval, and each group is prefixed, in ordinary text, with its interval index and step range. The backbone is not updated and the continuous memory is not enlarged; the prompt grows by sixteen short phrases. On a synthetic dense-event benchmark, interval localization rises from 60.7% to 98.0% over three training seeds, whereas learned interval embeddings added to the same tokens help much less. On question subsets from held-out OPPORTUNITY subjects, overall accuracy rises from 41.6% to 53.8% for the same shared-memory configuration, with gains that vary across subjects. A flat tokenizer given the same labels does not benefit. Time binding also fails to help in the tested Aria Everyday Activities configuration, where several transcript words share each interval. Explicit textual time labels are thus a simple and effective interface for reading time from a bounded memory, while preserving several items within one interval remains open.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.