acceptodds
Under review as a conference paper at ICLR 2027

Towards Multimodal Lifelong Understanding: A Multi-Scale Benchmark

Abstract

Multimodal systems can access increasingly long video histories, but their ability to connect observations of the same subject across days and months remains insufficiently evaluated. We formalize the Lifelong Horizon, distinguishing recorded duration from the physical temporal span of a persistent subject's history. To evaluate this setting, we introduce MM-Lifelong, a benchmark with 181.1 hours of video and 1289 questions across Day, Week, and Month scales. Its Month subset contains 105.6 hours of footage spanning 51 days. Every question is paired with manually annotated clue intervals. We evaluate answer quality and temporal grounding separately, using question-specific rubrics to accommodate valid alternative answers and partial correctness. Experiments with multimodal large language models (MLLMs) and agents reveal uneven gains from larger input budgets, with the best MLLM scoring 26.53 out of 100. Among agents, answer quality and grounding rankings diverge. Error analysis identifies lost answer-relevant details, confusion between event instances, and unresolved conflicting observations. These findings motivate memory systems that preserve event-specific evidence and verification that checks its support for the final answer.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.