acceptodds
Under review as a conference paper at ICLR 2027

PEARL: Personalized Streaming Video Understanding with Runtime-Defined Concepts

Abstract

Human cognition of new concepts is inherently a streaming process: we continuously recognize new objects or identities and update our memories over time. However, current methods are largely limited to static images or offline videos. This disconnects continuous visual input from instant real-world feedback, limiting their ability to provide the real-time, interactive personalized responses essential for future AI assistants. To bridge this gap, we first propose and formally define the novel task of Personalized Streaming Video Understanding (PSVU). To facilitate research in this new direction, we introduce PEARL-Bench, the first comprehensive benchmark designed specifically to evaluate this challenging setting. It evaluates a model's ability to respond to runtime-defined personalized concepts across Concept-Definition QA, Real-Time QA, and Past-Time QA, each further organized into fine-grained reasoning categories. Concept-Definition QA introduces user-defined concepts into the stream, Real-Time QA queries their current states and includes a new Action category for clip-grounded personalized motion concepts in addition to frame-grounded entity concepts, and Past-Time QA tests long-range retrieval of historical personalized evidence. PEARL-Bench comprises 132 unique videos and 2,173 fine-grained annotations with precise timestamps. To tackle this challenging new setting, we further propose PEARL, a plug-and-play strategy that serves as a strong baseline. Evaluations across 8 offline and online models demonstrate that it brings consistent PSVU improvements when applied to 3 distinct architectures, proving to be highly effective and robust. We hope this work advances vision-language model (VLM) personalization and inspires further research into streaming personalized AI assistants.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.