acceptodds
Under review as a conference paper at ICLR 2027

Beyond Reactive Visual Assistance: A Real-Time Scene-Interaction Agent for Blind and Low-Vision People

Abstract

Visual assistance for blind and low-vision (BLV) people remains largely query-driven: users must ask about cues they may not know exist. Real-time BLV assistance from continuous first-person video instead requires an agent that decides when evidence warrants speech, avoids redundant notifications, and delivers scene-appropriate guidance while that evidence remains actionable. We introduce a BLV-informed proactive scene-interaction dataset by densely annotating 439 multi-scene first-person video segments from BLV users' daily lives, yielding 7,482 timestamp-aligned PUSH/NO-PUSH decisions along continuous trajectories and assistance messages for 1,794 PUSH decisions. We present BLV-SceneAgent, a real-time observe–decide–act–commit agent built on a streaming VLM. Its Temporal Contextual Intervention (TCI) policy integrates the native response margin with a causal recurrent belief, signed state change, and a committed-notification ledger, and is refined under the agent's predicted action history. NO-PUSH terminates at the action boundary, whereas PUSH reuses the multimodal cache for content generation. On a segment-disjoint test set, TCI improves steady-state macro-F1 by 6.3 points over calibrated SFT, with a P95 action-boundary latency of only 478.8 ms. A live study across four settings shows that the agent surpasses baselines in usefulness, trust, and reduced annoyance. We will publicly release our dataset and model to support further research on proactive visual assistance for blind and low-vision people.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.