acceptodds
Under review as a conference paper at ICLR 2027

BLV-Gate: Benchmarking Proactive Communication in Blind and Low-Vision Assistance

Abstract

Continuous visual assistance for blind and low-vision (BLV) users requires an assistant to decide whether to communicate, when to respond, and what information is useful, while avoiding unnecessary output. Existing evaluations focus on isolated questions or predefined event recall, providing limited evidence about this sequence-level communication problem. We introduce \system, a benchmark built from 499 verified egocentric video sessions across five everyday scene categories, containing 5,863 temporally localized communication events with grounded visual facts and reference responses. \system evaluates complete causal response sequences across intervention regulation, timing and latency, grounding and response quality, and session-level information burden. We also introduce SenseBridge Agent, a companion baseline with causal visual reasoning and memory. Experiments with ten representative baselines reveal clear trade-offs among event coverage, communication burden, response quality, and latency: periodic systems achieve higher recall at the cost of frequent output, conservative streaming systems miss necessary events, proactive assistants show stronger potential, and stronger backbones improve grounding and response quality without resolving timely delivery. All experimental data and code are available at: https://anonymous.4open.science/r/BLV-Gate-7A03.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.