acceptodds
Under review as a conference paper at ICLR 2027

SURE-Bench: Sensor Understanding and Reasoning with Evidence Benchmark for Human Activity

Abstract

Recent human activity benchmarks increasingly incorporate heterogeneous sensors for multimodal understanding and reasoning. However, existing evaluations primarily verify the final task output or answer. Recent works indicate that correct answers do not necessarily mean that a model relies on the intended evidence, as VQA models can exploit shortcuts while still producing the expected answer. To address this issue, we introduce the Sensor Understanding and Reasoning with Evidence Benchmark (SURE-Bench), a multimodal and multi-view benchmark that evaluates both answer correctness and sensor-evidence correctness. SURE-Bench contains synchronized thermal, event camera, LiDAR, and radar observations across 61 indoor activity scenarios, with calibrated motion capture providing physical ground truth. We construct 687 sensor-exclusive atomic questions, each answerable from exactly one target modality, and compose them into 220 compound questions requiring specific sensor pairs. The required sensor evidence is therefore known by construction, allowing us to evaluate whether a model identifies the modalities required to support its answer. Experiments with recent VLMs show that answer correctness and sensor-evidence correctness do not consistently agree. Specifically, models may answer correctly without identifying the required sensor pair, or conversely, identify the required sensor pair without answering the question correctly. To the best of our knowledge, SURE-Bench is the first benchmark for indoor human activity to quantitatively evaluate answer correctness and sensor-evidence correctness separately, supporting research toward multimodal models whose answers are grounded in the required sensor evidence.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.