Trace-VSI: Benchmarking Progressive Spatial Reasoning Along Egocentric Traversals
Abstract
Understanding a home from egocentric video requires relating new views to previously seen places and accounting for movement between them. Answers to cumulative questions, such as how many distinct rooms have been entered, can therefore change as a traversal continues. End-of-video accuracy alone cannot reveal earlier mistakes or whether predictions track these changes. We introduce TRACE-VSI, a scan-grounded benchmark with 18 full-traversal question types about places, observed routes, and spatial configuration, and repeated questions about 5 cumulative quantities as observations accumulate. At each point, the model is queried independently using video from the start of the traversal up to that point. We score each answer for accuracy and compare predicted changes with changes in the correct answers using Spatial Change Accuracy (SCA). Across 30 systems, 22 exceed the answer-frequency baseline on place/time binding after the full traversal, but only 6 do so on observed-route questions. At earlier and later points, correct count changes often coexist with incorrect counts: Qwen3-VL-8B scores 55.3 SCA on Distinct Rooms but only 13.1% joint exact accuracy. Distance changes are also inaccurate: Qwen2.5-VL-72B has 23.6 m Path Length change MAE, compared with 9.0 m for a video-free prior. Larger frame budgets can improve change agreement without clear evidence of improved joint count correctness. These results reveal persistent count offsets and inaccurate distance increments, showing why accurate changes alone do not establish accurate spatial tracking.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.