SEER: Spatial-Embodied Evidence Reasoning for Long-Horizon Task Completion
Abstract
Embodied agents face a fundamental challenge in long-horizon tasks: maintaining coherent reasoning across extended sequences of observations and actions while grounding decisions in spatial evidence. Existing VisionLanguage-Action (VLA) models either process each timestep independently, losing temporal context, or rely on implicit memory that degrades over long horizons. We introduce SEER (Spatial-Embodied Evidence Reasoning), a framework that explicitly constructs and maintains structured evidence chains grounded in spatial coordinates. SEER comprises three components: (i) an Evidence Chain module that extracts verifiable spatial evidence from each observation; (ii) a Spatial Memory Bank (SMB) that persists evidence as key-value pairs indexed by 3D coordinates with efficient radius-based retrieval; and (iii) a Retrospective Verification mechanism that cross-checks current observations against historical evidence to detect and correct reasoning errors. Experiments on ALFRED (Shridhar et al., 2020), Habitat 2.0 (Szot et al., 2021), and BEHAVIOR-1K (Li et al., 2024) demonstrate that SEER achieves 48.7%, 42.3%, and 36.9% task success rates, outperforming OpenVLA (Kim et al., 2024) by 13.1, 13.0, and 12.2 points respectively. On tasks with 8+ sub-steps, SEER maintains 33.1% success compared to 15.2% for OpenVLA. Ablation studies confirm that all three components contribute meaningfully, with Retrospective Verification providing 8.6 points of improvement on error-prone long-horizon tasks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.