acceptodds
Under review as a conference paper at ICLR 2027

ReTAC: Residual Test-Time Agentic Computation for Video Spatial Reasoning

Abstract

Test-time agentic computation can improve multimodal large language models (MLLMs) through planning, tool use, iterative perception, and reflection, but indiscriminate computation may duplicate resolved work, introduce noisy evidence, and perturb correct predictions. We propose that such computation should be residual: allocate additional computation only to requirements unresolved by the base model and, after a localized failure, recompute only the affected states. We instantiate this formulation in ReTAC, a training-free multiagent framework for video spatial reasoning. ReTAC organizes supplementary computation around unresolved requirements. It plans ordered steps to acquire missing spatial evidence and combines the observations with compatible native evidence for reasoning. In addition, it uses stage-aligned verification to localize failures to planning, perception, or reasoning. Dependency-based recovery recomputes only affected states while reusing unaffected ones. Extensive experiments show that ReTAC achieves the highest overall score among the compared methods across three mainstream MLLMs and two widely adopted benchmarks. Further analyses show that gains from additional computation depend not only on how much is used, but also on where it is allocated and which valid intermediate states are preserved.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.