ASK THE WORLD: Correction-guided active perception for VLAs
Abstract
Vision-language-action (VLA) policies can fail when the spatial relation needed for an action decision is not visible from the current observations. Inference- time correction offers a way to improve frozen policies, but a proposed correc- tion can itself be uncertain: the available evidence may be insufficient to deter- mine whether applying it would actually help. We introduce Ask the World, a correction-guided active perception framework that uses a tentative correction to determine what the robot should observe before intervening. Our key idea is to treat the correction as a decision-directed visual query: it identifies the spatial re- lation that must be resolved to accept or reject the intervention. When the current evidence is insufficient, Ask the World selects a viewpoint that exposes this rela- tion, acquires a new observation, and reassesses the correction before execution. We instantiate the framework with LingBot as the frozen base VLA and Astra as an inference-time visual reasoning assistant. LingBot provides the nominal action chunk, while Astra proposes bounded local corrections, assesses whether the native observations provide sufficient evidence, and acquires additional views only when needed. The newly acquired evidence is used to reassess the correction against the same nominal action before execution.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.