acceptodds
Under review as a conference paper at ICLR 2027

RESCENE: Reasoning over Scene Edits for Verifiable Visual-Language-Action Grounding

Abstract

Vision-language-action (VLA) models unify perception, reasoning, and control by mapping visual observations and language instructions to actions. When responding to language-defined situations that are not directly observed, direct language–action supervision alone cannot distinguish visual grounding from an instruction-to-action shortcut. We propose RESCENE, a framework that makes V-L-A alignment through language-specified interventions in visual latent space explicit and testable. RESCENE combines two complementary, shared-weight components: a latent scene editor and a language–action response model. Given observation history, target regions, and an editing condition, the editor masks affected evidence and reconstructs visual features using context, ego motion, and spatial hints. Factual reconstruction anchors these features to observations, while counterfactual construction produces feature-space alternatives for object addition, removal, relocation, and state changes without generating pixels. The response model uses the constructed features as a shared interface for language reasoning and action prediction, without direct access to the edit condition. Response gradients stop at this interface, blocking direct policy-loss updates through latent construction and separating construction from the responses it supports. Reconstruction, language, and action are jointly learned, while normal driving retains a direct observed-feature pathway. On Bench2Drive, RESCENE improves driving score by points and route completion by percentage points over factual replay, with a higher success rate. Implemented in autonomous driving, RESCENE provides an explicit, visually mediated, and testable formulation of V-L-A alignment for language-defined scene changes. Project page: https://resceneiclr2027.github.io/.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.