acceptodds
Under review as a conference paper at ICLR 2027

Aimed Elsewhere: Understanding and Recovering Spatial Localization in MoE VLMs

Abstract

Supervised fine-tuning can adapt general-purpose vision-language models to embodied tasks by improving spatial reasoning and adherence to task requirements. In the sparse mixture-of-experts backbones we study, however, these gains coexist with substantial degradation in spatial localization, while general visual-language performance remains largely preserved. Complementary interventions reveal that spatial competence remains recoverable. Rephrasing the request recovers much of the localization loss without changing the weights. Subtracting an update along the fine-tuning trajectory partially recovers performance without additional training data, while adding the same update to a healthy checkpoint of an independent run reproduces the decline. Together, these findings support a re-aiming, not erasure account: adaptation alters how spatial competence is expressed. These interventions establish recoverability but do not ensure recovery across task requirements. Successful localization requires both accurate spatial targets and responses in the structure required by each task. Our controlled comparisons show that supervision covering only a subset of these requirements leaves recovery incomplete. We therefore introduce Joint Spatial Re-Aiming (JSR), which guides corrective updates through joint supervision of the affected tasks and their required response structures. With 384 distinct examples and 32 updates, JSR substantially restores spatial localization while largely preserving the gains of embodied adaptation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.