acceptodds
Under review as a conference paper at ICLR 2027

AfterImage: Object-Grounded Attention for Vision-Language-Action Policies

Abstract

Vision-language-action (VLA) policies must retrieve the right visual evidence for control even as a scene's appearance changes. Most VLA training leaves this selection to action supervision, while recent approaches guide decoder attention with task-relevant labels. Where such guidance should intervene remains a fundamental question. We identify a cross-modal positional asymmetry in rotary position embeddings: broad positional interventions severely impair control, whereas changes confined to action-to-image attention can preserve task performance. This finding motivates AfterImage, powered by a Conditional Object-Region Attention (CORA) mechanism. AfterImage treats the action stage following image encoding as a place to relearn where to look. It separates the selection of image regions from the total attention allocated to images, so object guidance redistributes an existing visual budget. We turn this principle into an adaptation system that learns from region priors during training. Deployment requires no external perception model and incurs negligible parameter/computational overhead. Attention measurements and closed-loop transfer evaluations connect this intervention pathway to more task-directed visual retrieval and improved manipulation robustness.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.