TokenGazer: Learning Where to Look in Token Space for Visuomotor Control
Abstract
Using dense tokens from pretrained vision transformers in robot policies preserves rich spatial structure and enables strong visuomotor reasoning, but dense representations are computationally expensive and contain substantial task-irrelevant information such as background clutter and distractors. We introduce TokenGazer, a policy for selecting a small subset of task-relevant tokens from frozen dense visual representations while preserving the information needed for control. Rather than training the selector with reconstruction, which favors visual fidelity instead of action relevance, TokenGazer scores candidate subsets according to how accurately a downstream action-prediction model can recover expert actions from the selected tokens. We optimize the discrete selection policy using Group Relative Policy Optimization (GRPO) and compare it with a Gumbel-Softmax relaxation. A downstream diffusion policy is then trained using only the selected tokens, reducing its conditioning sequence length without fine-tuning the pretrained vision encoder. Across PushT, Cube, cluttered variants, and language-conditioned LIBERO, TokenGazer matches or exceeds dense full-token policies while using substantially fewer visual tokens. On PushT, learned selection outperforms the full 256-token policy and same-budget baselines, while on Cube it reaches parity with dense conditioning using a fraction of the tokens. TokenGazer also shows improved robustness to visual clutter and reduces the recurring cost of downstream policy training and deployment. Code is available at https://anonymous.4open.science/r/tokengazer-release-E075.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.