GlimpseBC: Learning Saccadic Visual Policies from Global View Demonstrations
Abstract
Visual policies often process entire camera frames at every step, despite only small regions being relevant to control. In this paper, we study behavioral cloning subject to a hard pixel constraint. We introduce GlimpseBC, a simple learning algorithm in which a model jointly learns where to look and what to do from offline global view trajectories. Our model uses a differentiable crop to learn where to look from full-frame trajectories during training and a hard crop during inference, ensuring that it cannot access full-frame observations. On RoboMimic and Push-T, our models are competitive with full-frame policies (with comparable architecture) while reducing the computational cost in both training and inference. We additionally use the cluttered LIBERO-Spatial benchmark to study where narrow glimpses struggle.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.