acceptodds
Under review as a conference paper at ICLR 2027

Do You Need Proprioceptive States for Spatial Generalization in Visuomotor Policies?

Abstract

Visuomotor policies commonly combine images with proprioceptive states, yet such states can limit transfer to unseen spatial layouts. We empirically investigate State-free Policies, which remove proprioceptive state inputs, in the setting of spatial generalization from demonstrations collected under constrained spatial layouts. We study how state-input choices, action representations, and task-relevant visual coverage affect generalization. We identify that State-free Policies consistently outperforms the evaluated state-based alternatives, suggesting that reliance on state cues associated with training layouts can limit spatial generalization. This benefit depends on action representation, with relative end-effector (EEF) actions emerging as a key enabling condition. With this action representation, sufficient visual coverage for capturing task-relevant regions further supports State-free Policies' spatial generalization. Under these conditions, State-free Policies shows an increase in average success from 0% to 86% for height generalization and from 6% to 64% for horizontal generalization, compared with matched state-based policies. These findings characterize when state removal supports spatial generalization from demonstrations with constrained spatial coverage, and how action representation and visual coverage shape this benefit. Real-world robot videos are in supplementary materials.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.