CoCoON: Contrastive Consistency Over Environments for Robotic Manipulation
Abstract
Robotic manipulation policies often generalize poorly to unseen environments. We study cross-environment generalization, where policies must generalize across heterogeneous environments with substantial variations in appearance, geometry, and spatial layout. Our key insight is task-state-conditioned environment invari- ance: visual representations should preserve task-relevant state information while suppressing environment-specific variations. Specifically, observations at simi- lar manipulation states should have similar representations across environments. However, conventional contrastive learning is ill-suited to manipulation states that vary continuously rather than forming discrete instances or classes. We therefore construct continuous task-state similarities as soft contrastive supervision. This in- troduces two challenges: densely sampled trajectories produce diffuse and weakly discriminative soft targets, while similarity-based supervision does not explicitly align similar states across environments. We propose CoCoON (Contrastive Con- sistency Over Environments), a novel task-state-guided cross-environment con- trastive learning framework. Concretely, we first construct continuous task-state similarities from task-relevant information as soft contrastive targets. Secondly, we introduce a Task-State Similarity Calibration module to sharpen the soft targets through stage gating and a centered-sigmoid transformation. Thirdly, we intro- duce Cross-Environment Feature Alignment to explicitly align features of similar manipulation states across environments. Extensive experiments across diverse manipulation tasks and heterogeneous environments demonstrate that CoCoON achieves state-of-the-art zero-shot generalization to unseen environments.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.