EquiKey: Object-Centric SE(3)-Equivariant Keypose Guidance for One-Shot VLA Adaptation
Abstract
Vision-Language-Action (VLA) models acquire broad semantic and visuomotor priors from large-scale robot data, yet adapting them to novel manipulation tasks remains challenging when only a single demonstration is available. In particular, a generalist VLA may recognize task-relevant objects and reuse familiar motor skills while still failing to infer the precise three-dimensional configurations required for successful execution. We introduce EquiKey, an object-centric SE(3)-equivariant keypose guidance framework for one-shot VLA adaptation. Given a single demonstration of a novel task, a lightweight task-specific SE(3)-equivariant predictor learns to infer a sparse sequence of task-relevant keyposes. These predicted keyposes provide explicit geometric guidance to a pretrained keypose-conditioned EquiKey-VLA, whose backbone remains frozen during task adaptation. During execution, the predicted keypose sequence is temporally scheduled as high-level geometric targets, while EquiKey-VLA generates dense continuous actions conditioned on the latest observation. On 10 one-shot OOD tasks in RoboTwin 2.0, our method achieves average success rates of 77.1% and 76.9% under the Easy and Hard settings, improving the strongest VLA baseline by 42.5% and 44.6% and the strongest task-specific one-shot adaptation baseline by 26.5% and 33.6%, respectively. In real-world one-shot Task-OOD experiments, our method achieves a 73.3% success rate, substantially outperforming the strongest baseline at 11.7%. These results demonstrate that sparse SE(3)-equivariant keypose guidance provides an effective interface for combining geometric inductive biases with the semantic and visuomotor capabilities of pretrained VLA models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.