SELF: SE(3)-Equivariant Latent Flow Policy for Robot Manipulation
Abstract
Equivariant visuomotor policies exploit geometric symmetries to improve data efficiency in robot manipulation, but repeatedly evaluating expressive equivariant networks over action sequences can make inference expensive. In this paper, we propose , an E(3)-quivariant atent low policy that directly transports observation latents to action latents. SELF encodes observations and action chunks into the same representation of invariant scalars and equivariant vectors, enabling a lightweight velocity field to transport between endpoints with identical geometric transformation laws. To retain local geometry throughout transport, the velocity field accesses persistent scene context containing spatially anchored features and higher-order geometric information. This design reduces the cost of iterative generation while maintaining access to structured scene geometry. The latent transport is equivariant by construction under joint transformations of the source latent and scene context. The observation encoder, action autoencoder, and velocity field are trained jointly with reconstruction and flow-endpoint supervision. Experiments on eight simulated and three real-world manipulation tasks demonstrate that SELF matches strong equivariant policies in success rate, while achieving approximately 2.9 and 23.8 faster end-to-end inference than E3Flow and SDP-DDPM, respectively. Code will be released.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.