GeoLACE: Geometric Layerwise Action-Conditioned Enhancement for Vision-Language-Action Models
Abstract
Generalist vision–language–action (VLA) models benefit from vision–language pretraining but have limited geometric understanding for precise manipulation. Promising approaches use geometric foundation models (GFMs) to provide VLA policies with geometric priors from RGB observations. However, these methods often lack a systematic layer-selection strategy that jointly considers geometric complementarity and relevance to action prediction. This can lead to redundant or underutilized geometric features, limiting gains in manipulation performance. Incorporating geometric knowledge can also disrupt pretrained VLA representations, potentially compromising vision–language capabilities. To address these challenges, we propose GeoLACE, a plug-and-play framework that selects complementary GFM layers and action-relevant VLA layers. It performs layerwise geometric fusion using multi-level features extracted from current RGB observations. We further introduce a two-stage training strategy that enables the policy to better exploit these priors while preserving its pretrained capabilities. We integrate GeoLACE into representative VLA backbones with different architectures. Experiments on LIBERO, RoboTwin, and two distinct real-world bimanual platforms demonstrate improved manipulation performance, with particularly substantial gains in real-world manipulation. GQA evaluation further supports better retention of pretrained vision–language capabilities than the compared methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.