acceptodds
Under review as a conference paper at ICLR 2027

Gated Depth Attention: Preserving Vision-Language Understanding for Generalist Robot Policies

Abstract

Vision-language models (VLMs) underpin modern vision-language-action (VLA) policies, yet remain fragile on spatially grounded queries and manipulation scenarios where 3D geometry matters. Prior 3D enhancements typically either align visual representations to geometry-centric teachers or explicitly inject depth/geometry tokens, both of which can perturb the pretrained RGB–language interface. We propose Gated Depth Attention (GDA), a lightweight module that distills a latent geometry branch from a frozen 3D foundation teacher and fuses it as a gated residual. GDA retrieves geometry through cross-attention while a learned per-channel gate controls how strongly the retrieved correction modifies each visual token; the external 3D teacher is removed at inference. Across VLM benchmarks spanning relational grounding, VQA/OCR, and robustness, GDA consistently outperforms alignment and concatenation baselines and matches or improves strong RGB-only models, yielding a better trade-off between geometry sensitivity and generic vision-language utility. When transferred to VLA policies, GDA also improves manipulation success in LIBERO, RoboTwin 2.0, and targeted real-world spatial/contact tasks. Teacher-swap experiments further preserve the benefit across two distinct 3D foundation teachers. These results suggest that stronger 3D supervision does not automatically preserve vision-language compatibility, and that geometry is better integrated as a selective refinement than as an always-on replacement.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.