acceptodds
Under review as a conference paper at ICLR 2027

NDB-JEPA: Label-Efficient Grasp Detection from Frozen I-JEPA Features

Abstract

As robots become increasingly prevalent in real-world settings, their perception models must become label-efficient and robust to distributional shifts not represented in their labeled training data which poses a considerable challenge for end-to-end grasp detection models typically trained on limited or synthetic datasets. We introduce NDB-JEPA, a label-efficient grasp point detection perception model centered around I-JEPA. Our proposed model takes advantage of the robust, generalized feature extraction of I-JEPA, a self-supervised joint-embedding predictive foundation model. Our model trains lightweight grasp heads on top of the frozen image patch outputs of a 632M-parameter I-JEPA encoder, predicting grasp confidence, orientation, and width per embedded patch. We ask whether these frozen features can replace end-to-end training for the grasping problem. On the Cornell Grasping Dataset, NDB-JEPA reaches 86.9% Top-1 rectangle success, within 1.5 points of an end-to-end GR-ConvNet baseline, and among large pretrained encoders, I-JEPA features outperform MAE, supervised ViT, and DINOv2 alternatives. The frozen representation is also markedly more label-efficient and robust: with only 10% of the training data, NDB-JEPA attains roughly 74% success versus 40% for GR-ConvNet trained from scratch, and it retains about 94% of its clean accuracy under severe darkening while the end-to-end baseline collapses to 1.1%. While end-to-end models retain a small edge in clean accuracy, our results indicate that frozen joint-embedding predictive representations offer a data-efficient, robust basis for grasp perception.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.