acceptodds
Under review as a conference paper at ICLR 2027

Depth-Aware Visual Manipulation Relationship Reasoning with Gated Decoder-Layer Aggregation

Abstract

Object retrieval in cluttered scenes depends on removal order, since an object that appears graspable may still support or constrain its neighbors. Inferring this order from a single observation requires accurate object localization and recovery of a sparse, directed relationship graph over many possible ordered pairs. We propose a depth-aware model for visual manipulation relationship reasoning that jointly detects objects and predicts this graph from aligned RGB-D input. Specifically, DFormerV2-S produces multi-scale RGB-D features, while Deformable DETR generates object queries that are progressively refined across six decoder layers. For each ordered query pair, our relation head constructs a directed representation from every layer and combines the resulting evidence with pair-dependent sigmoid gates. Under the unified end-to-end protocol on the constructed ManipRelaD dataset, the model achieves 31.68% triplet F1 and 13.91% full-scene accuracy, improving upon the strongest compared baselines by 16.34 and 3.21 percentage points, respectively.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.