acceptodds
Under review as a conference paper at ICLR 2027

SAGE-Grasp: An End-to-End Target-Driven 6-DoF Grasping Framework via Cross-Modal Semantic Addressing and Gumbel-Routed Experts

Abstract

Language-conditioned 6-DoF grasping in clutter requires identifying the referred object while generating physically feasible grasps. Existing task-agnostic detectors disregard semantic intent, whereas segment-then-grasp pipelines are susceptible to upstream segmentation errors and often discard useful contextual information. Thus, we propose SAGE-Grasp, which establishes a unified semantic-geometric representation that couples target understanding with full-scene physical reasoning. Cross-Modal Semantic Addressing (CMSA) and a Gumbel-routed mixture-of-experts architecture provide complementary mechanisms for semantic grounding and expert-routed grasp prediction, enabling the model to focus on the instructed target while retaining the context required for feasible grasping. SAGE-Grasp is jointly trained with text, visual, and zero queries, enabling target-aware semantic grounding while preserving task-agnostic geometric affordances. Experiments on a target-driven extension of GraspNet-1Billion demonstrate stronger target adherence than cascaded baselines and competitive scene-level grasping performance. These results indicate that target-conditioned grasping can be integrated into full-scene prediction while retaining task-agnostic scene-level capability.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.