Reasoning Complex Relations: Object-level Cognitive Hypergraph for 3D Intention Grounding
Abstract
Understanding and grounding human intentions is an important direction in embodied intelligence. Recently, 3D Intention Grounding (3D-IG) has been proposed to localize target objects in 3D scenes according to abstract, ambiguous, and non-descriptive natural-language intentions. Unlike conventional 3D Visual Grounding, which aligns 3D objects with explicit referring descriptions, 3D-IG requires the model to infer latent functional semantics, such as object attributes and affordances, from human intentions. Therefore, 3D-IG requires reasoning beyond simple cross-modal matching, involving intention understanding, object relations, and scene-level constraints. However, existing methods still largely rely on implicit intent-object alignment, without explicitly modeling multi-object functional relations and scene-level context. To address these challenges, we propose a Cognitive Hypergraph framework that, inspired by the coarse-to-specific process of human cognition, reformulates 3D-IG from local matching into structured reasoning over intentions, functional object groups, and scene context. Specifically, candidate objects are represented as graph nodes, while higher-order hyperedges organize them into intention-conditioned functional groups. By introducing global intention and scene-level priors, the framework jointly models target, intention, functional groups, and scene context, and performs structured reasoning over the cognitive hypergraph to obtain intention-aware candidate representations for final grounding. Experiments on 3D Intention Grounding and 3D Visual Grounding benchmarks demonstrate the effectiveness of our method, showing strong localization performance with improved interpretability and generalization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.