From Partial Observation to Complete Understanding: Unified 3D Semantic and Affordance Reasoning from a Single Image
Abstract
Real-world robots often receive only partial observations but require complete 3D understanding of object semantics and affordances, including in occluded regions, to guide actions. Existing methods address only parts of this problem: 2D approaches are limited to visible regions, while 3D approaches typically model semantics and affordances separately, making cascaded solutions prone to error propagation. We formulate a unified task that jointly grounds semantic and affordance queries over the full 3D object from a single image. This task involves two sources of heterogeneity. First, features derived from 2D images and 3D shapes follow different representation mechanisms. We align and fuse them via optimal transport, enforcing spatial and cross-instance consistency. Second, semantic and affordance terms, such as arm and move, can be distant in language space despite referring to the same region. We inject spatial–functional cues into text embeddings to capture such part–function correspondence while preserving part distinctions. We further construct a multi-level, dual-attribute dataset with semantic and affordance annotations at the region, instance, and category levels. Experiments in simulation and the real world demonstrate consistent improvements over 2D- and 3D-based baselines under occlusion.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.