ActFun3D: Task-Conditioned Scene Narrowing for Fine-Grained 3D Functional Grounding
Abstract
Fine-grained 3D functional grounding aims to identify the physical region that supports a natural-language task, such as the handle to pull, the button to press, or the control to adjust. This is particularly challenging in large 3D scenes, where task-relevant regions occupy only a small fraction of the scene and are easily confused with visually or geometrically similar candidates. We introduce ActFun3D, a task-conditioned coarse-to-fine framework that uses task semantics to reduce the effective search space before fine-grained grounding. Given a task, ActFun3D first extracts structured functional semantics describing the target, its context, affordance, and spatial constraints, and uses them to localize a high-recall candidate region over language-aligned 3D features. A trainable local 3D grounder then integrates local geometry, appearance, multi-view visual features, and task semantics to predict the task-relevant functional region at the point level. To further resolve ambiguous predictions, ActFun3D retrieves a compact set of informative posed RGB-D views and projects verified visual evidence back into 3D for multi-view refinement. Experiments on SceneFun3D show that ActFun3D achieves 30.8 in fine-grained 3D functional grounding. Notably, task-conditioned narrowing retains 94.2% of the target while reducing the effective search space to only 15.9% of the scene, demonstrating the effectiveness of narrowing large 3D scenes before fine-grained grounding. Project page: https://anonymous.4open.science/w/ActFun3D-Web
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.