AcFunSG: Active Evidence Seeking for Open-Vocabulary Functional 3D Scene Graph Generation
Abstract
Functional 3D scene understanding goes beyond scene geometry and object-level semantics to model part-level affordances for embodied interaction. However, incomplete or ambiguous observations can introduce errors in cross-view association and functional reasoning that propagate into the scene graph. Existing methods typically construct graphs from a fixed observation set, with limited support for subsequent evidence acquisition and graph revision. To address this limitation, we propose AcFunSG, a training-free, closed-loop framework for open-vocabulary functional 3D scene graph generation that progressively completes and refines an initial graph through uncertainty-guided evidence acquisition. Given posed RGB-D observations, it builds an initial graph through cross-view association and relation-specific global reasoning while retaining entities and relations unresolved by the current evidence. These uncertainties guide further evidence acquisition. Under a limited budget, Ac-Agents selects informative views and queries, integrates the resulting evidence, and revises affected predictions. A deterministic verifier enforces the same consistency constraints used during construction and accepts only updates that pass structural and evidential checks. The graph is progressively completed and corrected, while indeterminate elements remain unresolved. Experiments on SceneFun3D, FunGraph3D, and FunThor show state-of-the-art overall instance recall, with a functional part mapping Recall@3 of 85.1% and relation F1 of 42.0% on FunThor.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.