Efficient 3DGS-Based Affordance Reasoning via Semantically Structured Shape Tokens
Abstract
3D affordance reasoning requires accurately mapping high-level interaction intent to fine-grained local functional regions. However, existing 3DGS-based affordance methods typically perform cross-modal interaction directly between spatially organized low-level Gaussian features and textual intent, without effectively modeling geometric relationships through hierarchical semantic structures. This can lead to incomplete localization and local fragmentation in object structures, while also introducing substantial computational redundancy. Meanwhile, reliance on low-level spatial patterns limits the abstraction of transferable functional structures and cross-category generalization. To address this issue, we propose ST-Afford, an efficient framework for 3DGS-based affordance reasoning that organizes dense Gaussian geometry into compact semantically structured representations and further dynamically selects task-relevant geometric information based on interaction intent. Specifically, we introduce a Semantically Structured Geometry Extractor, which compresses and encodes high-dimensional 3DGS features into compact shape-prior tokens. Based on these tokens, an Intent-Guided Dynamic Fusion module uses interaction text as a condition to dynamically select task-relevant geometric semantic information and fuse it with local Gaussian features, enabling accurate localization of fine-grained functional regions from high-level interaction intent. Experimental results show that ST-Afford substantially reduces inference overhead while generating more complete and accurate affordance reasoning, and achieves stronger zero-shot cross-category generalization on unseen categories.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.