AVA-3D: Attribute-based Visual Arbitration for Unified 3D Representation Learning
Abstract
Multimodal pre-training for 3D point clouds has shown strong potential by aligning 3D shapes with corresponding images and language descriptions. However, existing methods often integrate multi-view visual information at the global or view level, overlooking heterogeneous visual attributes across viewpoints and their varying relevance to object geometry. To address this issue, we propose AVA-3D, an Attribute-based Visual Arbitration framework. AVA-3D introduces a Mixture of Arbitration (MoA) module to reorganize view-indexed features into complementary latent visual attributes through dual-axis competitive arbitration, together with a Conditional Attribute Fusion (CAF) module that adaptively routes and transforms these attributes according to 3D geometry. By enabling attribute-level organization across views and geometry-conditioned fusion across modalities, AVA-3D learns more discriminative 3D representations. Extensive experiments on Objaverse, ModelNet40, and ScanObjectNN demonstrate consistent improvements across downstream tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.