Training-Free Spatially Grounded 3D Object Encoding
Abstract
We present X3DEnc, a training-free framework for spatially grounded 3D objects encoding. Our central insight is that such a 3D object is characterized not only by its geometry, but also by its appearance and spatial pose. Different downstream tasks may require different relative emphasis on these components. X3DEnc decomposes a 3D object into normalized geometry and color fields within the unit ball and a pose transformation. It then lifts the finite-dimensional pose vector to a harmonic field and projects geometry, color, and pose onto a shared 3D Zernike basis. This common spectral coordinate system supports separate or joint encoding. The geometry-color-pose joint encoding allow emphasis-adjustable encoding. X3DEnc accommodates voxel, signed-tetrahedral, and surface-triangle representations, covering volumetric solids, watertight meshes, hollow objects, and open surfaces. Theoretical analysis and various experiments demonstrate that X3DEnc can redirect similarity toward geometry, color, or pose while preserving a single encoding formulation. Rather than competing with neural encoders, X3DEnc provides an alternative when training data are limited or when interpretable, jointly controllable multi-component encoding is required.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.