Align-VAE: Aligning Attention and Geometry in VecSet VAEs for 3D Reconstruction
Abstract
Set-based variational autoencoders (VecSet VAEs) represent 3D shapes as compact sets of latent tokens. Although each token originates from a query at a sampled surface location, whether these latents truly align with geometry remains unclear. We investigate this alignment and identify four issues: (i) attention-sink-like tokens that attract queries from across the shape; (ii) spatially periodic repetitions in decoder attention responses; (iii) mismatched normal information in encoder and decoder query features; and (iv) missing supervision of the intrinsic relation between signed distance fields and surface normals. To address these issues, we introduce Align-VAE with four corresponding modifications: (i) a shared bank of learnable decoder tokens provides a workspace for global computation; (ii) distance-biased cross-attention guides retrieval toward nearby predicted spatial anchors; (iii) normal-free encoder queries align the two stages' access to normal information while retaining normals in the encoder keys and values; and (iv) a screened Poisson-inspired regularizer supervises surface values and normal-direction derivatives. These changes align latent retrieval and field prediction with geometry without enlarging the per-shape code. On the ABO dataset, Align-VAE reduces Chamfer distance by 26.6% compared to the standard VecSet VAE.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.