Pretraining registers with covisibility for long-sequence 3D reconstruction
Abstract
Feed-forward reconstruction transformers rely on all-to-all patch attention across views (global attention), whose quadratic cost in views and patches prevents scaling to a large number of views. Prior attempts to reduce this cost by routing cross-view information through a compact set of per-frame tokens while keeping patch attention intra-frame (register attention) report a significant accuracy drop. We show this gap is largely an initialization issue, not a capacity limitation of register attention. By pretraining the registers on a multi-view covisibility pretext task extended from a two-view, static-scene setting to long, dynamic sequences with a new view-conditioned, query-based segmentation head before finetuning on 3D reconstruction, our model recovers most of the global attention performance. On long sequences, where global attention’s quadratic cost becomes prohibitive, our approach scales efficiently and enables training on long sequences. This curriculum makes register attention a scalable alternative that outperforms chunked global attention on hundreds of views.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.