acceptodds
Under review as a conference paper at ICLR 2027

Pretraining registers with covisibility for long-sequence 3D reconstruction

Abstract

Feed-forward reconstruction transformers rely on all-to-all patch attention across views (global attention), whose quadratic cost in views and patches prevents scaling to a large number of views. Prior attempts to reduce this cost by routing cross-view information through a compact set of per-frame tokens while keeping patch attention intra-frame (register attention) report a significant accuracy drop. We show this gap is largely an initialization issue, not a capacity limitation of register attention. By pretraining the registers on a multi-view covisibility pretext task extended from a two-view, static-scene setting to long, dynamic sequences with a new view-conditioned, query-based segmentation head before finetuning on 3D reconstruction, our model recovers most of the global attention performance. On long sequences, where global attention’s quadratic cost becomes prohibitive, our approach scales efficiently and enables training on long sequences. This curriculum makes register attention a scalable alternative that outperforms chunked global attention on hundreds of views.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.