SMART-SLAM: Feed-Forward SLAM with self-adaptive submap memory
Abstract
Long-stream visual geometry needs access to past observations without keeping all scene tokens in active attention. We introduce SMART-SLAM, a feed-forward SLAM system with self-adaptive submap memory. Finalized token banks remain in a host-side archive; each online forward activates only the current submap, its predecessor, and up to k retrieved historical submaps. A detached Q⊤V feature- affinity signal adapts a lightweight retriever while the geometry backbone stays frozen. At end of sequence, Boundary-Pair Backend Re-query (BPBR) composes sparse adjacent relations into an open SE(3) chain for poses and point maps, without loop closure or iterative geometric optimization. On matched RTX A6000 prefixes, bounded-context online inference is 3.83× faster and uses 60.8% less resident GPU memory than native full-history attention. BPBR reduces the online pose error by 53.2% on TUM and 39.7% on Oxford. Indoor reconstruction reaches F-scores of 74.70% on NRGBD and 83.80% on 7-Scenes. Host-side storage grows with the stream while active attention remains bounded.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.