acceptodds
Under review as a conference paper at ICLR 2027

3AM: 3egment Anything with Adaptive Geometric Consistency in Videos

Abstract

Consistent video object tracking must preserve identity under large viewpoint changes, temporary absence, and nearby distractors. Memory-based trackers such as SAM2 rely on appearance matching and fail in these cases, whereas 3D-based methods preserve identity across views but need camera poses, depth, or a reconstructed point cloud, run offline, and are not promptable. We present 3AM, a promptable online tracker that achieves geometric consistency from RGB alone by coupling memory-based tracking with the online multi-view reconstruction model MUSt3R. We identify the MUSt3R layers whose features preserve instance identity across views and fuse them into the tracker's image features by feature-wise modulation whose parameters are decoded from routing tokens. Geometry alone is insufficient, since a location cue holds only while the object stays in place; the routing tokens therefore traverse the tracker's memory attention and decide per frame how much the fusion should rely on geometry, so that one model handles static and moving objects without a hand-designed fallback. On wide-baseline benchmarks built from ScanNet++ and Replica, 3AM reaches 90.6% IoU and 65.7% tracking recall on our selected subset of ScanNet++ (+12.3 and +28.2 points over the best prior tracker), matches or exceeds SAM2 on standard VOS and VOT benchmarks, and transfers zero-shot to ego-exo correspondence on Ego-Exo4D.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.