FastSAM2: Fully Convolutional Networks for Fast Segment Anything in Images and Videos
Abstract
Segment Anything Model 2 (SAM 2) provides a unified interface for promptable image and video segmentation, but repeated high-resolution encoding andglobal memory attention remain costly. We investigate whether a deliberatelylocal, content-conditioned memory operator can preserve quality when cross-frame correspondence is predominantly local, while trading flexible long-rangeretrieval for lower latency. We present FastSAM2, a unified fully convolutional image-video model whose Cross-frame Convolutional Aggregation (CCA) module combines temporal mixing, similarity gating, large-kernel depthwise propagation, memory-entry weighting, and gated fusion without forming a global query-key matrix. The model is trained only on public SA-1B and SA-V annotations,without pseudo-labels or distillation. Under a matched 1024 x 1024, batch-one,memory-write-inclusive protocol, FastSAM2 reaches 51.7 AP on LVIS and 132.5FPS on H20, versus 41.6 FPS for SAM 2 B+; its 3.18-3.20x speedup transfersacross three GPUs. FastSAM2 remains within 0.2-0.3 points of SAM 2 B+ onfour short/medium-duration video benchmarks, but trails SAM 2.1 B+ on SA-V(74.9 vs. 78.2) and long-term LVOS v2 (72.9 vs. 78.2). The resulting systemoffers a favorable measured efficiency-quality operating point, while its degra-dation under long absence and reappearance is consistent with the limitations oflocal spatial interaction.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.