FastSAM3:Accelerating SAM3 with Lazy Detection and Tracking
Abstract
SAM3 achieves robust performance in video referring segmentation, benefiting from the collaboration and mutual disambiguation between its detector and tracker. However, we observe that the dense per-frame computation of this dual-branch design introduces substantial temporal redundancy, resulting in high computational cost and inference latency: (1) Redundant detection execution. The detector is activated at every frame by default. However, in scenarios such as stable scenes, absence of newly appearing objects, or minimal visual changes, the semantic information becomes largely redundant. (2) Redundant tracking memory. The tracker establishes object associations via cross-attention over a large memory sequence. However, during phases with limited dynamics and distractions, the contribution of most memory frames is marginal, thus involving all of them in computation leads to unnecessary resource consumption. To address these issues, we propose FastSAM3, which introduces lazy detection and tracking to address temporal redundancy and accelerate inference. Specifically, for the detector, we propose Entrant-Predictive Gating, which leverages tracking signal and early-layer responses to assess the likelihood of new object emergence, thereby selectively trigger detection on demand. For the tracker, we propose Divergence-Aware Gating, which switches between sparse memory and full memory based on tracking divergence, performing full computation only when necessary. Without any training, FastSAM3 achieves adaptive computational budget allocation, delivering 1.77 end-to-end speedup with only 0.6 performance drop on the SA-V test set, where lazy detection and lazy tracking provide 2.17 and 1.57 speedup, respectively.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.