acceptodds
Under review as a conference paper at ICLR 2027

FlashTAM: Target-Conditioned Memory Tokenization for On-Device Video Object Segmentation

Abstract

On-device promptable Video Object Segmentation (VOS) aims to track a user-specified target on device-side, imposing even stricter requirements on real-time responsiveness and model compactness. Although on-device VOS has been studied for years, existing methods typically face a trade-off among model size, inference latency, and segmentation quality, and therefore may not fully preserve the segmentation quality of SAM 2 . SAM 2 provides stronger target propagation through its streaming memory, but this memory mechanism also introduces substantial computational overhead, limiting its direct deployment on mobile devices. This creates a central challenge: retaining target-related temporal information for accurate propagation while organizing it into a compact memory representation suitable for efficient computation on mobile hardware. To overcome this challenge, we present a Flash Target Conditioned Memory (FlashTAM) design. FlashTAM comprises two main components: (1) Mask-guided ROI Localization (MRL). MRL reuses the predicted mask to localize a target-centric ROI on the memory grid, avoiding extra parameters or attention matrices. (2) Window-wise Adaptive Pooling (WAP). WAP adaptively pools the target-centric ROI into a near- regular token budget for efficient mobile attention. In summary, MRL decides Where to Sample and WAP decides How to Compress, jointly preserving target-relevant temporal information while reducing memory redundancy. Experiments on several benchmarks (SA-V, DAVIS-2017, MOSE, and YouTube-VOS 2019) validate that FlashTAM achieves a favorable accuracy-efficiency trade-off, notably reaching J&F on SA-V test and FPS on an iPhone 16 Pro, exceeding the second-best on-device method (EdgeTAM) by at least 1.2 J&F and 2.3 FPS.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.