Where Precision Matters: Selection-Regret Learning for Quantized SAM 2
Abstract
Segment Anything Model 2 (SAM 2) generates multiple mask candidates for each prompt and selects one using its predicted-IoU head, but existing quantization objectives evaluate candidates independently and do not protect this selection process. We propose a mixed-precision quantization method that jointly learns the weights and a bit width for every layer, together with a teacher-relative selection-regret objective that penalizes only the additional mask quality forfeited by the quantized selector compared with its full-precision teacher. A controlled precision sweep over video propagation units shows that hard selection accuracy degrades significantly already at eight bits, even when candidate quality remains within half a percent of full precision, and that the resulting damage increases with propagation depth. Independently, precision learned using only the segmentation loss concentrates around the same predicted-IoU interface, identifying it as a critical quantization bottleneck. At an average precision of 4.79 bits, our method achieves 71.1 SA-V J&F and 0.743 COCO mIoU on SAM 2.1 Base+, whereas two mixed-precision baselines at the same budget preserve COCO accuracy but fall below 2.4 on video. At a matched bit budget, the regret objective raises candidate-set and selected-mask quality, each with a paired interval excluding zero.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.