Dense Checkpoints, Sparse Execution: Retraining-Free Acceleration of Query-Based Segmentation
Abstract
Query-based segmentation repeatedly materializes dense query-pixel tensors during decoding and output processing, generating substantial memory traffic at high image resolutions. We use a frozen checkpoint's predicted masks and logits as an input-dependent execution plan to reduce this traffic. Mask-guided sparse cross-attention with a fixed budget of 128 feature cells per query preserves dense-level panoptic, instance, and semantic accuracy across compatible checkpoints on Cityscapes, COCO, and ADE20K, without retraining or checkpoint-specific tuning. Coarse intermediate mask prediction at the next attention stride avoids temporary stride-4 masks. Task-sufficient output processing evaluates masks only in regions that can affect the final output. For panoptic segmentation, we find that masks can influence pixel assignments even where they predict background, and derive conditions for safely skipping computation. Given identical FP32 logits from a released Mask2Former R50 checkpoint, output processing exactly reproduces the original panoptic maps, instance masks, and labels on all 500 Cityscapes validation images. On Mask2Former at native Cityscapes resolution, these transformations reduce FP32 DRAM traffic by 67% in the transformer decoder and over 80% in output processing. On an AGX Orin, output processing achieves a panoptic speedup and a instance speedup over optimized dense processing. With a PPM-FPN pixel decoder under mixed precision, complete-model inference achieves speedups of up to for panoptic and for instance segmentation over an optimized dense reference. Dense segmentation checkpoints can therefore be deployed more efficiently without retraining by targeting the inference paths that dominate memory traffic.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.