PercepDraft for Efficient Generative Perception
Abstract
Generative perception models express different visual perception tasks as one generation process, in which sparse tasks such as detection are decoded autoregressively into structured symbol sequences and dense tasks such as segmentation, depth and surface normals are produced as pixel-level perception maps through iterative denoising. Both generation processes call the full backbone repeatedly, which makes inference expensive. We find that the structural regularity inherent in perception outputs makes their generation predictable from step to step: the serialization pattern of a structured sequence is largely fixed, and the global spatial layout of a dense perception map is essentially formed early in denoising. Based on this observation we propose , which keeps the pretrained backbone frozen and introduces two lightweight drafters, a token drafter and a trajectory drafter. They respectively propose several future tokens, or future denoising velocities with per-offset error estimates used to select a trusted horizon. The backbone then verifies the proposals in one call, so that fewer backbone calls are needed at inference. Experiments confirm the effectiveness of across multiple generative perception models. On , it achieves end-to-end speedups of up to 1.6, 2.0, 6.3, and 5.4 on detection, panoptic segmentation, depth estimation, and surface normal estimation, respectively, with limited changes in task performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.