SeqDETR: Bridging the Query-Feature Gap in DETR via Deformable State Space Models
Abstract
Detection transformers refine object queries through a fixed three-stage pipeline: self-attention models relationships between queries, cross-attention aligns each query independently to image features, and feed-forward networks refine predictions locally. Despite their success, this pipeline has a structural gap, as no single component jointly models inter-query context and image features at the same time. Self-attention sees queries but not the image, while cross-attention sees the image but treats each query in isolation, remaining blind to what neighboring queries are attending to. We introduce SeqDETR, which fills this gap with a novel Deformable State Space Module (DSSM). For each object query, DSSM deformably samples K reference features from the encoder output. It then constructs an interleaved sequence of (image feature, query) pairs across all queries. A selective state space model processes this sequence and propagates a joint inter-query and image-feature representation through its recurrent state. This makes DSSM the first decoder operation in DETR-based models to simultaneously reason over both. Our analysis shows that the performance gain comes from two main factors: the interleaved sequence design that connects queries with their local image features, and the selective context propagation of the Mamba model. Component analysis further demonstrates that DSSM is highly efficient, as it approximates most of the full attention pipeline’s performance on its own. It is also complementary, yielding additional gains when combined with standard modules. SeqDETR achieves a 59.5% relative improvement on LVIS rare categories (+5.5 AP) and +4.0 AP on the 13,204-category V3Det benchmark. It also improves performance on COCO across different backbones and baselines. Code will be available upon acceptance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.