ProtoTwinMamba: Prototype-Guided Spatial Ordering and Temporal Modeling for Video Polyp Segmentation
Abstract
Video polyp segmentation remains challenging due to foreground–background ambiguity and inter-frame appearance variations. In recurrent spatial modeling, token order shapes how contextual information propagates, yet fixed grid traversals do not explicitly account for foreground and background relevance. Motivated by this observation, we propose ProtoTwinMamba, which uses region guidance not only to refine features but also to organize their spatial processing order. Specifically, a coarse prediction provides soft pooling weights for constructing frame-wise foreground and background prototypes. Prototype similarities are combined with the coarse prior to produce region-specific guidance maps, shared by masked linear attention for feature refinement and two independently parameterized Mamba branches for spatial scan ordering. Sorting tokens by their respective guidance scores groups high-relevance tokens together, forming foreground- and background-conditioned sequences without hard token partitioning. Each branch then restores the original feature grid before bidirectional temporal aggregation at fixed spatial locations. Moreover, a shallow detail-enhancement pathway preserves local detail, while a convolutional decoder fuses shallow and deep features to produce the final segmentation masks. Extensive experiments on SUN-SEG and CVC-612 demonstrate that ProtoTwinMamba achieves state-of-the-art performance on both benchmarks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.