Coupling trajectories and Masks: An Efficient Sparse-to-Dense Annotation Framework via Bidirectional Enhancement for Event Motion Segmentation
Abstract
Event motion segmentation is bottlenecked by dense pixel-level masks, whose exhaustive annotation is labor-intensive and poorly matched to asynchronous event streams. Existing event-based segmentation approaches are poorly suited as scalable annotation tools: they either expose only discrete frame outputs or require costly mask, optical-flow, or auxiliary-modality supervision, limiting temporal coverage and annotation per interaction. We propose an efficient sparse-to-dense annotation framework that couples trajectories and masks through bidirectional enhancement. In Motion-to-Mask, a learned temporal trajectory field warps events to queried timestamps, forming motion-compensated Images of Warped Events (IWEs) that allow a frozen promptable segmentation model to generate masks from sparse bounding boxes. In Mask-to-Motion, the generated masks provide region-level constraints for mask-aware contrastive trajectory learning, refining the motion estimates used for IWE construction and mask propagation. Because the trajectory field is queryable within each annotation window and the promptable segmentation model propagates masks across query-time IWEs, the framework can produce masks at a denser temporal output rate than the box-prompt rate and transfer them to asynchronous events. Across three datasets, the generated masks enable fine-grained annotation with full or sparse prompts and no pixel-level mask or optical-flow GT; a user study shows up to 389.39% speedup at a 1:3 input-to-output frame-rate ratio across event-data experience levels.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.