FaSTra: A Faster Yet Stronger Multi-Trajectory Learning Framework for Reasoning Video Object Segmentation
Abstract
Reasoning Video Object Segmentation (RVOS) aims to identify target objects from complex language descriptions and generate temporally consistent pixel-level masks across video frames. Existing methods either rely on fixed keyframes, making propagation vulnerable to temporal variations, or iteratively search for keyframes at the cost of increased inference overhead. We observe that trajectories initialized from different temporal frames exhibit complementary reliability across different temporal regions. Motivated by this observation, we propose \ourmethod, a faster yet stronger multi-trajectory learning framework that reformulates keyframe selection as frame-wise trajectory selection. \ourmethod employs a temporally diverse frame selector to sample a small set of keyframes and generate multiple candidate trajectories with a lightweight propagator, while a complementary trajectory router dynamically selects the most reliable trajectory at each frame. To enhance multi-trajectory learning, we construct a best-of-trajectories teacher and introduce tri-level propagation distillation as well as utility-aware route distillation to strengthen propagation and routing, respectively. We further introduce GRPO-based frame-impact optimization to align frame-level routing decisions with segmentation improvements. \ourmethod-Qwen3-VL-4B outperforms Sa2VA-Qwen3-VL-4B by 12.9 and 6.5 & points on MeViS and Ref-DAVIS17, respectively, while improving frames per second (FPS) by 26.9-33.2% across videos of 60, 120, and 150 frames.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.