acceptodds
Under review as a conference paper at ICLR 2027

S4VY: Segment Anything in Feed-Forward 4D Visual Geometry

Abstract

Accurate instance segmentation in dynamic scenes is important for downstream applications such as robotics and autonomous driving. Existing Segment Any- thing models operate primarily on 2D image or video masks and preserve iden- tity through sequential memory, while promptable 4D instance segmentation built upon feed-forward visual geometry remains underexplored. We introduce S4VY, a Segment Anything model built on feed-forward 4D visual geometry. From a set of RGB observations, S4VY transforms shared visual-geometric features into an exhaustive set of class-agnostic 4D instance masks through a space-time query decoder, with each persistent object query binding one entity across all observa- tions. This representation supports prompt-independent segmentation as well as point- and box-conditioned selection, without requiring a seed mask or temporal ordering. We further develop an agentic harness for natural-language grounding in the large observation space of a 4D scene. Active tree search identifies rele- vant frames without scanning every fixed window; a dual-stream grounder com- bines fine-grained VLM visual priors with geometry-consistent instance features through complementary bounding-box prediction and object-query matching; and an independent critic selects the final 4D instance mask from their predictions. Extensive experiments demonstrate state-of-the-art 4D instance segmentation and strong language-guided grounding performance under a unified evaluation span- ning static and dynamic scenes.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.