acceptodds
Under review as a conference paper at ICLR 2027

SEVA: Self-Evolving VLM Agent for Video Segmentation

Abstract

Existing video object segmentation (VOS) methods process each video independently and discarding accumulated knowledge and re-initializing object representations from scratch, unlike human visual tracking. Separately, self-evolving agent frameworks have shown strong iterative improvement in coding and text domains, but remain unexplored for VOS tasks. We propose SEVA (Self-Evolving VLM Agent), which reformulates tracking as a continuous, memory-augmented reasoning process driven by a vision-language model. SEVA maintains evolving object representations across frames and videos, produces interpretable frame-by-frame reasoning, and self-corrects over batches of videos without additional fine-tuning, reporting a transparent summary of what it has learned after each batch. Evaluated on video benchmarks emphasizing 3D spatial reasoning under occlusion and viewpoint change, SEVA achieves state-of-the-art performance while showing measurable improvement as it observes more videos, demonstrating that self-evolving agentic reasoning extends applicably across the visual task domain.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.