VideoSpy: Evolving through Experience for Adaptive Video Understanding
Abstract
Recent approaches to agentic video understanding are often tailored to domain-specific evidence requirements and task settings, limiting their adaptability across domains. A central bottleneck is that their task-solving strategies remain largely fixed after design, without a systematic mechanism for learning and evolving from execution experience. In this work, we investigate whether a foundational video agent can evolve from execution experience to adapt to different domains without redesigning its framework. To this end, we first build a foundational VideoSpy agent with three complementary visual tools: overview, clip_skim, and frame_inspect. A Reflector combines task-grouped reflection and evidence-based review to extract reusable experience from training execution records and ground-truth references. An Agent Optimizer loads the resulting Skill and source-grounded Memory into the shared agent harness, enabling the same foundational agent to specialize independently for each domain. To assess this mechanism, we establish a unified protocol for domain-specific evolution and held-out evaluation on LVBench, EgoLifeQA, and MedVidBench, covering general long videos, egocentric daily life, and medical videos. With Qwen3.5-4B, evolution raises accuracy by 3.8 and 4.2 percentage points on LVBench and EgoLifeQA, improves nine of ten MedVidBench leaderboard metrics, and reduces inference tokens by 9.7–27.1%. These results show that a shared video agent can adapt to different domains by learning from execution experience.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.