acceptodds
Under review as a conference paper at ICLR 2027

VideoAgentBoost: Evolving Video Agents through Additive Capability Improvements

Abstract

Changing video sources and question demands create a recurring need for video agents to evolve. Multimodal perception, tool use, and long-horizon reasoning make execution feedback costly to obtain, motivating targeted capability improvements from limited execution evidence. We introduce VideoAgentBoost, a framework for additive video-agent evolution guided by capability residuals: reusable descriptions of missing or insufficient behaviors identified by comparing execution trajectories with reference solutions grounded in annotated answers and supporting video evidence. Each round develops an increment for one residual, expanding or refining the relevant capabilities within the current agent. Targeted research and design guide code generation to implement the increment in the affected modules, including the coordination needed to use the improved capability. Before integration, targeted probing examines whether the increment contributes to resolving the selected deficiency through improved evidence acquisition or reasoning. Subsequent executions guide further improvements. We independently evolve DVD, VideoARM, and VideoAgent for five rounds on a shared evolution set. All three agents improve over their starting points on LVBench, LongVideoBench, and Video-MME. DVD gains 6.59 percentage points on LVBench over five evolution rounds, and component ablations further support the effectiveness of the proposed method. Our code and evolution-set annotations will be made publicly available.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.