Retool-Video: Recursive Tool-Using Video Agents with Meta-Augmented Tool Grounding
Abstract
Tool-augmented video agents typically require the ability to locate evidence, inspect multiple modalities, and integrate observations across long temporal horizons. However, existing approaches face two key challenges: at the *tool-space* level, video agents typically expose only a small set of coarse retrieval and inspection tools, limiting fine-grained evidence processing; at the *action-space* level, hierarchical video agents often leave underspecified how abstract video-reasoning intents are grounded into executable tool calls. To address these limitations, we introduce the MetaAug-Video Tool Library (MVTL), comprising 26 base tools for multimodal evidence acquisition and 108 meta tools for fine-grained operations over intermediate results. Building on MVTL, we propose **ReTool-Video**, a recursive tool-using framework in which a planner issues either executable tool calls or abstract intents, a resolver repairs, substitutes, or recursively decomposes abstract intents into validated tool calls, and the planner is optimized with reinforcement learning. Experimental results show that **ReTool-Video** achieves 72.9, 81.5, and 76.6 accuracy on MVBench, MLVU, and Video-MME, respectively. Under the same Qwen3.5-9B backbone and tool library without planner reinforcement learning, Planner–Resolver grounding improves MLVU accuracy by 13.1 percentage points over planner-only tool selection, demonstrating the benefit of separating high-level intent expression from low-level tool execution. The code and tools are available at [the anonymous repository](https://anonymous.4open.science/r/ReTool-Video).
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.