SAVE: Self-Compressing Agents for Long Video Evidence via Multi-Question Pressure
Abstract
Long-video understanding agents are increasingly used to tackle complex long video question answering tasks through extensible tools and multi-step interactions. However, active evidence search in long videos introduces two key bottlenecks: high inference cost caused by repeated tool calls and lengthy observation sequences, and reasoning distraction caused by irrelevant or redundant evidence. Existing methods typically compress visual tokens or textual chains of thought, but overlook tool-mediated evidence search, which dominates the cost of long video agents. We propose (elf-compressing gents for Long ideo vidence), a training-time self-compression framework for long video agents. SAVE uses another question from the same video as a Pressure Question and jointly solves it with the Main Question within a shared multi-turn video-tool rollout. The resulting implicit resource pressure encourages the agent to allocate evidence-search steps more selectively. SAVE then converts the resulting joint rollouts into single-question supervision by retaining target-relevant and shared tool interactions while removing pressure-question-specific searches and answers. SAVE supports both supervised fine-tuning and reinforcement learning. Experiments on multiple long video benchmarks show that SAVE reduces tool calls and token consumption while maintaining competitive accuracy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.