VideoGym: Training Memory-Augmented Video Agents with Modality-Aware Curricula
Abstract
Long-video QA involves sparse evidence distributed over extended multimodal contexts, making retrieval-based memory a natural interface for revisiting relevant cues instead of relying on fixed keyframes. We instantiate this interface as VideoAgent, a memory-augmented agent that indexes a video into text, graph, and visual stores and retrieves evidence on demand through a multi-turn planner, with the planner trained by reinforcement learning (RL) to coordinate retrieval across modalities. This training, however, is bottlenecked not by capacity but by the modality skew of natural video QA data: text-heavy questions dominate at 58% of samples, while graph-, visual-, and time-range questions appear in only 15%, 17%, and 10% respectively. Fixed-distribution RL therefore concentrates updates on the already-frequent text mode and under-trains the remaining modalities, and naive reweighting of raw data alone is insufficient because rare modalities are scarce in absolute terms. We introduce VideoGym, a curriculum-based training environment that adapts VideoAgent's training distribution online: VideoGym treats per-batch sample selection as an adversarial bandit over (modality, mode) arms and uses EXP3 algorithm to shift sampling toward arms with low answer-level accuracy To supply the augmented arms, a TaskGenerator-Verifier pipeline produces hard-case variants that remain answerable within the same video. Across five long-video benchmarks, VideoGym lifts RL accuracy from to on average, surpassing standalone GPT-5 () with an 8B backbone. Gains concentrate on the under-represented modalities: 10.8 pp on graph-heavy and 7.8 pp on visual-heavy questions, at sample efficiency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.