Re-Omni: A Two-Pass Reinforcement Learning Recipe for Agentic Video Understanding
Abstract
Agentic video understanding commonly uses multi-turn loops to alternate reasoning and evidence acquisition. However, multi-turn configurations generate thousands of output tokens yet underperform the instruction-tuned supervised fine-tuning (SFT) baseline based on question answering (QA) data for limited model scales. To address these problems, we introduce Re-Omni, a two-pass reinforcement learning (RL) recipe enabling direct training from a strong instruction-tuned model without teacher-generated chain-of-thought (CoT) annotation or a separate CoT cold-start SFT stage. Re-Omni combines three mechanisms: re-watch revisits a predicted relevant interval at higher spatial fidelity; re-ask repeats the question after the new evidence; and re-answer refines an initial direct answer. Re-Omni improves accuracy over QA-SFT on all six multiple-choice video benchmarks. Against the full-history multi-turn configuration, it gains more than 4.1 percentage points across the benchmarks with two viewing rounds, reducing generated output tokens from 2,681.2 to 101.7 and inference GPU time from 18.3 to 9.5 seconds per question. Code, models, and data will be publicly released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.