acceptodds
Under review as a conference paper at ICLR 2027

Re-Omni: A Two-Pass Reinforcement Learning Recipe for Agentic Video Understanding

Abstract

Agentic video understanding commonly uses multi-turn loops to alternate reasoning and evidence acquisition. However, multi-turn configurations generate thousands of output tokens yet underperform the instruction-tuned supervised fine-tuning (SFT) baseline based on question answering (QA) data for limited model scales. To address these problems, we introduce Re-Omni, a two-pass reinforcement learning (RL) recipe enabling direct training from a strong instruction-tuned model without teacher-generated chain-of-thought (CoT) annotation or a separate CoT cold-start SFT stage. Re-Omni combines three mechanisms: re-watch revisits a predicted relevant interval at higher spatial fidelity; re-ask repeats the question after the new evidence; and re-answer refines an initial direct answer. Re-Omni improves accuracy over QA-SFT on all six multiple-choice video benchmarks. Against the full-history multi-turn configuration, it gains more than 4.1 percentage points across the benchmarks with two viewing rounds, reducing generated output tokens from 2,681.2 to 101.7 and inference GPU time from 18.3 to 9.5 seconds per question. Code, models, and data will be publicly released.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.