acceptodds
Under review as a conference paper at ICLR 2027

Freeze the Student, Train the Researcher: Learning the Value of Experiments for Streaming Video Understanding

Abstract

Streaming video systems are typically engineered against a fixed benchmark, yet deployment brings unfamiliar scenes, event statistics, and query types, where the tuning loop must start over. What survives such shifts is the research process itself: hypothesizing about failures, running diagnostic experiments, and revising the design. We propose StreamRSI, which keeps the deployed 7B video answerer frozen and instead trains a 27B research policy to conduct this process, learning the value of experiments rather than any single system. The central obstacle is credit: an unsuccessful experiment can rule out a misleading explanation and improve subsequent search, while an immediately successful candidate can teach nothing transferable. StreamRSI therefore assigns controlled evidence credit. Paired continuations restore the same candidate state and differ only in access to a diagnostic result, isolating that result's contribution to later research; a third continuation spends the diagnostic budget on direct search, pricing the experiment's opportunity cost. These comparisons supervise when to investigate, which probe to run, and when to reopen a diagnosis. The discovered inference-time program lifts the frozen answerer to 65.50 OVO-Bench Overall and 80.63 StreamingBench real-time accuracy. Under a matched 20,000 GPU-s budget, full StreamRSI reaches the target quality in all eight seeds and outperforms outcome-only and single-credit training. The resulting research policy generalizes to novel domains, accelerating search on unseen benchmarks even when stripped of its historical archive.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.