Annotation-Grounded Evaluation of Self-Generated Text Effects in Closed-Loop Streaming Video LLMs
Abstract
Streaming video LLMs can retain earlier generated text in the running context, allowing previous outputs to condition later predictions. We study this retained-output dependence as a distinct property of closed-loop streaming generation. Existing evaluations do not isolate it: supplied-history settings omit the model's own generated context, while end-to-end streaming metrics conflate retained-output effects with newly arriving visual evidence. We introduce an annotation-grounded protocol for evaluating retained-output effects during closed-loop generation. In-loop probing reads current-step predictions at annotated times, while narration judging evaluates generated output against the same timestamped reference timeline. Fixed-run reconstructions on COIN show that retained narration can reduce responsiveness to changing visual evidence, particularly across annotated step transitions. We further introduce Retained-Output Contrastive Decoding (ROCD), a training-free method that contrasts prefix-matched views with and without direct access to retained generated text. Across multiple streaming video LLMs, ROCD increases responsiveness to newly arriving visual evidence, with corresponding changes in both probe readouts and generated narration. Together, the protocol characterizes how retained output affects adaptation to changing visual evidence, while ROCD provides a training-free method for reducing this dependence during streaming inference.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.