acceptodds
Under review as a conference paper at ICLR 2027

FastBench: Can Streaming VLMs Perceive High-Dynamic Real-World Streams?

Abstract

Streaming Video Large Language Models (VLMs) have emerged as a promising paradigm for real-time, continuous video understanding. However, existing streaming benchmarks predominantly target low-dynamic scenarios, masking a critical challenge: under a bounded context budget, a model must trade off long temporal history, spatial resolution, and temporal granularity, and the sparse sampling rates (1–2 FPS) adopted by current systems inevitably miss fast-paced events. To bridge this gap, we introduce FastBench, a benchmark designed to evaluate streaming VLMs on high-dynamic real-world video streams. FastBench is built with a trajectory-grounded pipeline: a VLM proposes candidate QA pairs from native high-frame-rate clips, a low-FPS filter discards pseudo-dynamic questions that remain answerable at 2 FPS, and a Vision Expert Verification stage grounds answers on object trajectories extracted by SAM3 and CoTracker3 to suppress hallucinations, followed by three rounds of human inspection. The resulting 306 QA pairs span eight domains (e.g., sports, gaming, wildlife, and transportation), six capability dimensions, and forward, instant, and backward temporal scopes, each with human-annotated evidence intervals. Alongside the benchmark, we present ProactiveFrame, a training-free baseline that lets a model raise or restore the frame rate of incoming chunks through text tokens, supported by a dual-tier sliding window that buffers recent high-FPS chunks and gracefully degrades them into sparse history. Experiments show that high-dynamic perception remains far from solved: the strongest model, Gemini-3.5-Flash, reaches only 50.7%. Denser sampling helps substantially (Qwen3-VL-8B improves from 32.9% at 2 FPS to 44.6% at 24 FPS) but saturates as history is compressed. ProactiveFrame improves over sparse uniform sampling (+5.4% and +1.5%), yet stays well below oracle-guided focusing, revealing that current VLMs struggle to decide, from the visual stream alone, when finer temporal perception is needed. FastBench establishes a rigorous testbed and offers actionable insights for high-dynamic streaming video understanding.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.