CountRoute: Pattern Readout and Evidence-Grounded Counting for Streaming Video Counting
Abstract
Counting in a streaming video is usually treated as a single numerical inference over an increasingly long visual context. We argue that the first decision is how the count can be established. For some events a temporal pattern — the recurrence of a visual state, or an abrupt change of shot — is enough to read the count off directly. For others the count has to be grounded in semantic evidence: occurrences must be found, told apart, and added up. Video language models treat both as one direct prediction, and do poorly where a pattern would have sufficed. Evidence- grounded counting brings a second choice. Inspecting short visual blocks keeps local evidence, but each additional empty block is another opportunity for a false positive, and an event spanning blocks may be counted twice; reading a continuous textual record avoids this block-wise accumulation, but cannot recover events it never wrote down. We introduce CountRoute, a training-free streaming agent built on a query-agnostic memory of the stream. It reads a count directly off a temporal pattern only after verifying that the corresponding signal is present; otherwise it uses the record to select relevant cached blocks, and moves from inspecting their frames to reading the record as the number of selected blocks — and hence the exposure to accumulation error — grows. On SVCBench, CountRoute achieves the best overall score, ahead of existing offline and online methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.