acceptodds
Under review as a conference paper at ICLR 2027

TeQCo: Temporal Quota Consistency for Training-Free Video Token Compression

Abstract

Video token compression must control both temporal coverage and task relevance, but sequential stages can conflict: a content-driven allocation stage may be overwritten by a global ranking that removes entire temporal positions. We identify this as an allocation conflict and propose TeQCo, a training-free framework that resolves it through temporal quota consistency. TeQCo allocates candidate capacity from first-order changes in normalized frame representations and aggregates all tokens within each frame into a compact candidate set. Candidates are then scored by prior-corrected task attention and selected within integer frame quotas inherited from the first-stage counts. Content variation determines temporal capacity; question relevance determines which candidates fill it. This design guarantees an exact token budget and a per-frame coverage floor whenever feasible. On the 2,700-question Video-MME benchmark with Qwen3-VL-8B, TeQCo achieves 68.52% accuracy at 510 ms TTFT, a 32.9% TTFT reduction with a 1.00-percentage-point accuracy decrease.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.