Event-Centric Video Scene Graph Generation
Abstract
Top-leading solutions for Video Scene Graph Generation (VSGG) typically aggregate temporal context over frames, and are trained and evaluated one frame at a time. Though demonstrating steady gains in Recall@, they remain less reliable on brief relations and rare predicates, and the metric hardly notices: smoothing a perfect prediction over three frames keeps 92.5% of its Recall@ but only 21.2% of its single-frame relations. In this paper, we reveal the underlying nature of this phenomenon: counting frames weighs every relation by its duration, so the metric, the loss and the temporal context learned from it all favor long relations, although nearly half of all relation instances last a single frame. To this end, we introduce Kairos, which formulates VSGG in event time, a clock that advances only when a relation changes. Drawing inspiration from the inspection paradox, we show that Recall@ exceeds event recall, which counts every relation instance once, exactly by a covariance term that we call the duration bias. Specifically, Kairos: i) minimizes the risk over relation instances with Horvitz-Thompson weights; ii) accumulates a change hazard, learned from the relation changes in the annotations, into an event clock; and iii) reads out a trained model in event time with this clock. Extensive experiments on three setups of Action Genome demonstrate that its gains concentrate on short relations, transitions and the predicate tail at maintained frame-level recall. In a controlled study, it raises the event recall of the ten rarest predicates by 2.9 points (8.5%) and reveals that the duration bias of temporal models lies in what they count rather than in how they pool.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.