Relation Event Measure: Preserving Event Supervision under Partial Grounding for Panoptic Video Scene Graph Generation
Abstract
Panoptic video scene graph generation (PVSG) commonly constructs relation targets by matching annotations to predicted tracklets. When tracklets are incomplete, this matching process can discard entire events or truncate their positive temporal supervision. We introduce Relation Event Measure (REM), a mass-based formulation that decouples the full event annotation from partial visual groundings. To achieve this goal, we build a Relation Event Measure Prediction to separates event count and temporal allocation from grounding availability and tracklet-pair allocation. It preserves supervision for observed mass on matched pairs and unobserved mass in an explicit unobserved state. To supervise relation learning under REM, we propose the Annotation-Preserving Event Learning, which preserves event coverage and annotated temporal extent through measure supervision over all valid events and full-span temporal supervision. Furthermore, we develop the REM-based Relation Inference that predicts relations with integrated temporal mass and retained temporal localization. REM improves strict mR@100 over retained-event supervision by 3.43 and 0.37 percentage points under IPS+T and VPS, respectively.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.