acceptodds
Under review as a conference paper at ICLR 2027

AirRelComp: A Diagnostic Benchmark for Video-Text Compositional Understanding in Airport Surveillance

Abstract

Understanding continuous multi-entity airport activities requires resolving temporal order, spatial configuration, and entity participation beyond merely recognizing entities and actions, supporting operational analysis and safety assessment. Broad video-text matching can conceal errors in these relations. We introduce AirRelComp, a diagnostic benchmark for fine-grained compositional understanding in airport surface videos. AirRelComp organizes Temporal, Spatial, and Entity Relations under a common event schema. We manually annotate positive video-text pairs from surveillance videos and construct four types of controlled negatives using predefined rules. Each negative changes one designated relation while preserving the remaining event structure. AirRelComp contains nearly 4K controlled comparisons and supports similarity comparison for video-text models and binary choice for large multi-modal models (LMMs). Evaluations reveal limited zero-shot transfer in video-text models and uneven discrimination across disruption types in LMMs. We therefore develop a learning framework that integrates supervision from positive–negative correspondence and relation identity with representative visual patch selection. With the same backbone, the framework raises Overall Accuracy from 63.45% under positive-only fine-tuning to 78.98%. Ablations show that the structured use of controlled negatives outperforms generic negative augmentation. LoRA fine-tuning on AirRelComp further substantially improves LMM binary-choice accuracy. These results show that AirRelComp diagnoses compositional weaknesses and provides effective supervision for improving relation discrimination.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.