Dual-Cue Relation Modeling with Synthetic Cue Augmentation for Unbiased Scene Graph Generation
Abstract
Scene Graph Generation (SGG) represents visual scenes as structured subject-predicate-object triplets, yet relation prediction remains challenging due to semantic ambiguity, limited generalization to unseen triplet compositions, and long-tailed predicate distributions. To address these challenges of relational evidence in SGG, we propose a framework of Dual-Cue Relation Modeling (DCRM), which models visual-spatial and semantic-context evidence in separate relation branches. To enhance semantic modeling and compositional generalization, we further incorporate pretrained encoders into this dual-cue framework, leveraging its cross-modal image representations and encoded predicate descriptions to provide semantic guidance and fine-grained predicate prototypes. An adaptive router dynamically coordinates the predictions of the two branches, enabling effective integration of the resulting heterogeneous cues. To address the long-tailed distribution, we further introduce Synthetic Cue Augmentation (SCA), which performs semantically guided local variation transfer from data-rich to underrepresented predicates in both cue spaces. Extensive experiments on VG150, OIV6, and GQA200 demonstrate competitive overall and balanced performance of the proposed framework, with improved long-tailed predicate recognition and zero-shot generalization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.