acceptodds
Under review as a conference paper at ICLR 2027

Training-Free Test-Time Retrieval Augmentation for Visual Relationship Detection

Abstract

Visual relationship detectors learn relational knowledge from annotated examples, yet these examples are typically discarded as explicit evidence after training. We revisit their value at inference time and introduce Training-Free Test-Time Retrieval Augmentation, i.e., TF-TRA, a recipe that complements frozen detectors with non-parametric relation memory constructed solely from existing training data. TF-TRA organizes training instances into a relation memory that associates semantic and visual representations with relation annotations while maintaining broad configuration coverage. For each query, it retrieves exemplars through joint semantic and visual matching, with diversity-aware reranking encouraging complementary relation evidence. The retrieved annotations are aggregated into query-specific support and calibrated with the memory prior to refine the frozen detector's predictions. Extensive experiments on three distinct VRD tasks, i.e., human-object interaction (HOI) detection, scene graph generation (SGG), and nonverbal interaction (NVI) detection validate the effectiveness of TF-TRA. It improves overall mAP under both fully supervised and four zero-shot settings on HICO-DET, increases mean Recall and F-score on Visual Genome for SGG, and yields substantial gains in NVI detection. These results show that training examples remain valuable beyond parameter learning, providing explicit relational evidence that strengthens existing detectors at test time. Our soure code will be released.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.