Once Is Enough: Interaction Timing in Dual-Core Vision– Language Detection
Abstract
Vision–language detectors typically intertwine two architectural decisions: when visual and textual representations interact and how often this interaction is repeated. We study these variables independently through a controlled dual-core detector that first constructs semantic text cores and reference-aware visual cores, and then varies only the placement and frequency of cross-modal exchange. Holding representation size, decoder depth, initialization, and optimization fixed, a single interaction immediately after core construction consistently outperforms early fusion, middle or last-layer interaction, and four interactions distributed across the decoder. Importantly, this result persists when core capacity is increased: with 128 text cores, the single-interaction model matches the dense detector on COCO, whereas repeated interaction reduces AP by 1.23 points. The same ordering transfers to a larger Grounding DINO backbone, where repeated interaction incurs a 1.27-point drop, and remains compatible with COCO-to-LVIS zero-shot transfer. Mechanism ablations show that semantic text resampling and reference-aware visual refinement are both necessary, while preserving the original text tokens at the matching head protects the pretrained vision–language alignment interface. These findings establish interaction timing as a first-order design variable independent of interaction frequency and support a phase-separated decoder paradigm: construct modality-specific representations first, exchange information once, and continue visually grounded refinement afterward.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.