acceptodds
Under review as a conference paper at ICLR 2027

Logic Bus: When Operators Matter, Not Just Vectors

Abstract

Two-tower models are trained to align the vectors their encoders produce, while the encoders themselves—the operators—are left unused as a choice. Beyond the default of passing each side through its own tower, a trained system holds two alternatives, at no training cost, that beat it on part of the input space: the second tower applied to both sides, and the pretrained backbone applied to both. We call a system that chooses among these pairings a Logic Bus. Contrastive training fits the two towers on different marginals of the training pairs, so they differ in training coverage. We show that, when fitting follows data density, each tower's bias grows as its coverage thins. It follows that no fixed pairing is least biased everywhere, provided each tower covers some inputs more densely than the other, and the backbone is better than the towers on some inputs and worse on others; when either condition fails, the choice shrinks or disappears. In an image–text catalogue, reducing only training coverage moves the trained towers from 20 category-level MRR points above the backbone on the tail to 52 below it. In text retrieval the best pairing changes from dataset to dataset: the three fixed pairings score within about one point of each other in aggregate, while choosing the pairing per dataset gains 8.4 nDCG@10 over the best of them.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.