Learning Box Relations as Distances for Multi-modal Representation
Abstract
Multimodal representation learning seeks to capture semantic correspondence beyond pairwise similarity, including intersection, inclusion, and separation. Point embeddings model correspondence through proximity, whereas box embeddings can encode these relations and compose them into more complex structures. Existing box methods, however, primarily rely on volume-based objectives whose coordinatewise products attenuate learning signals as dimensionality or box separation increases. We introduce a new volume-free, distance-based objective that learns box-to-box relations through intersection-aware and directed inclusion losses. Our objective avoids this attenuation and theoretically guarantees box intersection and inclusion. Experiments show that our method reliably learns joint, disjoint, and inclusion relations, achieves the best cross-modal retrieval performance on COCO-1K, COCO-5K, and RedCaps, and shows the closest alignment with human semantic similarity on CxC. On MovieLens-1M, our learned relations compose to answer set-theoretic queries, including combinations not observed during learning. Together, these results establish distance-based box relation learning as a reliable foundation for multimodal correspondence and set-theoretic reasoning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.