acceptodds
Under review as a conference paper at ICLR 2027

Learning Box Relations as Distances for Multi-modal Representation

Abstract

Multimodal representation learning seeks to capture semantic correspondence beyond pairwise similarity, including intersection, inclusion, and separation. Point embeddings model correspondence through proximity, whereas box embeddings can encode these relations and compose them into more complex structures. Existing box methods, however, primarily rely on volume-based objectives whose coordinatewise products attenuate learning signals as dimensionality or box separation increases. We introduce a new volume-free, distance-based objective that learns box-to-box relations through intersection-aware and directed inclusion losses. Our objective avoids this attenuation and theoretically guarantees box intersection and inclusion. Experiments show that our method reliably learns joint, disjoint, and inclusion relations, achieves the best cross-modal retrieval performance on COCO-1K, COCO-5K, and RedCaps, and shows the closest alignment with human semantic similarity on CxC. On MovieLens-1M, our learned relations compose to answer set-theoretic queries, including combinations not observed during learning. Together, these results establish distance-based box relation learning as a reliable foundation for multimodal correspondence and set-theoretic reasoning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.