Unified Semantic Product Quantization with Category-Aware Implicit Codebooks for Cross-Modal Retrieval
Abstract
Cross-modal retrieval (CMR) requires compact and discriminative representations for efficient large-scale search. Although product quantization (PQ) is effective for compact representation learning, existing PQ-based CMR methods typically treat codebooks primarily as feature-approximation prototypes, limiting the direct incorporation of semantic supervision into codebook learning. To address this issue, we propose P2PQ, a unified semantic PQ framework with category-aware implicit codebooks for CMR. Instead of independently parameterizing quantization codebooks and semantic predictors, P2PQ shares a single matrix between these two roles, so that the codebook cardinality is induced by the dimensionality of the semantic prediction space rather than specified independently. This dual-role parameterization allows category-level supervision to directly shape the quantization prototypes without imposing label-constrained codeword assignments, thereby intrinsically coupling discrete quantization with semantic discrimination. Furthermore, P2PQ learns a shared latent semantic space for heterogeneous modality alignment and preserves geometric consistency between continuous and quantized representations through an implicit label-guided affinity formulation without explicit graph construction. Extensive experiments on three benchmark datasets demonstrate that P2PQ consistently outperforms representative shallow CMR methods under various code lengths, validating its effectiveness and scalability.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.