Factorized Semantic IDs for Controllable Coarse-to-Fine Outfit Retrieval
Abstract
We study controllable outfit completion: retrieving a compatible item for a partial outfit while letting the caller specify what the returned item must be, its role (e.g. “shoes”) and, optionally, its coarse color, without scoring every catalog item at query time. Our approach gives every catalog item a short ID made of three independent fields: role and color, assigned by fixed rules, and a compatibility code that groups items a frozen, role-conditioned Transformer teacher treats as interchangeable for the same kinds of outfits. This compatibility code comes from clustering the teacher’s own item embeddings, not from a learned decoder. At query time, the caller’s role and color constraints are enforced first, by intersecting these fields the way a database query intersects indexes on several columns, before any embedding similarity is computed; only the resulting small candidate set is then reranked with the teacher’s continuous compatibility score. Under ground-truth color control, this factorized routing at L=8 retains 98.8% of a role+color oracle’s Recall@10 while scoring only 1,032 candidates (11.3×compression). Under counterfactual color control, that is, deliberately requesting a color different from the true target’s so the system must actually honor an unfamiliar constraint rather than default to what it would retrieve anyway, factorized routing at L=8 recovers 88.2% of the top-10 items the exhaustive constrained teacher would itself select, against 10.8% for a matched-budget random control; both results hold stably across 3 independently trained teacher seeds (Section 7). Unlike prior compatibility-ranking methods, which offer no native mechanism to enforce such constraints, and unlike autoregressive Semantic-ID methods, which require a decoder and a collaborative-filtering training signal that outfit-compatibility data does not provide, our factorized index adds this controllability to an existing compatibility scorer using only fixed rules and clustering, at a small fraction of the cost of scoring the full catalog.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.