BreadBowl-Embed: Encode Once, Retrieve and Rerank
Abstract
Modern retrieval systems typically separate candidate generation from relevance refinement: bi-encoders provide efficient retrieval, while cross-encoders recover richer query-document interactions at substantially higher inference cost. We ask whether these two stages can instead operate on the same representation. We introduce BreadBowl-Embed, a retrieval model built around a Slot Encoder, which represents each query and document as a compact set of routing-value slots. The routing vectors provide an indexable representation for first-stage retrieval, while the same routing vectors subsequently define attention over a candidate's stored value vectors to produce query-dependent readouts for reranking. Retrieval and reranking therefore share not only an encoder, but the representation that determines both which documents are retrieved and how their contents are compared. Documents are encoded once and reused across both stages, so reranking requires neither a separate cross-encoder nor an additional backbone pass over candidate text. Across 15 BEIR tasks, value-based reranking raises macro nDCG@10 from 48.47 to 51.54 (scores scaled by 100) over routing alone on the same retrieved candidates. These results suggest that indexed retrieval and query-dependent relevance refinement can be supported by a single precomputed representation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.