acceptodds
Under review as a conference paper at ICLR 2027

Ravel: Separating Romanization Collisions in Multilingual Dense Retrieval

Abstract

Romanization erases script cues that separate names, allowing a faithful variant of one entity to share its surface form with another identity. We formulate script robustness as two coupled obligations: reconnect relevant variants and keep identity-colliding passages outside the retrieved neighborhood. The Fracture–Collapse Protocol (FCP) evaluates benchmark-wide noisy Recall@100 together with candidate-micro collision intrusion over a model-independent population, and connects retrieval behavior to neighborhood geometry, source-disjoint noise, real input-method errors, generated answers, and serving cost. Ravel realizes this constraint with BGE-M3 semantics, a 12M-parameter byte/character form channel, and identity-labeled collision negatives, while retaining one vector per query and passage. Across MKQA, mMARCO, and XOR-Retrieve, it gains 6.9–8.3 Recall@100 points over BGE-M3 and 3.5–3.7 points over the same-backbone invariance control, with 0.3–0.4 points of clean regression and 32 ms p95 latency. On MKQA, Ravel reaches 69.7 noisy recall at 1.8% candidate intrusion, compared with 66.2 and 4.1% for invariance; its 6.9-point source-disjoint advantage persists, and fixed-generator wrong-entity answers fall from 4.8% to 3.3%. These results establish identity separation as a co-equal criterion for training and selecting script-robust single-vector retrievers.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.