AUDITING TRAINING–EVALUATION OVERLAP, NEGATIVE SAMPLING, AND RERANKING IN MULTILINGUAL PATENT RETRIEVAL
Abstract
Cross-language retrieval systems are usually evaluated against a fixed benchmark, and the resulting numbers are read as evidence about representation quality. We show that in multilingual patent retrieval this reading is fragile: three ordinary design decisions—how negative examples are mined, which overlap definition an audit adopts, and whether a first-stage candidate space is validated—each change what a headline conclusion appears to be. We present CLARITY (Cross-Language Audited Retrieval with Integrity and Transparency), an audited data and evaluation workflow for multilingual patent counterpart retrieval spanning Chinese–English, Japanese–English, and Korean–English (2.692M bilingual records; 23,622 queries; 3,000–4,854 target-language documents per language pair). We contribute (i) a leakage audit that separates two definitions of training–benchmark exposure and shows they differ by 71–837× across language pairs: 70–76% of evaluation queries appear in the reranker training corpus by querytext match, but only 0.08–1.06% by exact query/positive pair; a control sample of non-benchmark titles matches at 79.7%, and the benchmark loose-exposure rates are 3.5–9.9 pp below this baseline, indicating that query-text overlap reflects corpus vocabulary intersection rather than answer memorization; (ii) a reranking evaluation corrected under the strict definition, where the reranker still improves NDCG@10 by +0.0249, +0.1446, and +0.1397 on the three pairs; and (iii) a repaired hard-negative pipeline: an earlier miner placed the query itself among the negatives in 2.51–9.14% of rows, and after excluding the query, same-component rows, and duplicates, hard negatives beat random negatives on all three language pairs. Baseline measurements for BM25, LaBSE, mE5, BGE-M3, and Qwen3-Embedding show that point-estimate rankings vary across language pairs, and direction asymmetry reaches 0.2619 NDCG@10 for mE5 on Korean–English.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.