acceptodds
Under review as a conference paper at ICLR 2027

LO-Analog: Leakage-Aware Retrieval of Transferable Reasoning Skills for Linguistics Olympiad Problems

Abstract

Retrieving a solved example can help a model reason about a new Linguistics Olympiad problem-but conventional similarity search systematically returns sibling subproblems and shared contexts that expose the target's reasoning path, contaminating evaluation. We introduce LO-Analog, an answer-free benchmark and constrained retrieval framework that exposes and eliminates this leakage. The benchmark indexes 160 Linguini tasks grouped into 73 original-problem and context clusters, represents candidates using answer-free structural signatures, and evaluates relevance, near-leakage, and result-set diversity jointly under a leave-one-task-out protocol. Eight retrieval configurations-spanning BM25, TF-IDF, hybrid scoring, leak filtering, skill supervision, and MMR diversification-are compared on five metrics. The contamination is severe: same-original or same-context items occupy 29.5% of top-five positions in raw rankings. Cluster-level exclusion eliminates every near-leak at both @5 and @10. Under answer-free skill supervision, filtered BM25 improves nDCG@5 from 0.494 to 0.848-a +0.354 gain confirmed by 2,000-resample paired bootstrap (95% interval [0.328, 0.382]). A diversity-aware skill-and-safety ranker simultaneously attains zero near-leak@10, full cluster diversity@10, and nDCG@5 within 0.02 of the oracle upper bound. Ablations show that each filter stage is necessary: context-only filtering still leaks 8.3% at @5, and removing skill supervision collapses nDCG@5 by 0.19. A preregistered, blinded evaluation of 388 unique top-ranked pairs (two annotators plus adjudication) confirms that leakage-aware retrieval improves human-judged analogy helpfulness while sharply reducing human-identified near-leaks; the conclusion survives removal of all public metadata features from the ranking signal. LO-Analog establishes analogy retrieval as a joint relevance-and-safety problem and demonstrates that a substantial fraction of conventional retrieval gains reflects leaked puzzle context rather than genuine structural transfer.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.