Beyond Single Positives: An Empirical Audit of Solution-Set Adaptation for Code Retrieval
Abstract
Competitive-programming problems provide multiple accepted solutions and rejected submissions for the same task. We audit whether adapting code encoders with this solution-set supervision improves retrieval beyond a single-positive baseline and the original pretrained encoder without adaptation (the frozen encoder). We evaluate MPR-Hybrid (Multi-Positive Retrieval Hybrid), a training system that trains language-specific LoRA adapters with a Multi-Positive Pooled Contrastive (MPPC) loss, linking each problem query to multiple accepted solutions and sampled rejected submissions selected from candidate pools using submission verdicts and embedding similarity. We compare it with a single-positive in-batch contrastive-learning system trained with the InfoNCE loss, using the other examples in each batch as negatives on six encoder backbones, trained with CodeContests supervision and evaluated on its separate test set, with zero-shot CodeSearchNet transfer, a fixed pool of code candidates selected by keyword matching (BM25), including lexically similar incorrect distractors, to test retrieval when keyword overlap is misleading, and CodeXGLUE AdvTest. We also evaluate the original pretrained encoders without fine-tuning as frozen references on the available CodeContests and CodeSearchNet evaluations, to assess whether adaptation improves retrieval beyond the pretrained representation. The observed benefits vary across encoder backbones and evaluation protocols. On the fixed CodeContests official-test protocol for unseen problems, MPR-Hybrid has higher average scores for some backbones, but no result meets our two-metric confirmation rule for a MPR-Hybrid win. In zero-shot CodeSearchNet Java transfer, where candidate code almost never contains the query text, MPR-Hybrid beats the single-positive baseline on three of six backbones, but none improves on its frozen encoder under the same two-metric rule. On Python, every protocol with a MPR-Hybrid win also contains the query text verbatim in candidate documentation. Removing documentation and comments leaves two confirmed MPR-Hybrid gains over the single-positive baseline: i) one other confirmed win becomes a loss, one formerly inconclusive comparison becomes a loss, and ii) two confirmed wins become inconclusive. The audit shows how several apparent gains change interpretation under frozen-encoder comparisons and documentation removal, while identifying settings where gains over the single-positive baseline remain. The comparison evaluates complete adaptation systems and therefore does not independently isolate the effects of the loss function, candidate sampling, or which encoder layers receive the LoRA adapters.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.