PairDistill: Amortizing Test-Time Pairwise LLM Oracles for Code Selection via Pairwise Judgment Distillation
Abstract
Pairwise LLM oracles can select the best candidate from a set of sampled programs: ExPairT-LLM poses membership and equivalence queries to a frontier LLM at test time and attains strong pass@1, but pays the cost of dozens of frontier-model queries per task at inference time. We conduct a critical analysis of distilling such pairwise oracles into compact trained verifiers. As the analysis vehicle, we build PairDistill, which amortizes this oracle into a compact trained verifier: a frontier teacher (Qwen3-8B) annotates pairwise queries once during training; because code differentiating inputs are executable, we score teacher labels against ground-truth execution outcomes before distillation (execution-scored distillation); a 1.5B student is then trained on the teacher's judgments with swap augmentation and answer-token loss masking. At inference the verifier replaces teacher generation with a single forward pass per query (a 5.3 smaller model, no decoding), using 8 verifier queries per task. On HumanEval+MBPP and TACO with Qwen2.5-3B candidates, the 1.5B verifier improves pass@1 over random selection by (HumanEval+MBPP, borderline), (TACO-fn) and (TACO-stdin) points, recovering 13–15% of the oracle–random gap (23% with a 3B verifier) at 14% of the teacher's per-task query count (8 vs. 56 queries; under 8% in token terms, since each student query is a single 1.5B forward pass over 4.1K tokens while each teacher query requires 8B generation of a differentiating input and verdict), and exceeds the teacher's own test-time selection quality (64.8% teacher vs. 67.7% student on the full HumanEval+MBPP evaluation; paired studentteacher, CI ). Gains over random are statistically significant on both TACO benchmarks (paired bootstrap, 95% CI). Ablations show that distillation volume matters far less than dev metrics suggest (25% of data recovers most selection quality), that verifier scale exhibits a threshold effect (0.5B verifiers match random selection despite higher dev accuracy than 1.5B), that task-level leakage between teacher annotation and evaluation inflates apparent selection gains by several points, and — via a de-confound — that the execution-filtering step itself is not responsible for end-to-end gains: under a unified training schedule, the class prior (not label filtering) is the largest axis of variation, and the headline configuration (filtered + balanced) is the only cell whose selection gain over random is individually significant, so the distillation recipe with a balanced query prior carries the result.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.