Beyond Execution Agreement with Few Tests: Relational Transmission for Code Generation
Abstract
Large language models (LLMs) have become increasingly capable code generators, yet reliable assessment of generated code remains essential to improving inference and learning. Existing approaches draw on model likelihood for independent assessment and execution agreement for relational assessment on available tests. However, inadequate tests can cause correct and incorrect candidates to exhibit identical observed behavior, leaving execution agreement alone unable to rank the correct candidates higher. We therefore seek finer relational evidence in code differences to move beyond execution agreement with few tests. To this end, we propose Trace-Conditioned Relational Transmission (TCRT). For each candidate pair, TCRT uses each response's reasoning trace or other non-code text in turn to score changed tokens on both sides of the code diff. For each response's context, the mean token log probability defines a directed local log potential. These potentials, with unary model likelihood and execution agreement, define a relational kernel. TCRT applies finite-step relational transmission through this kernel to redistribute evidence and rank the candidates. Theoretically, our unified analysis explains why execution agreement can leave likelihood-based misrankings unresolved. We prove that TCRT can reverse these misrankings when informative contexts receive sufficient support. Empirically, TCRT consistently outperforms strong baselines across LLMs and code-generation benchmarks, with particularly strong gains when few tests are available.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.