acceptodds
Under review as a conference paper at ICLR 2027

RepoEquiv: Benchmarking Large Language Models' Understanding for Repository-Level Semantic Equivalence

Abstract

Large language models (LLMs) for code and agentic coding systems are rapidly improving at complex repository-level programming tasks. However, training and evaluating these systems largely rely on execution-based signals from unit tests, which are costly to obtain and difficult to scale across web-scale repositories and heterogeneous execution environments. We advocate a complementary approach: learning an LLM-based verifier for semantic equivalence that determines whether two pieces of code produce identical outputs for all valid inputs, thereby reducing the need for explicit execution. We introduce RepoEquiv, a benchmark and dataset suite for repository-level semantic equivalence constructed from repository-level code infilling datasets spanning both general-purpose and scientific programming domains. Unlike prior benchmarks focused on standalone functions with synthetic perturbations, RepoEquiv features repository-level context and dependencies, as well as realistic candidate code snippets generated by diverse coding LLMs. Evaluating 5 representative LLMs on RepoEquiv shows that semantic equivalence verification in repository contexts remains challenging for off-the-shelf models. Fine-tuning on RepoEquiv substantially improves verifier performance on general coding repositories, demonstrating the feasibility of learned semantic verifiers for scalable, execution-free evaluation, and yields uniform gains across all four scientific coding benchmarks, indicating promising out-of-domain generalization for learned verifiers. We further conduct calibration analysis and find that fine-tuning reduces calibration error, and we measure the resulting cost-accuracy trade-off directly: routing low-confidence cases to execution reaches 95% accuracy on the scientific benchmarks while executing 58% of cases, against 80% under confidence-agnostic routing. RepoEquiv establishes a foundation for developing learned semantic verifiers that can significantly reduce the cost of large-scale evaluation and learning for repository-level code generation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.