Evaluation Is an Intervention: How Evaluation Protocols Shape Measured LLM Program-Repair Performance
Abstract
Benchmark scores are often used to measure a large language model’s capability. We argue that the resulting number can be affected by the evaluation procedure as it determines the opportunities and information available to the model. In this paper, we investigate how evaluation protocol choice changes measured program-repair performance. The effects of repeated sampling without feedback, refinement with generic feedback and refinement with structured feedback are compared under matched maximum attempt budgets. Experiments on the synthetic BenchCraft benchmark show substantial protocol effects across three models. The estimated improvement in visible-test success with structured feedback over repeated sampling without feedback is about 11 percentage points for Claude Haiku, 23 for GPT-5.6-terra and 31 for Claude Sonnet 5. Each protocol allowed up to three attempts per task. The results support a substantial improvement, although uncertainty about its size exceeds the pre-specified precision target. In the real-world BugsInPy benchmark experiment, refinement with structured feedback shows no clear advantage over repeated sampling without feedback at a comparable attempt budget. In this experiment, additional attempts produced successful repairs even without feedback. Success rates also varied substantially across defects. The structured-feedback advantage observed on BenchCraft therefore does not clearly extend to BugsInPy under the execution conditions studied. This leads us to conclude that protocol effects depend on the tasks and execution conditions under which they are measured. Evaluation protocol should therefore be treated as a consequential experimental factor. An LLM benchmark score is under-specified without the protocol and execution regime that produced it.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.