acceptodds
Under review as a conference paper at ICLR 2027

Graded Execution Feedback for Co-Evolving Code and Test Generators

Abstract

Execution feedback for code generation is often represented as a binary pass/fail signal. This is natural for unit-test style correctness checking, but it discards useful information in tasks where multiple outputs are feasible and differ in solution quality. We study code generation for graded programming tasks, including combinatorial optimization and sequential decision-making, where programs are evaluated by normalized objective or rollout scores. We introduce Continuous Co-Evolution (CCE), a framework that trains a shared language model under both solution-program and test-generation prompts. Candidate programs are rewarded by their continuous performance on reference tests, while generated tests are rewarded when the program rankings they induce agree with rankings from reference tests. This ranking-based objective enables generated tests to provide informative feedback while preserving graded differences among candidate programs. Experiments on Traveling Salesman Problem (TSP), Set Cover, Parallel Machine Scheduling (PMS), and Game 2048 show that continuous test-feedback co-evolution improves over coder-only training and binary-feedback co-evolution baselines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.