SeeThat: Learning to Optimize Code through Recursive Policy–Data Co-Evolution
Abstract
Learning to optimize code requires feedback on both correctness and runtime, which program execution can provide. Beyond supplying rewards, execution also identifies valid program variants that can serve as inputs for further optimization. Recursively reusing these variants allows a policy to recursively expand its own training data, but also creates more candidate inputs than a finite training budget can afford. The resulting challenge is to identify which candidates offer useful learning opportunities before spending the budget to generate and execute their refinements. We introduce SeeThat, which trains the code-optimization policy itself to predict whether an input will yield a rollout group with variation in validity or performance outcomes. These predictions are supervised by execution feedback already collected during training, with region-specific supervision for prediction and code generation. The policy then uses its predictions to rank newly admitted variants for future training, without generating or executing additional refinements or requiring a separate value network. In this way, policy learning and recursive training-data construction co-evolve. Across systems-code and competitive-programming optimization with both dense and mixture-of-experts models, SeeThat achieves the lowest P90 and P95 runtimes among the compared training methods in all four model–benchmark settings, and reaches the lowest P75 runtimes in three settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.