Learning Instance-Adaptive Solver Configuration from Execution Feedback
Abstract
Mixed-integer programming solvers expose many parameters whose best settings depend on the structure and search dynamics of each instance. We study instance-adaptive solver configuration: a policy observes diagnostics from a short solver run, proposes sparse parameter edits, and refines its configuration using execution feedback. We formulate this process as a budgeted sequential decision problem and train a language-model policy with an incumbent-aware reinforcement learning objective. The evaluation compares against default SCIP, Random Search, SMAC3, and a prompt-only Base Qwen agent under a matched four-turn interaction protocol. On 53 held-out instances at each of four preregistered 100, 200, 300, and 400 second solver limits, the trained policy has more paired wins than losses in all 16 method–budget comparisons. Even against the strongest baseline at each limit, its net win margins are , , , and , respectively. Across six frozen prompt-only GPT and Claude agents, it wins 23 of 24 model–budget pairings; the only negative margin is against Claude Opus 5 at 400 seconds. Removing candidate feedback and incumbent reuse yields 37 wins for the full policy at both 100 and 400 seconds, with six and four losses, respectively. Compared with the no-regression-penalty training-and-selection pipeline, the full policy has net margins of and , but does not consistently reduce observed episode-best-to-last regressions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.