HeuriGuide: Calibrating LLM Authority for Test-Time Heuristic Design
Abstract
Large language models (LLMs) have recently emerged as a powerful tool for automated heuristic design (AHD) in combinatorial optimization. Most existing LLM-based AHD algorithms employ an outer agent to optimize a strategy, an algorithmic procedure that constructs a solution, during development and then deploy it unchanged, missing the opportunity to adapt the strategy to individual test-time instances. We therefore systematically investigate placing an inner LLM agent directly into the test-time solving process. Since LLM-generated output can be unstable, the central design question is how to calibrate the inner agent's authority at test time. By imposing progressively greater procedural guidance on the inner agent, we propose three inner-agent organizations: a solution generator that directly generates a solution for each test instance, a strategy adapter that selects and adapts existing strategies, and a strategy router that only selects among strategies curated by the outer agent. We evaluate these three agent organizations on HeuriGym, a benchmark comprising nine combinatorial optimization problems across four domains. Averaged across all nine problems, QCI increases monotonically with procedural guidance, from 0.599 for the solution generator to 0.787 for the strategy adapter and 0.805 for the strategy router, compared with 0.728 for the outer-agent-only baseline. The strategy router also performs best on most individual problems. However, on E-graph Extraction, a problem for which the baseline fails entirely, success requires allowing the inner agent to generate beyond the curated strategies. Together, these results suggest using strong procedural guidance by default while relaxing procedural guidance when curated strategies prove insufficient.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.