UCE: Uncertainty-Calibrated Evolution for Automatic Heuristic Design
Abstract
In iterative LLM-based heuristic design, evaluation scores serve not only to rank candidate implementations, but also to shape the search directions that an LLM continues to explore or abandons. However, a raw score evaluates only one implementation in isolation and can be misleading when interpreted without considering the performance of other candidates generated from the same parent. In particular, an isolated low score may cause the LLM to prematurely discard a potentially productive search direction. We introduce Uncertainty Calibrated Evolution (UCE), a stage-aware scoring paradigm that calibrates candidate evaluations according to the collective performance within each fixed-parent stage. Specifically, UCE contextualizes each candidate's score using the mean and dispersion of its stage-level candidate pool, transforming an isolated evaluation into a relative signal for subsequent selection and feedback. This design reduces the disproportionate influence of low or otherwise atypical scores on the LLM's iterative decisions, without modifying the program generator or requiring additional model training. We evaluate UCE across diverse operations-research datasets and LLM-AHD baselines. UCE improves overall performance and reduces premature acceptance and rejection in counterfactual replay, suggesting that stage-level score calibration can provide more reliable guidance for iterative heuristic search.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.