RADAR: RUNTIME-AWARE DIAGNOSIS AND RACING FOR AGENTIC SOLVER DESIGN
Abstract
LLM-based automated algorithm design often uses a solver's final scalar score both to rank candidates and to guide code revisions. Yet similar scores can conceal different runtime behavior that calls for different edits, a problem we term behavioral aliasing. In controlled studies of MCTS-AHD and ReEvo, we find no detectable performance gain from exposing exact score magnitudes or using run-specific reflections. We introduce RADAR (Runtime-Aware Diagnosis and Racing), an agentic framework that separates the two roles: execution traces guide revisions, and scores allocate evaluation effort. A trusted evaluator records when valid solutions improve, how long the solver runs without further improvement, and why outputs are rejected; a diagnostic agent turns this trace and the search history into one measured bottleneck and one code edit for a writer agent. Racing first runs each new program on a single training sample and stops weak programs before they are tested on the remaining samples, and Branch-Adaptive Reflection changes the edit strategy when a search branch stalls. Across seven benchmarks, RADAR with Qwen3-32B achieves the best mean score among the evaluated search baselines on every task; on CVRP-200 its gap to the reference is 7.20% versus 11.76% for the strongest baseline. With GPT-5.4-mini, RADAR leads or ties on five of the seven tasks. Attaching our racing and diagnosis mechanisms to MCTS-AHD and ReEvo improves both controllers on all three tested tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.