Understanding Framework Advantages in LLM-Driven Program Evolution
Abstract
LLM-driven program evolution frameworks are attracting growing attention, yet how their task-specific advantages develop and vary with models and search budgets remains insufficiently understood. We systematically compare four frameworks across five tasks in 1,560 runs under a common 50M-token request-start threshold. Pairwise and within-pool nearest-neighbor distances characterize the dispersion and local structure of task-specific execution profiles, while size-matched random subsets provide a reference for high-quality candidate concentration. We combine these measures with reconstructed budget trajectories and program-level controls to examine search improvements. Framework advantages are task-specific and broadly consistent across model configurations, although some pairings reverse rankings. Selecting a framework for each configuration does not consistently outperform a task-wide choice on held-out search repeats with the same inputs. Lower response dispersion alone does not consistently identify the best-performing framework. Reconstructed trajectories show that additional budget narrows quality gaps on some tasks and increases attainment of common quality targets. These findings support framework and model selection based jointly on task fit and available budget, rather than configuration-level rankings or response dispersion alone.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.