acceptodds
Under review as a conference paper at ICLR 2027

EvoORSkill: When Optimization Skills Evolve, Who Checks The Math?

Abstract

Large language models can translate operational requirements into optimization models and solver code, but code that runs successfully may still solve the wrong decision problem. Reusing generated procedures reduces repeated modeling work, yet adapting them to new requirements can change which decisions the model allows without causing execution errors. We introduce EvoORSkill, an agent framework that evolves reusable operations research skills together with their revision strategies, without retraining the underlying LLM. Each skill combines an executable modeling procedure with explicit mathematical relationships, applicability conditions, and version information. EvoORSkill organizes these skills in an evolving skill hypergraph (ESH), which connects the procedures assembled for a task through the mathematical objects they share. When the agent proposes a revision in response to task feedback, ESH traces the affected relationships in both the accepted and candidate hypergraphs. Specification, requirement, and retention checks then determine whether the revised skills can enter the reusable library. Historical replay compares revision strategies using records of past attempts with compatible versions, avoiding repeated candidate generation and evaluation. The selected strategy guides further skill revisions, whose outcomes extend the history used to refine the strategy. Across 518 problems from six datasets, EvoORSkill achieves 81.5% average solution accuracy, exceeding the strongest evaluated baseline by 5.0 percentage points. Removing cross-skill dependency checking or specification and requirement checks reduces accuracy by 6.6 and 13.3 percentage points, respectively. Historical replay reduces library-building tokens by 14.8% compared with generating and evaluating candidates afresh. With its evolved library frozen, EvoORSkill achieves 57.0% accuracy on 100 OptMATH problems, outperforming the same agent without a library by 25 percentage points while using 28.3% fewer online tokens.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.