CO-Agent: Cost-Aware Combinatorial Optimization via Verifier-Guided Distillation and Skill Evolution
Abstract
Language-model agents can generate optimization programs, invoke mathematical solvers, and repair execution failures, but strong performance often comes at the cost of large models and repeated inference. We investigate how small language models can coordinate these capabilities to solve combinatorial optimization (CO) problems under limited computational budgets. We introduce CO-Agent, a framework that combines a compact planner and generator with a solver harness and an evolving repository of reusable optimization skills. The framework integrates three complementary learning mechanisms. First, cost-aware planner optimization learns to allocate generation, solver, and refinement resources from the outcomes of complete agent trajectories. Second, verifier-guided distillation trains the generator using re-executed and verified teacher corrections on states visited by the student. Third, repository-level skill evolution evaluates candidate skills by their marginal contribution to deployed-agent utility under natural retrieval, accounting for retrieval interference and negative transfer. A shared execution–verification loop grounds all three mechanisms in feasibility, solution quality, and computational cost. Experiments on CO-Bench, FrontierCO, and HeuriGym demonstrate that CO-Agent achieves competitive solution quality with lower inference cost, while ablation and continual-evolution studies show the complementary benefits of the three mechanisms and sustained improvements in skill reuse.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.