From Agent Harnesses to COP Experts: Automated Skill Evolution for Cross-Problem Combinatorial Optimization
Abstract
General-purpose agents built on agent loops and execution harnesses are becoming the default interface for complex tasks, yet they remain generalists. In combinatorial optimization (CO), existing agentic systems typically rely on manually engineered workflows tailored to each problem, which limits reuse and scalability. This gap raises a central question: can a general agent automatically acquire skills that turn it into an expert solver across heterogeneous COPs? We propose an automated skill-evolution method that learns reusable behavioral guidance from solver-development experience, allowing the agent to discover and refine its own skills from indirect feedback rather than relying on handcrafted workflows. Our key result is that a skill evolved solely from TSP development evidence transfers to CVRP, JSSP, and 3D-BPP. Across these three unseen COPs, it substantially outperforms no skill, the initial skill, and state-of-the-art skill-evolution methods such as SkillOpt, which optimizes agent skills as a trainable external state through validation-gated text-space edits. These findings suggest a promising route to general agents that can automatically become experts across diverse combinatorial optimization problems.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.