Think Less, Evolve Better: Unlocking Gate-Free Skill Self-Evolution in LLM Agents
Abstract
Skill self-evolution enables language model agents to refine reusable instructions from execution experience without updating model parameters. Existing methods often rely on costly validation gates to filter harmful revisions, yet removing these gates allows harmful updates to accumulate. We observe that accuracy-degrading revisions often lengthen reasoning, while fixed guidance against redundant reasoning mitigates degradation during ungated evolution. These observations suggest that reasoning efficiency can identify harmful guidance from existing execution trajectories. Motivated by this insight, we propose FreeEvo, a framework that eliminates validation gates by turning reasoning efficiency feedback into instruction selection. FreeEvo decomposes each skill into task-specific accuracy atoms and a fixed efficiency atom that discourages redundant reasoning. The optimizer uses this guidance to attribute unnecessary reasoning, useful procedures, and task errors to individual accuracy atoms. Evidence accumulated across distinct examples then determines atom activation and permanent retirement without additional candidate validation rollouts. Across three models and seven benchmarks spanning four task domains, FreeEvo matches or outperforms existing SOTA methods while reducing evolution tokens by up to 93.6%. At test time, the evolved agent improves task performance by up to 23.1 percentage points while using 6.4%–77.9% fewer output tokens. Our code will be released upon publication.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.