The Harness Learns to Lie While Evolving: Measuring and Preventing False Success in Harness Self-Evolution
Abstract
Harness self-evolution repeatedly rewrites the prompts, tools, skills, middleware, and long-term memory around a frozen model over multiple iterations. Although existing self-evolving harness methods and benchmarks can improve usability, self-evolving harness methods cannot distinguish between "tasks that are actually completed" and "fabricated deliverables that pass the same check," and benchmarks cannot measure whether false successes increase with each iteration. We introduce two frameworks, EvoTrap-Bench and Verifiably Honest Harness Evolution (VHE). EvoTrap-Bench measures how false success evolves over the iterations of an arbitrary self-evolving harness. EvoTrap-Bench uses provably impossible tasks, on which any claimed success is false, and keeps the ground truth hidden from the loop. Verifiably Honest Harness Evolution (VHE) is a self-evolving method in which each step is determined by execution rather than by the model's assessment of its own performance. VHE derives its rewards from trajectories rather than reports, and its experience is derived mechanically without asking the model why a particular trial failed. Notably, we found through EvoTrap-Bench that false successes increase with each round of harness self-evolution. This phenomenon occurs in most existing harness self-evolution methods. In the others, false success stays high and shows no downward trend. VHE overcomes the false success phenomenon, keeping false success at an extremely low level in EvoTrap-Bench. Furthermore, VHE performs even better on two benchmarks where it has never evolved before. These results show that measuring and overcoming false success enables scalable, efficient, and self-improving LLM systems.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.