GrokRSI: Few-Shot Agentic Recursive Self-Improvement Through Qualified Task–Harness Co-Evolution
Abstract
Language-model agents can improve their behavior without updating model weights by modifying the external procedures that govern reasoning, memory, and tool use. However, under few-shot supervision, Harness evolution repeatedly learns from a small and largely fixed set of situations. Simply generating more tasks is not sufficient: unreliable self-generated supervision can be recursively amplified as both the curriculum and the agent evolve. We present GrokRSI, a framework for qualified Task–Harness co-evolution. The current executable Har- ness conditions subsequent task proposals through its public behavior, while qual- ified outcomes from those tasks determine the evidence available for constructing the next Harness. A qualification and evidence-routing protocol separates verified learning signals from provisional observations, and frozen promotion prevents val- idation and held-out outcomes from feeding back into subsequent updates. Across two few-shot benchmarks, the joint procedure yields higher held-out point esti- mates than protocol-matched Harness-only evolution; a complementary standard- scale study tests the same task-side strategy beyond the few-shot regime. Together with the mechanism ablations, these results provide evidence that adapting what an agent learns from together with how it acts is a promising direction for recursive improvement under scarce supervision.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.