RepRSI: Can Internal Learning Progress Guide Recursive Self-Improvement?
Abstract
Recursive self-improvement relies on generated training experiences that support further learning. Immediate task performance can overlook internal changes that predict later capability acquisition. We study semantic cohesion: agreement between hidden representations of equivalent intermediate computations across contexts, relative to representations of different computations. In a controlled compositional task, semantic patching links these representations to reusable computation, and curriculum interventions show that cross-context exposure shapes cohesion and subsequent learning. Adding cohesion progress to behavioral predictors reduces mean absolute error in predicting subsequent accuracy from 19.3 to 10.1 percentage points. We introduce RepRSI, a Representation-guided Recursive Self-Improvement framework that uses curriculum-induced semantic cohesion progress to reward a curriculum-generating teacher and select student branches for continuation, while training the student with verifiable task rewards. At matched total training compute, it improves pass@32 over a teacher rewarded by target-performance gains by 13.8, 3.1, and 5.7 percentage points on Manufactoria–HAS, MATH, and HARP, respectively. Its curricula also improve fresh students, and the recursive system reaches the comparison method's final greedy performance using 65–75% of the common full budget. Cohesion progress thus provides useful feedback for recursive self-improvement beyond immediate task performance. The code is available at https://anonymous.4open.science/r/RERSI-BB36/
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.