Before the Breakthrough: Bottleneck Crossings as Precursors in Long-Horizon Agent RL
Abstract
When validation success plateaus, it is difficult to tell whether a long-horizon LLM agent is still making hidden progress toward a breakthrough or is genuinely stalled. We introduce the Crack-Front Monitor, which tracks clusters of verifier-confirmed bottleneck crossings as candidate precursors of breakthroughs, without additional model inference or environment interaction. We characterized the signal on 15 discovery runs of 8–14-turn executable tool tasks. We then froze the monitor, its thresholds, and its predictions before any crack or breakthrough event in a ten-run prospective cohort spanning independent seeds and a second task distribution in the same environment family. In the completed cohort, crack-fronts preceded all nine observed sustained breakthroughs across both task distributions and both training arms, with no false alarms. The one run that never fronted never broke through. In an exploratory comparison with thresholds frozen before application, the monitor's median online warning lead was 20 steps versus 5 for a near-miss baseline; a post-hoc crossing-rate rule achieved a similar lead, consistent with the signal residing in the crossing events themselves. The frozen timing window and restart budgets transferred poorly: only four of nine lead times fell within the preregistered 30–60-step window, and the restart budgets would have stopped three runs that later broke through. These results provide initial prospective evidence that bottleneck crossings can reveal progress before validation improves, but they do not yet support reliable breakthrough timing or stopping rules. Code and data are available at https://anonymous.4open.science/r/crack-front-monitor-8997.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.