acceptodds
Under review as a conference paper at ICLR 2027

InScribe: Internalizing Scaffolds through Agentic Reinforcement Learning

Abstract

Reinforcement learning with verifiable rewards trains coding and terminal agents from expensive rollouts, yet a task on which every rollout fails yields no group-relative solver signal and is discarded. Such failures need not mark the model's limits: with the weights frozen, different harnesses—the programs that organize the agent's planning, tool use, and verification—recover a substantial share of them, and a short teacher diagnosis of the failed attempts recovers more of the remainder. Can these gains outlast the support that made them possible? We treat harnesses and hints as scaffolds: supports put up during training and taken down once the policy has learned what they made reachable. InScribe is a two-stage on-policy RL recipe for internalizing the benefits of both. In Stage I the policy writes its own harness before each task, so that choosing how to work becomes one of its actions; each harness is used once and discarded. In Stage II, training continues on the tasks that a Stage-I probe leaves unsolved under both the stock harness and sampled rewrites, keeping harness generation and adding a verifier-informed teacher diagnosis as context rather than as a prediction target. Main evaluations run under fixed harnesses, with no generated harness and no teacher hint. A 9B student trained this way gains 10.1 points on Terminal-Bench 2.0 and 11.0 on SWE-bench Verified over its base model; the gains hold under harnesses it never trained with and transfer to function-calling and interactive-agent benchmarks it never saw. What the scaffolds taught lives in the weights.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.