InverseHarness: Internalizing Agentic Reasoning through Reinforcement Learning
Abstract
Reinforcement learning (RL) improves the reasoning capability of large language models (LLMs) by searching and then memorizing successful trajectories, but progress can stall when effective strategies are rarely discovered due to the difficulty of problems through native sampling. Agentic harnesses can elicit stronger reasoning through external orchestration, yet their executions are not directly usable as standalone training trajectories. We introduce HAIR (Harness-Assisted Internalization via Reinforcement Learning), which combines an agentic harness with InverseHarness to reconstruct successful executions into self-contained trajectories while preserving the current policy’s generated reasoning. Interleaved with standard RL, this process targets problems where native rollouts fail and regenerates training trajectories from the evolving policy. On mathematical reasoning with Intern-S2-Preview at 35B and 397B, HAIR outperforms GRPO across most evaluated benchmarks under matched backbones, training corpora, and optimizer steps. At 397B, pass@1 increases from 55.5 to 68.8 on IMO-Proof-Bench and from 29.2 to 38.4 on PhD-Paper. Across rounds, harnessed performance improves with the policy, while training on successful executions improves standalone reasoning, supporting a closed loop of self-improvement. Further analysis shows that generating trajectories from the current policy matters beyond correctness alone. These findings establish harness-assisted training as a promising route to self-improvement beyond native sampling.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.