ProContract: Decoupling Responsibility and Progress from the Harness in Long-Horizon Agents
Abstract
Long-horizon agents are increasingly used to turn user-assigned responsibility into artifacts over hours of execution. Yet even with capable underlying models, their reliability is undermined by two persistent failure modes: responsibility dilution and progress misjudgment. Both arise because long-running agents continuously rewrite their execution state, such as delegating subtasks, compacting context, restarting, and even modifying their harness, while the responsibility and its progress remain entangled with that state. We introduce PROCONTRACT, whose progress-contract kernel independently keeps responsibilities in a Contract outside the agent and tracks progress against that Contract through authorized judgments. The agent’s harness is deliberately disentangled from the kernel, so changes in execution state cannot alter assigned responsibilities or interfere with progress judgment. This separation also makes self-improvement an ordinary responsibility: the agent can rewrite its own harness while the kernel preserves the task it was asked to carry out. With GPT-5.6 Terra, PROCONTRACT raises the mean pass rate from 72.35% to 77.24% over the same model’s public mini-SWE-agent run. The gain comes from the bottom: the number of tasks below 10% halves. A version written by the agent itself produced a successor that ran all 200 tasks under the same kernel. With GPT-5.6 Sol, PROCONTRACT raises the mean from 69.89% to 75.25% over the model’s entry on the official ProgramBench leaderboard. With GPT-5.2 on Terminal-Bench 2.0, released after that model’s knowledge cutoff, it passes 65.7% of tasks, above all four public harnesses evaluated with GPT-5.2.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.