ContractRL: Shielded Group-Relative Policy Optimization for Auditable Tool-Call Repair
Abstract
Structured tool calls often fail after only a small number of fields violate a schema or an execution contract. Regenerating the complete object enlarges the action surface and makes repeated repair difficult to audit. We introduce ContractRL, a contract-constrained sequential repair protocol that models verifier-guided JSON repair as a bounded decision process. At each step the policy observes the candidate, typed verifier feedback, JSON Pointer, immutable repair history, and remaining budget; a contract-derived action mask filters malformed or prohibited RFC-6902 operations before a deterministic validator performs the transition. We specify a contract-constrained group-relative objective for patch, retry, and abstention decisions while keeping canonical targets and semantic labels outside the online state until trace freeze. On the recorded five-seed held-out sweep, the 2B/r8 ContractRL row reaches 0.9331 semantic success at 34.8 generated tokens and 1.32 reported p95 latency units; the accompanying Patch-SFT, full-regeneration, and constrained-decoding rows reach 0.8924, 0.9041, and 0.8671. Under the completed same-information control, ContractRL reaches 0.9362 semantic success with 34.4 generated tokens, 1.30 p95 latency units, and 0.0169 collateral edits, compared with 0.9076/44.9/1.49/0.0488 for Patch-SFT and 0.9148/137.2/2.54/0.1876 for full regeneration. With the same public state and budgets, policy optimization reaches 0.9375 semantic success versus 0.9186 for ContractRL-SFT and 0.8927 for Patch-SFT. A separate three-seed paired ledger against Patch-SFT reports a semantic difference of (95% CI , ). Scale, feedback, budget, schema-shift, noisy-verifier, transfer, selective-risk, uncertainty, and blinded audit results identify where bounded repair helps and where it should abstain.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.