CrashDiag: Mechanically Verified Reinforcement Learning for Executable Infrastructure Repair
Abstract
Executable agent environments can verify terminal states, but terminal verification alone does not ensure that an episode validly tests diagnosis and repair. We audit CrashDiag, an infrastructure-repair environment with 52 multi-fault workflows and mechanically checked state transitions. The audit identifies a scenario-generation defect: a history action intended as an unsuccessful remediation partially repairs an active hidden sub-fault before the policy receives its observation. Reconstructing 192 retained held-out composite episodes from their seeds shows a hard boundary: the relevant sub-fault is resolved in all 117 noisy or shifted-noisy episodes and in none of the 75 redacted episodes. Thus, 34 of 228 retained Group Relative Policy Optimization (GRPO) exact successes and 8 of 16 base-policy exact successes are affected by pre-inference state leakage. We invalidate the prior policy comparison rather than post-hoc adjusting it. We formalize a pre-inference task-integrity invariant and give an audit protocol, regression-test requirements, and reporting requirements for mechanically verified agent environments.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.