acceptodds
Under review as a conference paper at ICLR 2027

Guarded Commits: Transactional Human Approvals for LLM Workflows

Abstract

Large language model (LLM) workflows often require human approval before an irreversible action, such as merging a pull request that a coding agent opened. Many agent frameworks record the approval apart from the component that performs the action. Unless the code-hosting platform is configured to enforce it, nothing checks that the approval covers what is merged. We ask whether this matters and whether approval could be automated instead. In 30,796 agent-authored pull requests, 83.4% of merges carry no approving review, and 12.6% of 3,880 approved merges went in at a code version other than the one approved. We then use Learn-then-Test to bound, with stated confidence, the share of auto-approved items that a human rejected. On 8,077 reviewed agent pull requests and 17,007 writing-assistant decisions, no method that sees only the proposed change and its origin certifies any automation at a 5% bound on the full held-out sets at our sample size, and LLM judges certify at most 3.1% on a stricter subset. This holds for trained predictors and for frontier LLM judges, the best of which reaches an area under the ROC curve (AUROC) of 0.77. The reviewers' own pre-decision signals, which partly encode the decision, raise AUROC to 0.88. These results motivate guarded commits, which store each approval with its evidence in an append-only ledger. The only component that can perform the action checks the approval against the exact artifact, policy version, and executed steps. Automated decisions also need a recorded risk certificate, and reused decisions must point to the human decision they reuse.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.