ObliGate: Obligation-Gated Runtime Repair for LLM Workflow Optimization
Abstract
LLM workflow optimization has largely focused on selecting or constructing workflows before execution, while failures after partial execution are commonly handled by retrying or regenerating the failed output. We study how to transition an already-executed workflow from a failed state to a verified state while preserving valid computation. We formulate post-failure repair as a transactional state transition over a versioned artifact DAG: runtime evidence opens a typed obligation that restricts legal interventions; cloned root probes test candidate repairs; dependency-closed replay restores consistency; and verifier improvement determines promotion. On strict MBPP+ (378 tasks, three seeds), ObliGate reaches 80.95±0.70% raw pass@1 at 325 task-action tokens, reducing tokens by 50.5% relative to fixed execution. In controlled latent-root failures, it recovers 76.9% of actionable cases versus 28.1% for draft-only repair. Across 228 naturally occurring failures, ObliGate reaches 33.3% accepted recovery, close to full-artifact replay with root probing (33.8%) and above direct draft repair (31.6%) and full rerun (20.6%); dependency-closed replay lowers normalized unnecessary recomputation by 0.083 in paired contrasts. The same evidence-to-action interface is instantiated for code, mathematics, and QA under adapter-registered, single-writer workflows.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.