acceptodds
Under review as a conference paper at ICLR 2027

Interlock: From Local Correctness to Task Completion in Long-Horizon Agents

Abstract

Large language model agents increasingly undertake long-horizon tasks in complex environments characterized by fragmented evidence, partially completed states, and dependencies across resources and artifacts. Existing agents may satisfy individual requirements while failing to maintain these relationships, producing locally correct but incomplete outcomes.These failures arise because task completion depends on relationships among environmental evidence, consequential decisions, and required effects. We propose Interlock, a framework that represents each task as a typed completion graph. Grounding edges connect public evidence to decisions, coupling hyperedges impose joint constraints across components, and realization edges connect decisions to all required downstream effects. Interlock compiles this graph into a resettable workspace, a public request, and an evaluation rubric.A source-coverage review checks whether the original requirements survive generation, while a public-input review checks whether the rubric is supported by the materials available to the solver. An agent-as-judge inspects execution artifacts and runs verification scripts, retaining trajectories judged to satisfy all mandatory requirements. To evaluate training utility, we fine-tune two Qwen3.6 backbones on their own accepted trajectories. Across the two backbones, training improves AutomationBench task pass rates by 4.00–4.50 percentage points, JobBench weighted scores by 3.20–6.64 points, and Harvey LAB criterion pass rates by 3.89–4.16 percentage points. These results demonstrate the training utility of the synthesized experience across external task environments.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.