Learning to Select Constraints for Coding Agents
Abstract
Coding agents must follow written constraints on how they implement, test, and report their work. Different constraints apply at different steps, so a monitor must decide what to check before each action. We use OctoBench, a benchmark that provides coding tasks, their written instructions, and a checklist of requirements for each task. The task instructions provide the text to retrieve, and the checklists specify the requirements used to evaluate constraint following. We propose Applicable Constraint Retrieval (ACR), which learns to select instructions according to the proposed action and earlier interactions. We further design Applicable Constraint Monitoring (AConM), which uses the retrieved instructions to check proposed actions and request revisions when needed. On OctoBench, ACR retrieves instructions that fully express 66.96% of applicable constraints on average across tasks, compared with 56.90% for a pretrained retriever given the same inputs and budgets. On OctoBench, 85.34% of requirements are marked as met with AConM, compared with 79.72% without monitoring, averaged equally across three coding models. In an exploratory evaluation on 50 SWE-Gate tasks, we reuse the same retriever without further training. With DeepSeek-V4-Flash, 38% of tasks pass both functional and constraint tests with monitoring, compared with 32% without it. The source code is available at https://anonymous.4open.science/r/aconm-3219.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.