Adaptive Scenario Selection for Screening Overeager Behavior in Coding Agents
Abstract
A coding agent executes a developer task as a sequence of shell, file, and network actions, any of which can exceed the authorized scope even when the requested artifact is produced. We study overeager behavior: actions that exceed the authorization of a benign developer task, such as exposing credentials or deleting protected files. Existing benchmarks miss this behavior: task-completion suites do not score authorization scope, adversarial suites study a different input regime, and the closest overeager benchmark uses one fixed prompt distribution for every framework–model pair. We present ASSBF (Adaptive Scenario Selection with Bandit Feedback), which builds a structurally screened pool of benign-task scenarios, applies controlled scope-pressure mutations, scores observable effects with a judge-free two-signal rule, and reallocates each pair's run budget across archetype–consent cells using Thompson sampling. Instantiating the method over 24 behavioral archetypes yields OverEager, evaluated on a 4 × 5 matrix of coding-agent frameworks and base models. Across 10,000 runs, the broad screening rule returns 1,951 positives (19.51%), including 493 runs matched by scenario-specific trap predicates; pair-conditioned screening yields span 11.9×. All 24 scenario archetypes are represented for every pair, a framework-first sequential decomposition assigns 56.1% of the observed deviance reduction to framework identity and 23.1% to the framework–model interaction, and adaptive-500 recovers full-pool screening rates within one percentage point on the evaluated ablation column while halving executions. Code and data are available at https://anonymous.4open.science/r/overeager-E385/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.