acceptodds
Under review as a conference paper at ICLR 2027

ActCI-Bench: Evaluating Selective Decisions in CI Maintenance Agents

Abstract

A software agent responding to a failed continuous integration (CI) run must decide whether to modify the repository, take a non-patch action, or defer to another owner. Repair benchmarks that start from known patch tasks do not evaluate this decision. We present ActCI-Bench, an evaluation of selective CI maintenance decisions over 1,543 real failure episodes. Each episode links failure evidence to three targets: actionability for the current author, root cause, and appropriate response. The benchmark combines logs and repository context with available CI history and developer handling, and evaluates both classification quality and false repairs on non-actionable episodes. In our evaluation, 509 episodes (33.0%) are non-actionable for the current author. Repository-aware agents achieve higher macro-F1 than specialized triage baselines, yet still propose patches on 41–48% of non-actionable cases. Richer context and an explicit diagnosis-before-response workflow improve response selection in the reported retrospective setting. These findings identify response selection and restraint as distinct dimensions of coding-agent capability. Because some evaluated context includes later developer handling, full-context scores should be read as retrospective evidence-rich results rather than estimates of performance at the moment a CI run fails.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.