acceptodds
Under review as a conference paper at ICLR 2027

ProcCtrlBench: Evaluating Process-Level Defects and Control Preservation in LLM Coding Agents

Abstract

Outcome-based evaluation of LLM coding agents provides limited information about defects that arise during execution. We present ProcCtrlBench, a benchmark for execution-process evaluation in LLM coding agents. ProcCtrlBench organizes recurrent execution defects into a reusable ontology covering 11 defect types in 4 categories, and evaluates agent trajectories through standardized process evidence rather than final outcomes alone. To support cross-platform comparison, ProcCtrlBench standardizes raw logs into a shared trajectory representation and reports calibrated scorecards over process-level findings. In addition, ProcCtrlBench uses control preservation to quantify execution-process quality, capturing whether execution remains interpretable, interruptible, correctable, reversible, and able to hand back authority when needed. We evaluate ProcCtrlBench on 200 execution trajectories sampled from three benchmarks, corresponding to 2,200 defect-level evaluation points under the full 11-defect annotation protocol. Results show that ProcCtrlBench can be instantiated with useful reliability, provides more stable semantics than direct thresholding, and reveals distinct differences in execution quality often overlooked by conventional outcome-based metrics. Our code are available at https://anonymous.4open.science/r/ProcCtrlBench-655C.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.