SysRepair-Bench: Measuring Whether Agents Can Fix a Vulnerability Without Breaking the Service
Abstract
Vulnerability remediation is the operator's last step: change a running system so a disclosed weakness is no longer exploitable while the protected service keeps serving. No benchmark measures this. Offensive benchmarks reward exploits with no availability obligation, and repair benchmarks edit source and re-validate a rebuilt target. We present SysRepair-Bench, remediation scenarios, each a live vulnerable service an agent must fix through a shell; the binary scenarios are graded by a two-component verdict: the exploit is blocked and the service stays healthy. Every scenario in the main grid runs in two conditions, one supplying the vulnerability report and one requiring the agent to find the fault itself. Our central result is the gap between them. Given the report, the median model and suite cell reach , and the strongest models approach the ceiling; required to locate the fault first, the same models on the same hosts with the same budget reach a median of . Every model loses ground, by a median of points. The report-informed condition is an anchor: it measures execution once a fault is named, and the gap is the cost of withholding the report. Collateral damage follows the same axis. The Collateral-Damage Rate, the share of exploit-blocking successes achieved by breaking the service, reframes remediation as specification gaming; without the report it is higher for seven of the eight models. We release the benchmark, the two-component verifiers, the scripted validity agents, and the evaluation harness.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.