SecRespond: Benchmarking and Training AI Agents for Real-World Post-Compromise Incident Response
Abstract
Large Language Model (LLM) agents are increasingly adopted in real-world security operations with access to host artifacts and command-line interfaces (CLIs), making it critical to assess their security capabilities. However, existing cybersecurity benchmarks mostly focus on pre-compromise settings where agents operate in a clean environment before an attack occurs. This leaves the post-compromise setting underexplored. To address this gap, we introduce SecRespond, the first benchmark for evaluating LLM agents to autonomously conduct the post-compromise incident-response workflow in an end-to-end manner. Given a forensic disk snapshot of a compromised host together with the alerts, vulnerability scans, and baseline checks reported by a host security product, agents are required to produce forensic reports on intrusions, baseline risks, and vulnerability risks, together with a remediation plan. We instantiate this task across 10 cyber ranges, each constructed from a distinct compromised cloud host, spanning 4 entry-point types, 21 ATT&CK techniques, and 5 operating systems. We evaluate 23 frontier LLMs on the OpenCode agent harness. Experimental results show that although current agents can reliably uncover the problems exposed by alerts, they struggle to proactively investigate the disk for silent intrusions and to produce comprehensive, verified remediation plans. We further propose a training recipe for post-compromise agents with skill internalization and RETrace, a report-anchored credit assignment, which assigns each rubric checkpoint's score to the corresponding output span and the investigation step. On out-of-domain ranges, RETrace improves over outcome-reward RL by 8.8% in detection and 7.2% in planning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.