acceptodds
Under review as a conference paper at ICLR 2027

CyFER-Bench: Cybersecurity Fault-trace Execution Reasoning Benchmark

Abstract

Large language model (LLM)-based agents are increasingly used in cybersecurity, where they must interpret threat evidence, reason about relevant weaknesses and vulnerabilities, and translate those conclusions into actions in a target environment. Evaluating these agents therefore requires assessing both their stage-specific capabilities and whether those capabilities compose across the full workflow, while also evaluating whether LLMs can reason about their own failures. More broadly, cybersecurity provides a verifiable setting for studying a fundamental question about LLM agents: can models apply available knowledge correctly in context, translate it into effective action, and diagnose failures in their own reasoning? We present CyFER-Bench, a linked benchmark spanning (1) malware-report tactics, techniques, and procedures (TTP) extraction, (2) Common Weakness Enumeration (CWE) identification and explanation, (3) executable weakness exploitation, (4) multi-stage attack planning and execution, and (5) attacker-intent inference. CyFER-Bench provides validated intermediate targets and sandboxed execution, enabling stage-wise evaluation of knowledge, contextual reasoning, planning, action, and failure localization. Our findings highlight recurring challenges for current LLMs on complex multi-stage tasks: strong performance at one stage does not reliably carry to the next, local knowledge does not always translate into coherent reasoning, planning, and action, and models struggle to identify where their own reasoning fails.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.