FragBench: Multi-Session Cyber Attacks Hidden in Benign-Looking Fragments
Abstract
Production LLM safety guardrails check one request at a time, yet documented cyber campaigns are already decomposed across many requests. A request that appears legitimate can therefore contribute to a harmful operation whose purpose is only apparent from the broader campaign. We introduce FragBench, a benchmark that includes attacks based on reported cyber campaigns, split across multiple model API requests to be executed in separate agent sessions. A catalogue of 35 cyber-incident and threat-report entries informs 24 campaign templates. The construction pipeline generates fragments and records their dependencies, safety-judge verdicts, tool activity, and execution-checker outcomes. Judge-guided iterated rewriting increases mean bypass from to on the evaluated subset. Classifiers trained on malicious and benign MCP executions achieve event F1 up to . FragBench supports analysis of both attack execution and detectability by retaining failed attempts alongside successful runs. We release the benchmark and evaluation code at https://anonymous.4open.science/r/fragbench-FE86/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.