acceptodds
Under review as a conference paper at ICLR 2027

FragBench: Multi-Session Cyber Attacks Hidden in Benign-Looking Fragments

Abstract

Production LLM safety guardrails check one request at a time, yet documented cyber campaigns are already decomposed across many requests. A request that appears legitimate can therefore contribute to a harmful operation whose purpose is only apparent from the broader campaign. We introduce FragBench, a benchmark that includes attacks based on reported cyber campaigns, split across multiple model API requests to be executed in separate agent sessions. A catalogue of 35 cyber-incident and threat-report entries informs 24 campaign templates. The construction pipeline generates fragments and records their dependencies, safety-judge verdicts, tool activity, and execution-checker outcomes. Judge-guided iterated rewriting increases mean bypass from to on the evaluated subset. Classifiers trained on malicious and benign MCP executions achieve event F1 up to . FragBench supports analysis of both attack execution and detectability by retaining failed attempts alongside successful runs. We release the benchmark and evaluation code at https://anonymous.4open.science/r/fragbench-FE86/.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.