Atoms to Agents: Efficient Proxies for Cybersecurity Agentic Evaluation
Abstract
End-to-end benchmarks for cybersecurity large language model (LLM) agents are costly to run: they require interactive environments, tool execution, and multi-turn reasoning, yet ultimately provide only a coarse measure of overall task success. This makes them expensive for repeated model comparison and provides limited visibility into the underlying operations that contribute to success or failure. We introduce ATOMS TO AGENTS, a framework that decomposes agent trajectories into recurring atomic tasks, organizes them into a trajectory-derived capability taxonomy, and converts each into an isolated evaluation with an independently verifiable answer. By posing each task from a fixed trajectory state, the framework prevents upstream errors from propagating and avoids full multi-turn rollouts. The resulting tasks assess both operation execution and, where applicable, action selection in context, localizing model weaknesses to specific capabilities. On a panel of 13 models evaluated on ExCyTIn and CTI-REALM—the two benchmarks represented in bank construction—atomic-task scores correlate with full agentic performance at Spearman ρ ≈ 0.9. Compared with uniformly sampled benchmark subsets that achieve the same Spearman correlation with the full-benchmark model ranking, our best observed 25-item proxy uses 76× fewer tokens on ExCyTIn and 96× fewer on CTI-REALM. Transfer to a new workflow is weaker when relevant capabilities are missing from the evaluation set, but the framework’s taxonomy helps identify these gaps and guide targeted expansion, hence improving ranking agreement. These findings support efficient model comparison and capability diagnosis while highlighting the need for coverage-aware maintenance and end-to-end validation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.