EviTrace-DE: Evidence-Gated Generation of Traceable and Executable Detection Rules from CTI with LLM Agents
Abstract
Detection engineering converts attack behaviors described in cyber threat intelli- gence (CTI) into executable detection rules over security telemetry, but it remains expert-intensive and manual. Tool-use LLM agents can retrieve reports, inspect schemas, and query telemetry, but retrieval alone does not prevent unsupported conditions or event relations from entering the final rule, producing false or unexe- cutable rules. We study this failure mode as the evidence-commit gap and present EVITRACE-DE, an evidence-gated framework that separates open-ended LLM reasoning from rule construction through explicit evidence verification, enabling accurate and traceable rule generation. An LLM agent proposes candidate predi- cates and event relations, and programmatic validators verify candidate evidence and close the corresponding behavioral, schema, telemetry, or relational precondi- tions. Only candidates that pass all required verifications are admitted to P-DIR, an intermediate representation backed by evidence that binds each admitted con- dition to its supporting evidence. Deterministic compilers then generate KQL and Sigma rules using only the predicates and event relations admitted into P-DIR. The evidence-verification state also provides each LLM call with only the context needed for the current verification step, avoiding repeated full-history prompts. In CTI-REALM-50, EVITRACE-DE improves the normalized score from 0.537 to 0.752 and KQL F1 from 0.496 to 0.804, achieves a valid-submission rate of 100% and reduces the mean token usage per task by 98.7% compared to ReAct. In an independent AWS CTI-to-Sigma task, adapting verification to threat reports and API field definitions improves API-call F1 by 35.9% over a ReAct-style agent using the same model.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.