acceptodds
Under review as a conference paper at ICLR 2027

Learning representations of operating system level data and designing interpretable AI for defensive cybersecurity

Abstract

Frontier models have developed the emergent capability to formulate cybersecu- rity exploits and use them agentically with many recent high profile incidents. In large part due to these incidents, many surveys indicate a majority of Americans now think of AI as an existential safety risk. With frontier models speeding up the formulation and execution time of exploits beyond human reaction time in multiplicative effect with increasing the scale of attacks through agentic swarms, the value of high precision detection of exploits as a defensive counterbalance becomes greater; both to safeguard valuable computing systems and to allow the larger field of Artificial Intelligence to safely advance. Toward these ends, vec- tor representations of cybersecurity exploits were learned on operating system level data, including command lines and executables used in host operating sys- tems. These representations were trained on by interpretable ensemble machine learning approaches to make interpretable detectors. Some of the detectors had greater than 90% precision when tested on reserved data, even when the targets were a minority class with other logs dwarfing the numbers of true positives by orders of magnitude. This high precision for detecting minority classes in large swaths of operating system level data is very valuable in defensive cybersecurity because most networks on average generate thousands of logged events per com- puter per day with many false positives using up responders’ time thereby delaying responses to and often completely burying the true positives. In the new era of AI and cybersecurity where agentic AI swarms are conducting offensive cyberattacks faster than current defensive systems can respond, in large part due to the lower precision and lack of interpretability of previous systems, it seems imperative for safety to use AI, and preferably interpretable AI, defensively as a counterbalance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.