SeekOrStop: Benchmarking Tool-Mediated Blue-Team Investigation under Matched Evidence Access
Abstract
Can a model that makes correct decisions from supplied evidence also identify missing evidence, acquire it, and know when to stop? We introduce SeekOrStop, a benchmark that distinguishes evidence acquisition, decision quality, and completion using matched Passive and Active views of the same security investigation episodes. Episodes are generated through controlled execution in a real Kubernetes environment, with ground truth derived independently of sensor labels. Passive provides a registered reference evidence view, whereas Active begins with an initial observation and decides whether to retrieve additional records through finite read-only tools. Across 420 Expansion episodes and four model configurations, end-to-end Active accuracy is consistently 14–32 percentage points lower than Passive accuracy. Moreover, registered-route completion does not guarantee successful investigation: agents may complete the registered route yet make an incorrect decision, fail to submit, or continue retrieving after the registered decision requirements have been met. Controlled contrast and equivalence pairs further test joint correctness under specified relevant and irrelevant changes. SeekOrStop therefore evaluates evidence acquisition, evidence use, completion, and stopping separately, showing why supplied-evidence accuracy alone gives an incomplete account of agentic investigation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.