acceptodds
Under review as a conference paper at ICLR 2027

SeekOrStop: Benchmarking Tool-Mediated Blue-Team Investigation under Matched Evidence Access

Abstract

Can a model that makes correct decisions from supplied evidence also identify missing evidence, acquire it, and know when to stop? We introduce SeekOrStop, a benchmark that distinguishes evidence acquisition, decision quality, and completion using matched Passive and Active views of the same security investigation episodes. Episodes are generated through controlled execution in a real Kubernetes environment, with ground truth derived independently of sensor labels. Passive provides a registered reference evidence view, whereas Active begins with an initial observation and decides whether to retrieve additional records through finite read-only tools. Across 420 Expansion episodes and four model configurations, end-to-end Active accuracy is consistently 14–32 percentage points lower than Passive accuracy. Moreover, registered-route completion does not guarantee successful investigation: agents may complete the registered route yet make an incorrect decision, fail to submit, or continue retrieving after the registered decision requirements have been met. Controlled contrast and equivalence pairs further test joint correctness under specified relevant and irrelevant changes. SeekOrStop therefore evaluates evidence acquisition, evidence use, completion, and stopping separately, showing why supplied-evidence accuracy alone gives an incomplete account of agentic investigation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.