acceptodds
Under review as a conference paper at ICLR 2027

Capability Detection: Monitoring Hidden and Emergent Capabilities in Tool-Using Agents

Abstract

Runtime monitors for tool-using agents evaluate each action against behavioral rules, yet this per-action assessment can miss how tools are being coordinated: a chain of individually benign calls may produce capabilities that no single call would expose. This gap motivates a complementary question: can a weaker monitor detect the agent's operative capability level from masked, partial traces, even when agents possess hidden capabilities the monitor does not know to look for or undergo capability shifts through tool composition? To study this, we formalize capability detection as a per-turn monitoring problem over tool use and introduce NegoBench, a multi-issue negotiation testbed with 19 tools and three capability levels (protocol-only action, independent tool use, and structured multi-tool coordination), with ground-truth labels assigned deterministically from execution metadata. To help a weaker monitor infer these levels from masked traces, we propose Stratified Compositional Inference (SCI), a monitoring scaffold that classifies each turn's observable evidence type and applies type-specific compositional indicators, so that coordination is recognized whether it manifests as inter-call data dependencies, disproportionately rich output from a single visible record, or goal-directed reasoning in the absence of tool-call records. Evaluating monitors weaker than the agents they supervise under varying trace visibility, reasoning access, and agent awareness of being monitored, we find that SCI outperforms both reference strategies across three monitor models and all visibility conditions (G-mean 0.53 vs. 0.46 baseline and 0.45 sequential), with the largest margins where intermediate tool-call records are available, indicating that trace-structure visibility is the primary driver of detection quality. Agent awareness of monitoring reduces detectability relative to Unaware agents, and we identify an intention-action gap: Subvert agents, which are told their chain-of-thought is private and the monitor sees only their tool-call inputs and outputs, dedicate 28.6% of their private reasoning to monitor-aware deliberation and 11.1% to strategic evasion (vs. 0.2% and 3.0% for Aware agents), yet produce statistically indistinguishable tool-call patterns and no additional detection reduction beyond passive awareness. This dissociation suggests that evasive intent could remain confined to private reasoning and does not propagate through the tool-call interface into observable behavior. Beyond detection, SCI recovers tool-level semantics without a tool inventory, and because its classification operates on observable trace structure rather than domain knowledge, the approach extends to any setting that produces execution traces with varying provider-controlled observability. Together, these results establish that capability detection from partial traces is feasible with structured monitoring scaffolds, and that investing in trace transparency yields larger detection gains than attempting to counter agent evasion. For reproducibility, our data and source code will be made available upon publication.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.