acceptodds
Under review as a conference paper at ICLR 2027

TraceSpec: Active Evidence Acquisition for Executable Monitor Synthesis

Abstract

Executable monitor synthesis from language contains two distinct uncertainties: semantic uncertainty about how an alert maps to deployment-local signals, and policy uncertainty about stateful parameters that language alone does not identify. We formulate active monitor synthesis, where a model receives alert language, an opaque local metric/tool catalog, and a measurement budget, and must choose which evidence to acquire before emitting an executable stateful monitor. We introduce TraceSpec, which combines a compact language model for semantic binding and value-of-information decisions with bounded symbolic inference for temporal-policy identification and compilation to MonitorIR. Activation evidence constrains onset-side behavior, whereas recovery evidence resolves parameters that remain observationally equivalent after activation alone. On controlled deployments derived from held-out Thanos and Mimir rules, Qwen3.5-4B reaches 0.904/0.929 execution agreement, versus 0.686/0.738 for the strongest compact controls and 0.912/0.942 for a gold-semantics/oracle-tools upper bound; Phi-4 reaches 0.903/0.922. One informative phase raises execution from about 0.48 to 0.80, and a complementary second phase adds another 10.2–13.9 points. The learned acquisition policy also transfers to an unseen cost value and budget. Within the bounded monitor family studied here, these results show that semantic reasoning, active evidence acquisition, and bounded policy identification form complementary parts of executable synthesis.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.