acceptodds
Under review as a conference paper at ICLR 2027

From One Monitor to One Joint: Tractable Probabilistic Inference for Agent Safety with Tensor Networks

Abstract

A safety monitor typically estimates the risk of a completed agent trajectory, but runtime safety must reason as the available evidence changes: observations may be missing, actions assessed before execution, or future violations anticipated from a prefix. Existing methods can answer these queries individually, but do not generally ensure that their answers are probabilistically consistent. The missing ingredient is a shared joint distribution. We therefore propose TTMonitor, which extends a black-box completed-trajectory monitor into a joint distribution over trajectories and safety outcomes, so that different runtime queries become conditionals of a unified probabilistic model. TTMonitor combines an estimated trajectory distribution with monitor distillation and represents the resulting joint as a tensor train, enabling tractable marginalization over various types of evidence. Experiments on InjecAgent and AgentDojo show that TTMonitor maintains high fidelity to the original monitor while ensuring cross-query consistency. Exact tensor-network contraction further avoids the accuracy–latency tradeoff of completion sampling and remains practical relative to direct neural prediction.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.