acceptodds
Under review as a conference paper at ICLR 2027

Auditing Inference-Time Safety Gates in LLM Trading Agents with Fixed-Outcome Replay

Abstract

Inference-time gates are increasingly used to abstain from or override actions produced by agentic language models, yet end-to-end utility comparisons conflate changes in proposal generation with changes caused by the gate itself. We introduce Fixed-Outcome Replay (FOR), a diagnostic protocol that logs each proposal, gate trigger, deployed action, and realized outcome, and separates whole-system contrasts from proposal-fixed gate contrasts. We evaluate FOR on 3,214 MAG7 stock-day decisions collected from January 2024–October 2025. The unprotected baseline obtains +23.56% cumulative signed utility, compared with +6.37% for the full system, yielding a +17.19-point end-to-end gap that cannot be attributed to the gate alone. When the A1 proposal stream is held fixed, the default gate changes utility by −1.81 points (95% bootstrap interval [−6.40, 2.15]); trigger-level replay reports +3.70 points for low-confidence interventions and −5.48 points for divergence-only interventions. Thus, the same gate has heterogeneous value across trigger types, and aggregate utility is insufficient to characterize its behavior. Fixed-Outcome Replay provides a reproducible accounting diagnostic for coverage, trigger composition, conditional risk, and rejected-action value; its interpretation is limited to the fixed-outcome setting and does not estimate deployment-time causal effects.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.