Auditing Inference-Time Safety Gates in LLM Trading Agents with Fixed-Outcome Replay
Abstract
Inference-time gates are increasingly used to abstain from or override actions produced by agentic language models, yet end-to-end utility comparisons conflate changes in proposal generation with changes caused by the gate itself. We introduce Fixed-Outcome Replay (FOR), a diagnostic protocol that logs each proposal, gate trigger, deployed action, and realized outcome, and separates whole-system contrasts from proposal-fixed gate contrasts. We evaluate FOR on 3,214 MAG7 stock-day decisions collected from January 2024–October 2025. The unprotected baseline obtains +23.56% cumulative signed utility, compared with +6.37% for the full system, yielding a +17.19-point end-to-end gap that cannot be attributed to the gate alone. When the A1 proposal stream is held fixed, the default gate changes utility by −1.81 points (95% bootstrap interval [−6.40, 2.15]); trigger-level replay reports +3.70 points for low-confidence interventions and −5.48 points for divergence-only interventions. Thus, the same gate has heterogeneous value across trigger types, and aggregate utility is insufficient to characterize its behavior. Fixed-Outcome Replay provides a reproducible accounting diagnostic for coverage, trigger composition, conditional risk, and rejected-action value; its interpretation is limited to the fixed-outcome setting and does not estimate deployment-time causal effects.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.