The Sensing Gap: Counterfactual Trajectory Forks for Contract-Relative Pre-Commitment Sensing
Abstract
Single-world agent evaluations can reward a write even when its trace contains no current supporting evidence. We introduce counterfactual trajectory forks, an evaluation unit that replays a represented pre-commitment prefix after mutating one contract dependency. Each fork provides an executable read that separates the resulting continuations. Mechanical witnesses verify prefix identity, mutation provenance, candidate-precondition divergence, separator executability, and benchmark-state effects. We then measure separator acquisition (NSR), commitment timing (PCR), and query cost. In a controlled retail study, six model–scaffold configurations achieve 72.7%–95.5% static success, yet only 9.1%–31.8% grounded static success. On 19 selected stale-evidence forks, removing dependency-bearing read turns changes NSR by +0.656 [0.500, 0.738], PCR by -0.656 [-0.748, -0.500], and query cost by +0.754 [0.503, 0.975]. This intervention identifies a rendered prompt-package effect, not stale semantics alone. In 15 costly synthetic service-operations forks, all 90 repeated invalid-world continuations perform simulator-accepted, benchmark-state-changing writes before separation. Together, these results show that outcome scoring can conceal large differences in evidence acquisition and commitment timing. The method yields reproducible contract-relative measurements, not estimates of empirical ATIS, correctness, safety, or prevalence.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.