TracePact: Contract-Indexed Replay Effects with Support-Conditioned Emission
Abstract
Local replay scores drive failure localization and corrective updates, but their meaning depends on which replacement is sampled, how much of the trace is rerun, and which terminal event is measured; weak two-arm support can make an equally precise-looking scalar unusable. TracePact makes these choices part of the output: an estimate-or-refuse record binds a typed catalog-average effect to its replay contract, five measured support diagnostics, and calibrated direction-agreement confidence. The record is instantiated by randomized keep-versus-replace logging, a replacement-aware cross-fitted doubly robust learner, and a factorized emission predicate. With identical 5,000-replay budgets on ToolBench and API-Bank, TracePact raises task success from 70.1% to 71.4% and emitted top-1 localization from 53.8% to 62.8%, while reducing direction-agreement ECE from 0.084 to 0.038 relative to forced typed replay. At matched 91.6% coverage, it reaches 57.6% all-request localization versus 54.4% for calibrated empirical-score filtering, a paired 3.2-point gain (95% CI [0.8, 5.6]); on a disjoint 300-request study, all-request top-1 accuracy is 54.7% versus 49.0% with catalog-matched replay labels and 54.3% versus 48.0% with blinded labels. Controlled catalog variants change the fitted top-ranked segment in 27.3% of 13,842 comparisons and its effect sign in 16.8%, while backend support erosion raises abstention from 8.4% to 29.3% at fixed assignment propensity. These results establish a concrete interface for comparing, selecting, and consuming local replay effects in the evaluated agent-trace regimes.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.