TraceGate: Randomized Support Calibration for Local Retrieval Edits in Repeated Agents
Abstract
Should a repeated agent place a previously useful fragment into its next retrieval slate after tool reliability or execution policy changes? TraceGate answers with a randomized support-calibration experiment at the candidate occurrence: forced inclusion and exclusion define the verifier-pass contrast, and task-held-out Horvitz–Thompson bin effects evaluate frozen scores without imputing individual potential outcomes. A signed-clipped AIPW estimate may change slate membership only after overlap, a stressed interval, uncertainty, and set-size checks agree; otherwise a development-fixed scorer retains control, while an optional bounded probe can resolve ambiguity. The resulting frontier is sharp: over 15,600 retained opportunities, scores attain 87.4% sign agreement and 1.18 pp MAE, while the no-probe rule promotes 64.0% with 6.1% bin-sign disagreement. Bypassing individual gates raises disagreement to 7.4–9.8%; matched-coverage score-only DR and same-pool CRM/DR yield descriptive rates of 10.9% and 9.2%. These local decisions improve completed work: across 9,750 paired episodes, verifier-confirmed success rises by 1.8/2.2/2.6 pp over the setting-wise strongest non-causal reranker under tool confounding, policy shift, and correlational traps, with all three paired intervals excluding zero; repository-blocked folds and Qwen2.5 retain positive same-pool point differences. By tying each override to a randomized effect target and explicit support evidence, TraceGate converts logged causal estimates into reliable retrieval reuse without additional probes at its primary operating point.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.