acceptodds
Under review as a conference paper at ICLR 2027

TRUCE: Conflict-Conditioned Interventional Routing for Tool Agents After Updates

Abstract

After a tool or fact update, an agent can simultaneously retrieve a current semantic record, an obsolete procedural recipe, and a successful episodic trace for the retired behavior; exposing all three can turn relevance into the wrong next call. Aggregate retrieval accuracy does not answer the resulting systems question: when is the conditional reliability gain from resolving typed disagreement large enough to justify randomized estimation, uncertainty calibration, and routing latency? The conflict–value profile answers this question by indexing held-out policy gain by pre-routing disagreement, support, latency, and exploration cost. Its realization, TRUCE (Typed-memory Routing Under Conflicting Evidence), estimates same-background singleton and pairwise store contrasts from exact post-clipping action propensities, gates weakly supported scores, composes a latency-feasible subset, and resolves revision contradictions. Across three public update benchmarks, TRUCE gains success points over an exposure-matched supervised router in conflict-heavy episodes, while low-conflict differences are points; gains over the strongest same-feature subset policy are . Episode-wise randomization of store identities reduces the learned-router advantage to points (95% CI ), and greedy routing agrees with exact fitted-score search on 94.1%/93.4% of public-core decisions while finishing within 0.2/0.2 success points. At of the reference exploration-and-training cost, the full stack retains of its public-core lift, making typed conflict a measurable trigger for spending routing complexity after updates.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.