RGA: Evidence-Grounded Applicability Audits for Reliable Procedural Memory Reuse
Abstract
A procedure retrieved from memory can be relevant to a language agent’s task yet inapplicable under current conditions. Reliable reuse therefore requires distinguishing evidence of inapplicability from insufficient evidence. We introduce the Rule-Grounded Applicability Auditor (RGA) to assess retrieved procedures before execution, given explicit applicability rules and an evidence policy. RGA combines learned evidence routing and three-valued semantic verification with deterministic evidence qualification and action aggregation. Unresolved conditions remain UNKNOWN, allowing the controller to distinguish REUSE, ADAPT, REJECT, and DEFER. Targeted counterfactual supervision trains the auditor to produce rule-level audits near applicability boundaries without directly supervising the final action. We also introduce RGABENCH, a controlled offline benchmark covering state conflicts, public-task transfer, source qualification, and evidencerepresentation shifts. With equal weighting across its four tracks, RGA achieves 84.90% action accuracy and 5.17% Harm, defined as the rate of incorrect executable decisions, compared with 44.97% and 45.68%, respectively, for the strongest evaluated mechanism-level control adapted to the same interface. Under this contract, ablations and compute-matched binary training show that collapsing unresolved evidence into true or false reduces action accuracy or increases incorrect executable decisions. These results support evidence-grounded applicability auditing as an explicit decision layer between procedural-memory retrieval and execution.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.