When Unsupported Claims Become Working State: Measuring Ratification in Sequential Agents
Abstract
The same unsupported claim can be challenged and disappear in one agent trace, yet be accepted as a premise and redirect retrieval or action in another. We call the latter transition ratification and represent it as a claim–consumer state change, with separate injection, endorsement, and operational-reuse events and event-specific risk sets. This representation supports matched policy tests and EndorseGuard, a plan selector whose epistemic critic predicts whether a candidate will carry unsupported content across the state-admission boundary. On 200 held-out Persuasion-Conflict scenarios, assertion-time and adoption-time guards yield 8.2% and 4.6% conditional propagation, respectively, under identical planner, detector, feedback, proposal, and refinement settings. The ordering survives blinded human adjudication (a 3.6-point gap), a component-disjoint DeBERTa/BGE evaluator (3.5 points), and direct annotation of all 500 TopiOCQA test dialogues (6.1% versus 4.1%). Under matched planner-call budgets, EndorseGuard reduces Persuasion-Conflict propagation from MARCH's 4.1% to 2.1% while improving success from 78.4% to 81.0%; on TopiOCQA it raises answer F1 from Utility-only's 55.6 to 60.4, with the largest propagation reduction at topic switches (9.8% to 3.7%). Treating state admission as an explicit reliability event therefore reveals consequential contamination that response-local factuality misses and supplies a concrete control target for sequential agents.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.