ReplayClaim: Executable Returned Support for Persistent Multimodal Claims
Abstract
A claim stored from an invoice can return with its original OCR crop and still be unsafe to reuse: the read path must bind the amount to the requested field, validate current service state, and suppress contradictory claims before an answer or tool consumes it. This missing read contract defines the persistence gap; ReplayClaim implements it as a returned-support object that stores evidence with typed verifier routes, calibration and model versions, write-time outputs, and provenance. Its read operator executes query-conditioned verification, constructs a contradiction-consistent served set, and sends invalid state to quarantine or re-grounding; decision completeness denotes whether the object reconstructs this declared read contract. Across 74 tool workflows with downstream queries at 5- and 10-session horizons, ReplayClaim raises \(K=10\) success from \(70.8\pm1.9%\) to \(72.6\pm1.8%\) versus an equally provisioned evidence cache (\(p=0.002\)) and reduces unsupported actions from 47/900 to 38/900. A complementary one-write/one-read benchmark on DocVQA, ChartQA, and ScienceQA compares returned objects inside a predeclared \(83.5%\pm0.5%\) coverage band: unsupported claims fall from \(8.7\pm1.0%\) to \(7.6\pm1.0%\) (paired 95% CI [0.5,1.7], \(p=0.0018\)), and method-blind truth error falls from 8.2% to 6.4%. Lifecycle enforcement further reduces stale-state attack success from \(28.2%\) to \(9.4%\), while the cache gap persists with LLaVA-1.5-13B (\(9.8\pm1.2%\) to \(8.6\pm1.1%\)) and after manual parser cleanup. The results make returned support an executable decision object and expose safe claim reuse as a property of the read contract, not evidence retention alone.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.