Certified Adaptive Refresh: Anytime-Valid Monitoring for Federated Conformal RAG
Abstract
Question-answering services built on retrieval-augmented generation (RAG), in which a language model answers from retrieved documents, are inspected continuously and upgraded repeatedly, so their reliability guarantee must survive both. We study federated conformal RAG: retrieval nodes holding private corpora score a shared list of candidate answers with a shared language model and send only compressed scores to a hub, which returns an answer set; a miss is a set that omits the true answer. We formulate monitoring as a sequential test: raise an alarm when misses exceed a certified upper bound on the miss rate (the allowance), while keeping the probability of any false alarm over the whole deployment below a preset budget; an alarm is false when no deployed configuration has drifted above its allowance. Existing tools solve only parts of this problem. A conformal certificate bounds the miss rate of one frozen configuration at one look fixed in advance, so it supplies no anytime-valid alarm, one that stays valid however often it is checked; a threshold calibrated for one model certifies nothing about its replacement; and re-certifying each upgrade at the full error level compounds failures. We propose Anytime-FC-RAG, a protocol with three auditable rules: pick the next configuration before drawing its fresh calibration data, charge every certificate attempt to one trajectory-wide budget , and fix the deployed threshold, allowance and monitoring bet before each query arrives. One betting wealth process , evidence that grows in expectation only when misses exceed the allowance, then runs across every deployment without reset. Our certified adaptive-refresh composition theorem shows that, for an evidence budget , alarming when first reaches has false-alarm probability under a random number of history-selected, evidence-driven refreshes of the model, retriever, corpus, score or threshold; pays for certificates that are wrong by chance. With an exact order-statistic certificate the guarantee is distribution-free and architecture-agnostic; it concerns misses against the allowance, not per-input coverage or distribution shift as such. Replaying Qwen2.5 on MMLU-Pro at a 5% total budget, all 210 runs whose shift an independent audit rated harmful alarm within 5,000 queries, and none of 222 benign or unshifted runs does. In a swap of the deployed model, calibrating the incoming model on fresh data holds a downgrade's miss rate at .089, under its .10 target, whereas calibrating it on the outgoing model's scores lets it rise to .154.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.