acceptodds
Under review as a conference paper at ICLR 2027

Byzantine-Robust Federated RAG via Aligned Calibration and Fixed-Membership Conformal Prediction

Abstract

Language models answer questions more accurately when they can consult relevant documents, an approach called retrieval-augmented generation (RAG). Many valuable collections, such as medical records or company files, cannot be pooled in one place because of privacy rules or ownership. Federated RAG leaves each collection with its owner: each owner, or node, searches its own documents and scores candidate answers, and a central hub combines the scores. Because the hub cannot inspect the nodes, some of them, called Byzantine, may be compromised, faulty, or misled by malicious instructions hidden in documents, and report arbitrary scores. We want the hub to return a set of candidate answers that contains the correct one with a chosen probability, the guarantee offered by conformal prediction. Conformal prediction keeps every answer whose score passes a cutoff, and the hub sets that cutoff in a calibration step, using questions with known answers. An unknown group of nodes, no larger than a declared bound, may misreport both in this step and when new questions are answered. Existing methods assume every node is honest or protect only the calibration step. Our method rests on a simple observation: the honest nodes are the same in both steps. The hub therefore has all nodes score the same calibration questions. It keeps a candidate answer only if some plausible group of honest nodes, using its own scores in both steps, would keep it. We prove three guarantees. The resulting sets contain the correct answer with the chosen probability in finite samples, whatever the Byzantine nodes report. No method using the same information can return smaller sets without risking the loss of an answer the honest nodes support. If nodes fail at random, the guarantee weakens only by the probability that more nodes fail than declared. We tested the method in simulations and on real question-answering tasks, including medical exams. We also used a panel of language models as nodes, some of them hijacked by injected instructions. Our answer sets reached the target probability whenever no more nodes misbehaved than declared. They were also clearly smaller than those of simpler methods offering the same protection, most of all when the declared bound was generous. Simply averaging the nodes' scores could miss the target. In practice, an operator can therefore declare a cautious bound on the number of bad nodes without paying much for it in the size of the answer sets.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.