Can Federated Agents Agree and Still Be Unsafe? FedSafeBelief for Repairing False Safety Consensus
Abstract
Federated agents may agree on a safety decision because they share the same blind spot. Pooling such decisions can preserve an unsafe rule and return it to clients as apparently reliable feedback. At the same time, differences between local policies can make a rule valid for one client and invalid for another. We study this combination as false safety consensus and propose FedSafeBelief, a correction procedure over client scoped safety beliefs. Clients expose abstract decision tuples while retaining their raw interaction traces. The coordinator aligns these tuples, checks policy scope, searches for counterevidence to both disputed and unanimously allowed decisions, and returns a field level correction. Its central restriction is fix not add: correction may revise an existing client belief but may not create a new client key. We characterize when union aggregation dilutes client scoped precision, why shared errors survive voting, and when verified correction contracts the remaining error. The evaluation design separates belief repair, unsafe feedback, response safety, and legitimate policy heterogeneity. Experiments across controlled belief correction, harmful request, and heterogeneous policy settings show that FedSafeBelief consistently reduces unsafe allows and improves client scoped belief precision while preserving legitimate policy differences.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.