acceptodds
Under review as a conference paper at ICLR 2027

VerDICT: Verifiable Decentralised Inference-only Collective Training

Abstract

As AI systems increasingly serve users with differing interests, training shared models requires both efficient participation and principled preference aggregation. We introduce VerDICT, a decentralised training approach that allows participants to contribute computation while expressing their preferences over model outputs. VerDICT combines EGGROLL with a scalar pairwise preference objective targeting maximal lotteries, which accommodate cyclic preferences and select the Condorcet winner when one exists. Training operates entirely through inference, requiring no backward passes or separate reference model, and shares baseline responses across workers to reduce computational costs. Each worker sends only its votes, and every update is a deterministic function of the shared seed and the logged votes, so it can be recomputed and checked. We assume workers report their preferences honestly. On controlled preference games VerDICT approaches the maximal lottery where DPO under the same optimiser does not, and on summarisation and two human preference datasets it improves judged quality over the initial policy at lower per-step cost than EGGROLL-DPO.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.