Emergent Specialization in Populations of Self-Supervised Collaborative Vision Experts Without a Shared Gate or Cross-Agent Gradients
Abstract
Consider a population of neural networks with no shared gate and no gradients flowing between them: each network has its own copy of the weights and is trained on its own, all see the same heterogeneous data, and any network may ask another for help through a forward pass. Does a division of labor emerge in such a population, and is it of any use? This is a candidate regime for training populations of models at scale, and it reverses how machine learning has organized a division of labor so far: in conventional approaches, such as mixtures of experts, one gate trained jointly with all experts decides who processes what, and specialization is typically assumed rather than measured. We answer the question in a controlled, small-scale proxy for predictive visual pretraining. Initially identical agents are finetuned with one shared objective, masked prediction of frozen DINOv3 features, on an unlabeled mixture of six visual domains, with no roles or central router. We measure specialization (does the best agent for an input align with the latent domain?) and utilization (is responsibility spread across the population?) along a ladder of regimes that removes central control step by step, ending in DISCO (DIStributed COllaboration), a fully local protocol in which each agent chooses a helper, reads its internal state through a gradient-free channel, and rewards its own router only for the local improvement the help produces. Specialization emerges, and it is useful: (i) a randomly routed population of the same size is worse than a single generalist while semantically routed populations are better, so specialization rather than population size is what helps; (ii) specialization survives the removal of the central router; (iii) the gradient-free exchange makes emergent expertise usable by non-experts: in DISCO alone a randomly chosen agent, helped by the expert, matches the solo generalist, while the experts themselves surpass it, and the effect persists on data outside the specialization mixture; (iv) local routers trained only on their own improvement select the emergent expert for 98% of inputs; (v) the effects persist across population size, agent capacity, data imbalance, and fine-tuning seeds. These results provide small-scale, measurable evidence for the population dynamics that decentralized predictive pretraining would require.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.