acceptodds
Under review as a conference paper at ICLR 2027

Extracting, Removing and Shrinking Domains from One Jointly Trained Network

Abstract

Deploying language models for different sets of text domains and memory budgets usually means training a separate model for each. Dropout already trains many subnetworks inside one network but uses only their average. We ask whether single subnetworks can be deployed on their own, for one domain, several domains or a smaller budget, without a router or further training. We introduce region dropout, which assigns each domain a region of every feed-forward layer: within a region, the drop rate rises with a unit's position, so that early units form smaller models; across regions, other domains' regions are included and trained on each domain's data with a probability that rises during training, which we call exposure. Without exposure, combined models containing every unit of an arithmetic member solved 5 and 53 of 192 problems at 30M and 124M parameters that the member solved completely; with exposure, they solved all 192, and combined members at full width stayed within 0.11 nats of their single-domain members on external text and within 0.008 nats at 1.3B parameters, while the code member lost 0.06 to 0.13 nats against dense models. In a matched study, neither joint inclusion alone nor training regions on other domains' data alone kept the largest loss increase from combining below 0.10 nats. Weighting the regions of models trained without exposure by a domain posterior, as in expert mixtures, gave 0.03 to 0.12 nats lower text loss than exposure at 5.6 times the scoring time but solved at most 60 of the 192 problems. Removing a region reduced the exact answers on its task from 188 and 187 to 1 and 0 of 192 and left the other domains within 0.003 nats; retraining on the task restored most answers. A code member exported at 75% width kept 32.3M of 51.3M parameters at 0.05 nats above full width; narrowed combinations kept fewer exact answers. One network trained with region dropout thus provides models for each subset of its domains, without a domain, or for single domains at smaller width, on text from these domains.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.