acceptodds
Under review as a conference paper at ICLR 2027

Sabha: Sparse Latent Coordination of Specialist Language Models

Abstract

We introduce Sabha, a modular language model that coordinates separately pre- trained domain specialists through a compact learned interface. Each specialist, or island, is a dense transformer. A router selects a subset of islands, which emit pooled hidden-state messages into a shared latent bus; a Council decoder uses these messages to generate the system’s output. The architecture separates specialist transformer blocks from the coordination interface and does not require one pre- trained island to serve as the output hub. With three 234M-parameter islands and a 25M-parameter Council, the top-2 configuration achieves aggregate multitask cross-entropy of 7.708, compared with 7.721 for a 729M dense baseline and 7.873 for our Branch-Train-Stitch implementation. Its estimated inference FLOPs per token are 74.5% of the dense baseline’s. An individual code-island update reduces system-level code loss by 0.12, with a 0.01 increase on reasoning and a 0.29 de- crease on general text. These findings support sparse whole-model coordination as a practical design point for modular language modeling. They establish neither a general advantage over dense models nor behavioral isolation under arbitrary updates: the quality margin is small, routing underuses one specialist, and the current sequential implementation is slower despite its lower estimated arithmetic cost.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.