Global Cache Workspace: Towards Robust Reasoning through Internal Computation
Abstract
Language models are usually adapted by updating their weights, trading broad, out-of-domain competence from large pre-training runs for in-domain gains. We show that adaptation can instead live in a model's memory. We introduce the Global Cache Workspace (GCW), a module inspired by Global Workspace Theory that selects information from specialised heads across the layers of a frozen backbone, integrates it in a shared workspace, and broadcasts an update back into their memory, covering both attention and linear-attention layers. Trained with next-token prediction while every backbone parameter stays fixed, GCW reaches in-domain gains competitive with full fine-tuning and LoRA while forgetting far less: across every adaptation setting we study, it loses the least out-of-domain reasoning accuracy on average, and after mathematical adaptation its loss is less than half that of weight-space methods. Its adaptation is also selective. With thinking enabled, GCW outperforms full fine-tuning in domain while shortening in-domain reasoning by up to 82%, a saving comparable to weight-space methods, yet keeps reasoning on unfamiliar tasks close to its original length; weight-space methods shorten that reasoning too, and lose far more accuracy there. Adapting what a frozen model remembers, rather than its weights, thus adds capability while preserving far more of what the model already knows.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.