acceptodds
Under review as a conference paper at ICLR 2027

Does the Workspace Do the Work? A Causal Protocol for Verifying Global-Workspace Computation in a Language Model

Abstract

Global Workspace Theory (GWT) is a leading theory of consciousness. It holds that information becomes conscious when it enters a limited-capacity workspace and is broadcast to specialized processes that use it. As AI is increasingly used to build and test such theories, an essential step is to determine whether a workspace built into a network is actually used or merely present. We introduce a causal protocol that answers this in stages. It asks whether the model depends on its workspace, whether receiving modules compute with the broadcast content, how that content acts, and whether it reaches several receivers at once. We apply it to a small language model whose modules can communicate only through a learned workspace. We confirm the results on four models trained under a protocol fixed in advance. Switching the broadcast off lowers accuracy by 41–52% on questions that need two modules' knowledge, while changing accuracy on single-module questions by at most 1.4 points. Dependence alone does not show use, so we then replace the broadcast with one taken from a different input. The receiver then combines the new content with its own memory, rather than copying the other input's answer, in 61–73% of cases, compared with 5–13% without intervention. Undoing only the resulting change in what the receiver retrieves restores the original output in 2,047 of 2,048 cases. This shows that the content acts through memory retrieval. A separately trained variant shows one broadcast serving three receivers. The protocol turns GWT's central claim into a testable property of trained models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.