acceptodds
Under review as a conference paper at ICLR 2027

NEURAL POPULATION EMBEDDINGS AS A TEACHER FOR BELIEF-GUIDED DECISION MAKING

Abstract

Low-dimensional structure in neural population activity often resembles the representational geometry that emerges in artificial networks trained on the same task, which suggests that recorded neural activity could serve as a teacher for artificial agents. Prior brain-tuning work, including its single application to decision-making, has transferred representations tied to the stimulus presented on each trial. To our knowledge, no study has asked whether neural activity can transmit a latent context that is absent from the input and must be inferred across trials from the history of choices and outcomes. Here we address this question using the International Brain Laboratory (IBL) task, in which mice choose the side of a visual stimulus of varying contrast while the prior probability of that side switches between uncued biased blocks, such that on zero-contrast trials the animal must rely on its inferred belief alone. Specifically, we train a recurrent agent on a simulated version of the same protocol and, building on the Brain2Model paradigm, add a transfer loss that aligns a low-dimensional belief embedding within the agent to a teacher constructed from CEBRA embeddings of brain-wide IBL population activity averaged by trial condition. We compare this brain teacher against a noise-less Bayesian ideal-observer teacher and a shuffled control that retains the same embedding vectors under a permuted condition assignment. On zero-contrast trials, the brain teacher matches the ideal-observer teacher across the mid-range of transfer weights despite receiving only a coarse belief label rather than the observer’s posterior, and it outperforms the shuffled control by a margin that grows with transfer weight while also accelerating learning. Importantly, this advantage cannot be attributed to regularization, since all three teachers produce the same small gain at low transfer weight, and it replicates with a self-supervised embedding constructed without any belief labels. Moreover, within the neural embedding itself, the separation between low- and high-prior states is largest at zero contrast, where the prior is the only information available to the animal. Together, these results provide the first demonstration that neural population data can substitute for a hand-designed teacher when the task-relevant variable is latent and must be inferred across trials.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.