acceptodds
Under review as a conference paper at ICLR 2027

Sensory Geometry Sets the Reference Frame of Multimodal Spatial Codes

Abstract

Recurrent networks trained to predict their upcoming sensory input develop spatial representations resembling those of the hippocampus, constructing a world-centred (allocentric) map from egocentric input. Such models have almost always been given a single sense, usually vision, yet hippocampal representations draw on many senses whose signals depend differently on position and heading. Here we asked how the available senses determine the reference frame of the learned representation. An agent explored a virtual arena with first-person vision and binaural audio from four speakers; recurrent networks received one or both senses plus self-motion and were trained to predict their input one second ahead. Because egocentric signals change with the direction the agent faces, we measured spatial allocentricity: the consistency of each unit's spatial tuning across headings (1, identical at every heading; 0, unrelated). All trained networks became more allocentric, but audition far more than vision (median allocentricity vs ). This difference reflected the geometry of the signals rather than the senses themselves: what the agent sees depends on heading as well as position, whereas what it hears from the surrounding speakers depends mainly on position. Rearranging the speakers so that sound also depended on heading along one axis reduced auditory allocentricity by 40%, removing interaural cues at test time increased allocentricity. Moreover, with visual input fixed, making the prediction target heading-invariant or heading-dependent moved allocentricity up or down. With both senses, vision resolved the positional ambiguity created by rearranging the speakers without raising allocentricity, and removing vision at test time could raise or lower allocentricity, depending on the geometry of the signals the network had to predict. The reference frame of a learned spatial representation is therefore set by how its afferent signals depend on heading, not by their modality. For predictive models, this suggests that latent invariances are inherited from the geometry of inputs and targets and cannot be read from decoding accuracy alone; for the brain, that place field directionality should track how the cues an animal relies on depend on heading.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.