BiSURE: Toward a Unified Neural Representation for Visual Encoding and Decoding
Abstract
If visual encoding and decoding truly meet in the same interchangeable latent representation \(Z\), then a neural response from one participant should be readable, through \(Z\), as the response of another. We build and test such a representation. **BiSURE** (**B**idirectional **SU**fficient **RE**presentation) maps images and EEG, MEG, and fMRI responses into a common latent, from which the stimulus can be retrieved and every participant's measured response predicted, while retaining private interfaces for each instrument and participant. A single consistency term ties the image-derived and brain-derived routes, making their latents interchangeable. A response from one participant, passed through \(Z\) to another participant's response head, predicts that participant's response at 87–96% of the level reached from the image itself, with no stimulus and no parameter fitted for the pair, and more accurately than a linear map fitted for that pair on paired data. The same latent carries EEG responses into MEG at 84%, despite no paired EEG–MEG recordings, and transmits more than the model's own decoded CLIP features. Without the latent tie, retrieval remains essentially unchanged, yet the two routes become unrelated and translation falls to at most 6%; the tie also improves encoding by 15–31%. The same model matches or exceeds the strongest specialist in retrieval on every instrument, predicts every participant's responses above an established encoding baseline, and reconstructs images from the frozen latent.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.