acceptodds
Under review as a conference paper at ICLR 2027

Connecting Brain Activity to Language through Mixed-Resolution Visual Representations

Abstract

Brain-to-text decoding aims to describe viewed scenes from fMRI recordings. Pretrained vision-language models provide rich visual and linguistic priors, and their visual representations determine the spatial detail that brain decoders are trained to predict. Inspired by the non-uniform allocation of visual information in human perception, we propose MindMRQ, which uses adaptive visual granularity as a bridge between cortical representations and language models. Its mixed-resolution visual representations retain finer detail in regions sensitive to spatial compression and encode surrounding context more coarsely while maintaining full image coverage. We then connect brain activity to these visual targets through dual-level visual-semantic alignment, combining region-level matching with query-level semantic supervision to support language generation from brain-predicted representations. Extensive experiments demonstrate competitive decoding performance, with controlled comparisons and ablations supporting the effectiveness of the proposed representations and alignment strategy.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.