acceptodds
Under review as a conference paper at ICLR 2027

Reinforcement Learning for Brain Image Captioning by Imposing Visual Posterior Matching

Abstract

Decoding visual semantics from brain activity can provide a natural-language account of perceived stimuli. While existing approaches mainly focus on image classification, retrieval, or reconstruction, brain image captioning provides a compositional readout of objects, attributes, and relations. However, powerful vision-language priors can turn incomplete and noisy neural evidence into plausible yet stimulus-inaccurate descriptions. We propose Visual Posterior Matching (VPM), a reinforcement learning framework for brain image captioning. VPM first aligns neural representations with image tokens to initialize a brain-conditioned language policy, then trains a conditional flow model on paired image and caption features. During reinforcement learning, VPM constructs shared noisy anchors from the paired stimulus and rewards agreement between the clean visual-feature estimates induced by generated and reference captions at the same anchors. This target-anchored reward is designed to keep generation tied to stimulus-specific evidence while allowing diverse phrasings. Experiments on fMRI and EEG show that VPM improves core captioning metrics over supervised initialization and outperforms existing brain-captioning systems on the evaluated NSD dataset.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.