acceptodds
Under review as a conference paper at ICLR 2027

CALIBRA-fMRI: Text for Signals Without Captions, Read by a Frozen Video-Language Model

Abstract

Multimodal language models answer many kinds of question in one output space, and extending them to scientific signals requires text that rarely exists and must be generated. Generated text can often be answered without the non-text input, a shortcut usually found only after a model has been trained on it. Task fMRI comes with a controlled experiment and rich metadata, from which answers can be computed and the shortcut controlled at construction. We build CALIBRA-fMRI, whose items span the experimental design, the signal, behavior and the participant in a design space defined by a rule-based item generator. Each item carries a signal-free accuracy, fixed before training by a rule that answers it without the recording, and a model trained on the text alone stays at this level (.252 against .259). Without paired fMRI–text data at pretraining scale, we first verify that the image encoder of a frozen video-language model, never trained on fMRI, is competitive with fMRI foundation models under a linear probe. Building on it, we develop an fMRI aggregator that pools each four-dimensional volume through a cortical atlas, compressing a 15-minute recording from millions of patch tokens into thousands. Coupled with the frozen language model, it forms RIVET, which takes fMRI inputs of any length while training only 21.2M parameters. Two frozen fMRI foundation models with trained heads each answer one side, relations within a segment or participant identity, whereas RIVET answers both when trained on each and, trained on all levels at once, exceeds the signal-free accuracy on 9 of 14 items. It orders states in time without per-step labels and, from three segments on, learns to compare segments given separately, a comparison it carries to unseen paradigms. Trained on sex instead, it retains much less of the moment-to-moment state, consistent with how the two foundation models split as released, so what a model can answer depends on what its input keeps and what its items ask. The construction needs only a record that fixes the answers, offering multimodal language models a way to build corpora that require their non-text input by design.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.