From Activation to Specificity: Automating Counterfactual Testing of Visual Representations in the Human Brain
Abstract
Identifying which brain regions represent a visual concept in the human brain is a central neuroscience challenge. Existing approaches have localized coarse functional regions (e.g., faces, places) through activation maximization, identifying regions that activate strongly for a target concept relative to others. Yet strong activation alone does not establish that a region represents the concept itself, as responses may instead be driven by correlated visual or semantic cues. We introduce BrainTRACE (Testing Representations through Counterfactual Evidence), an automated framework that combines generative and brain models to synthesize controlled stimuli and test candidate neural representations through targeted counterfactual-specificity testing. Given a query specifying a concept, our framework constructs targeted stimulus sets comprising concept images, counterfactual edits that remove the target concept while preserving other image content, and images with correlated distractors. It then uses an image-to-fMRI encoding model to predict brain responses and searches for representations that respond specifically to the target concept over correlated alternatives. BrainTRACE returns candidate representations supported by counterfactual-specificity evidence and proposes follow-up fMRI experiments to further test or extend these findings. Our approach recovers known functional localizations and identifies new candidate representations across dozens of concepts, supported by both predicted and measured fMRI data. Critically, we show that without counterfactual evaluation, a large fraction of localizations would be identified as false positives by our counterfactual-specificity testing, suggesting that activation alone is insufficient evidence of representation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.