Learning Visual Corrections from EEG
Abstract
EEG recorded during visual perception carries information about viewed objects and their attributes, offering a basis for neural guidance of image editing. The challenge is to translate this information into precise semantic and appearance changes while preserving unrelated source content. We introduce EEGRAFT, which learns to align neural and visual representations to guide image editing from EEG. The editor’s source-conditioned visual states query the aligned EEG summary and unpooled features through separate attention pathways, adapting neural guidance throughout generation through time-gated residual updates. At inference, the editor requires only a source image and EEG recorded while viewing the target, without access to the target image, text instructions, or edit masks. To support systematic study of this task, we construct Neve, a benchmark covering appearance and semantic editing with 29,275 source–target pairs and 2.71 million EEG trials, by generating source images while preserving the measured correspondence between the original viewed stimuli and their EEG recordings. Our method outperforms six baselines across all six aggregate metrics, reducing C-LPIPS by 22.7% and increasing the recovery–preservation score R-Edit from 0.386 to 0.585 relative to the strongest baseline for each metric. With the editor and source image fixed, correctly paired EEG produces more accurate edits than mismatched or disabled EEG conditioning. We hope this work advances the use of neural signals for visual content creation and interaction.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.