acceptodds
Under review as a conference paper at ICLR 2027

BrainBench: Benchmarking Large Language Models for Comprehensive EEG Understanding

Abstract

Large language models (LLMs) are increasingly being explored for electroencephalography (EEG) analysis, yet their ability to support comprehensive EEG understanding remains poorly understood. Existing evaluations largely focus on isolated decoding tasks or demonstrations tailored to specific systems, capturing only limited aspects of EEG competence. In practice, EEG analysis requires a system to interpret analysis instructions, operate on physiological signals, derive quantitative evidence, and translate the results into scientifically grounded conclusions. To systematically evaluate this broader capability, we introduce BrainBench, a unified benchmark for comprehensive, instruction-conditioned EEG understanding. It comprises four subsets—Foundational Analysis, Sleep Assessment, Neurocognitive Assessment, and Physiological Integration—covering 17 datasets, 172 tasks, and over 4K real-data instances. Given an instruction and EEG recordings with optional physiological signals, a system must perform the requested analysis and produce scientifically grounded, verifiable outputs. We evaluate 13 representative LLMs across more than 100K executions under two paradigms: autonomous code execution and structured agentic analysis. Results vary substantially across models, subsets, difficulty levels, and execution paradigms, demonstrating that EEG competence depends not only on the underlying model, but also on how it is operationalized. BrainBench advances the exploration of artificial intelligence systems that reason over neural signals and support scalable, reproducible discovery in neuroscience.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.