Beyond Full-Context Accuracy: Context-Marginalized Evaluation of Test-Time Structure-Aware Inference in EEG-to-Image Retrieval
Abstract
Test-time structure-aware inference can improve cross-subject EEG-to-image retrieval by exploiting relations among unlabeled test queries, but it also makes predictions depend on query-set composition. Existing evaluations typically report accuracy under a single full-query context, leaving unclear whether gains and comparative conclusions remain reliable when the available context changes. We introduce Context-Marginalized Retrieval Evaluation (CMRE), which repeatedly evaluates each target query within sampled subsets of companion queries, applying each inference rule to the corresponding rows of a fixed query–gallery similarity matrix. CMRE distinguishes structural exploitability, quantified by gains over independent retrieval, from context robustness, measured by same-query Top-1 context consistency. Under leave-one-subject-out evaluation on THINGS-EEG2 across six frozen representations and six structure-aware inference rules, structure-aware inference improves full-context Top-1 accuracy by up to 20.40 percentage points over independent retrieval. However, the rule selected under full-context evaluation is no longer optimal in 17 of 24 representation–context-size conditions (70.8%). Selecting the full-context winner incurs an average loss of 2.70 percentage points and a maximum loss of 8.29 percentage points under reduced contexts. Representation rankings are likewise inference- and context-dependent: the ordering under independent retrieval is not consistently preserved under structure-aware inference. Moreover, larger retrieval gains do not necessarily imply greater context robustness, showing that retrieval effectiveness and prediction stability provide complementary views of structure-aware inference. Together, these results show that full-context accuracy represents only one operating point of context-dependent inference. CMRE complements conventional evaluation by characterizing structural gain, context consistency, and the context dependence of conclusions about inference rules and frozen EEG-to-image representations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.