acceptodds
Under review as a conference paper at ICLR 2027

What Do Hypergraph Benchmarks Measure? Feature Paths, Evaluation Protocols, and Modular Constraints

Abstract

Hypergraph neural networks (HNNs) are compared by node-classification accuracy on a fixed set of benchmarks, and the differences are read as evidence about their aggregation operators. We test that reading from two sides. First, we ask what the benchmarks measure. Ablating the released ED-HNN code at its published configurations, and a factorial over datasets, hyperedge constructions and architectures, show that accuracy depends mainly on whether a node's own features reach its output — often through a single preprocessing line — while the aggregation operator moves it by under two points. Second, we ask what the benchmarks cannot measure. We build HyperCSP, constraint hypergraphs with known answers, a majority-class floor and a non-neural reference, evaluated so that input labels cannot be copied and no test label touches model selection. On modular constraints, additive aggregators with linear messages solve none of 720 runs, while an exclusive-product aggregator (ProdSet) solves 149 of 240. Controls locate the difference in a receiver-specific interaction among the other members: the same product computed from additive power sums performs like ProdSet, averaging inside the same pipeline fails, a receiver-exclusive Deep Sets learns the binary constraint but rarely the ternary one, and every additive arm learns a cardinality constraint. A naive evaluation protocol hides this picture, overstating ProdSet and erasing the cardinality result. Carried back to seven benchmarks, ProdSet is never better than mean aggregation, even under a protocol in which it recovers parity constraints injected into a real hypergraph; on LDPC decoding, where such constraints exist, it recovers much of belief propagation's gain and additive arms none. Under the probes and protocols we test, the capability that separates these operators is one the benchmarks do not exercise. We release HyperCSP, the protocol and all code.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.