acceptodds
Under review as a conference paper at ICLR 2027

When Do Graph Neural Networks Help in Cross-Sectional Stock Ranking? A Multi-Seed, Multi-Universe, Cost-Aware Study of the U.S. S&P 500

Abstract

Papers often report that graph neural networks (GNNs) improve cross-sectional stock ranking by encoding relations among stocks as graph edges. Many reported gains, however, come from single-seed, single-split evaluations that do not adjust for multiple testing. We ask a narrower question: when, if at all, does the graph help? We evaluate an equal-budget tuned ladder from LightGBM to a heterogeneous relational GNN (L0–L7) over twelve expanding quarterly walk-forward folds and ten seeds on the U.S. S&P 500, with two pre-registered confirmatory families. The predictive family asks whether any tuned model beats tuned LightGBM. Hansen's SPA test does not reject in either feature universe ( and ), and the benchmark comparisons are underpowered, so the result is a non-rejection and does not establish equality. Local DM contrasts place the observable signal in the non-graph MLP, and the best-supported contrast runs *against* the correlation graph. Adding a correlation-GAT to the MLP reduces IC (); this penalty is BH-significant in *both* feature universes, and its sign holds when we drop any single quarter. The penalty is specific to that graph at its tuned operating point: in Universe C the sector, full-attention, HATS, and combined-edge arms all recover above the correlation-GAT, though Universe-C positives rest on a leak-selected basis. The MLP's gain over LightGBM () is significant only in the leak-selected universe and falls below its detectable-effect threshold, so we read it as suggestive. The edge-attribution family freezes the tuned correlation-GAT operating point and varies only the edge set. None of six edge contrasts survives Benjamini–Hochberg control, and all six are underpowered. A transaction-cost crosswalk (each IC contrast re-expressed as net Sharpe at 10 basis points) points the same way. The two Universe-C ladder contrasts we examine most closely (MLP - LightGBM and correlation-GAT - MLP) survive 10 bps; the tuned news-edge IC underperformance is near zero in net Sharpe, fragile across folds, and absent in the fixed-operating-point test. We therefore report conditional findings and failure modes.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.