acceptodds
Under review as a conference paper at ICLR 2027

Cross-Asset Graphs in Reinforcement-Learning Trading: Derivation, Validation, Decision-Relevance and Reporting

Abstract

Reinforcement-learning traders increasingly condition on a graph over assets and report gains from doing so. Two questions go unasked: nobody applies multiple-testing control to the thousands of candidate edges a graph is selected from, and nobody reports how much the graph changes what the agent does. This paper addresses both. For a mean-variance-with-costs reward under an approximate factor model , it proves that the reward decomposes exactly into per-asset terms, pairwise terms supported precisely on , and one rank-one global term, so the validated graph is not an architectural choice but the coupling structure the reward already has. It also bounds the value error by the false-discovery rate of edge selection and fixes the construction's scope by a third-order criterion. On 116 US large caps over 3,772 trading days that coupling changes a median of 0 of 116 actions, and forty pre-registered studies leave it intact, including two regimes built to revive it, a second architecture, and a published system whose released code was run. The mechanism is measurable and does not depend on the agent: the graph's information is contemporaneous at daily resolution, with out-of-sample of 0.0907 same-day against 0.0030 for a degree-matched random graph, but -0.0010 at one day's lag, for every structure tried. Same-day co-movement is a risk fact, and risk enters the decision marginal one order in position size below alpha and cost. That diagnosis names a remedy, which the paper tests: in an objective with no return term, the same edges as sparse precision support beat a density-matched random support out of sample by -10.42% in realised volatility. A pre-registered adversarial pass that nets out sector structure retains 48% of that effect and 27% on log-likelihood, and the retained figures are the ones claimed. This research recommends three numbers, which no graph-RL paper currently reports: the count of actions the graph changes, the fraction of decisions beyond any graph's reach, and the number of seeds the comparison can resolve (19,261 here against a cap of 100).

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.