How Reliable Are Reward-Objective Conclusions in Financial Reinforcement Learning?
Abstract
Financial reinforcement learning (RL) studies use materially different reward objec- tives, while separate work shows that repeated RL training can yield substantially different outcomes across random seeds. Neither observation is treated as new here. Instead, this study asks how reliable conclusions attributed to reward choice are when evaluated across independent stochastic retraining and different subsets of observed runs. Six reward families—portfolio-value change, log wealth growth, differential Sharpe ratio, rolling Sortino ratio, negative maximum drawdown, and Sharpe minus conditional-value-at-risk penalty—are evaluated with Proximal Pol- icy Optimization (PPO) on a controlled single-asset task using 10 matched seeds per reward. Omnibus tests do not detect a significant reward effect on annualized return, Sharpe ratio, Sortino ratio, or maximum drawdown; reward-associated differences are detected for turnover and long exposure, although most policies remain predominantly long. Counterfactual single-seed comparisons produce dif- ferent apparent reward leaders. Robust summaries reinforce that caution: replacing the mean with the median or interquartile mean (IQM) substantially compresses the apparent performance separation among rewards, and matched-seed bootstrap resampling does not yield a stable IQM leader. Exhaustive enumeration of matched seed subsets further shows that the non-significant return and Sharpe conclusions are stable, whereas the long-exposure conclusion becomes increasingly stable as more runs are included: 99.2% of seven-seed subsets and all subsets of eight or more seeds yield a significant Friedman reward effect. Turnover stabilizes more slowly. Reward rankings also become progressively more similar to the 10-seed reference as additional runs are included, while winner claims remain fragile when leading estimates are close. No significant performance reward effect is detected across three temporal windows or with a second actor–critic algorithm. The contribution is therefore a reward-specific conclusion-reliability analysis: the results support reporting full run distributions and directly examining whether conclusions attributed to reward choice remain stable under stochastic retraining
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.