Distributional Epistemic Robustification for Sequential Portfolio Optimization
Abstract
Portfolio reinforcement-learning systems must operate with policies frozen before future market regimes are observed, because return dynamics shift between historical training and deployment. Overcoming this requires separating ordinary market-outcome risk from uncertainty about which portfolio action remains valuable, rather than treating all predictive dispersion as one risk signal. Distributional value learning is well suited to representing stochastic portfolio returns, yet its within-model return law leaves epistemic ambiguity unresolved once historical data support multiple plausible value functions, especially in offline frozen-policy extraction. Within that subcase, distributional critic ensembles address model uncertainty through predictive disagreement, but aggregate dispersion can mix within-critic outcome variation with between-critic disagreement about expected action values and thereby penalize the wrong uncertainty. We therefore propose Distributional Epistemic Robustification (DER), which separates return variation from disagreement across learned value functions and yields exact structural guarantees for the extractor together with a conditional shift-aware bound. Experiments on a strictly time-ordered, cost-aware A-share benchmark show that DER attains the highest aggregate and worst-calendar Active ARR among the external baselines in the primary comparison. A matched uncertainty-target control shows that penalizing cross-critic expected-value disagreement preserves substantially more return than penalizing within-critic or total dispersion while remaining competitive with ensemble-mean extraction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.