acceptodds
Under review as a conference paper at ICLR 2027

Decision Confusion: Learning More Can Decide Worse under Weak Signal

Abstract

Expressive learned policies are increasingly used in sequential decision-making because they can discover complex representations and decision rules directly from experience. However, when decision-relevant predictive signals are weak relative to noise, additional flexibility may increase finite-sample estimation error without providing commensurate gains in exploitable structure. We investigate how flexibility in policy learning affects performance under weak signals. We argue that the appropriate extent of learning is governed by predictive information: under weak signals, additional flexibility induces decision confusion, making some structure better left unlearned. Specifically, we demonstrate that, in environments characterized by high noise and bounded drift, a restricted rule can outperform an expressive learner, and that increasing the sample size may not eliminate this advantage. Using a signal-to-noise ratio (SNR)-controlled Heston-style environment, we show that learning only a selector over fixed rules matches the strongest fully learned policies at low SNR, whereas end-to-end reinforcement learning catches up only at high SNR. Building on this finding, we propose Diverse-Cue Hierarchical Reinforcement Learning, a method designed for real financial markets. Experiments on eight indices from the Shanghai and Shenzhen stock exchanges show that our method attains the best total return on six of the eight markets and the best Sharpe ratio on seven. This paper offers a new perspective on architecture design under weak signals: the granularity of learning should scale with predictive information rather than model capacity.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.