acceptodds
Under review as a conference paper at ICLR 2027

DIHAS: Estimating Conditional Action Contributions for Reinforcement Learning in Trade Execution

Abstract

Executing a fixed parent order requires balancing immediate trading costs against exposure to future prices. Because final cost also depends on subsequent decisions and market innovations, it provides limited evidence about the contribution of one current action. We introduce Data-Induced High-Dimensional Action-Conditional Ambiguity Suppression (DIHAS), an action-conditional evidence layer that compares a candidate option with frozen-reference continuation under shared market innovations. A local financial surrogate and bounded residual estimate the resulting paired saving; a paired inventory control variate improves conditional-mean estimation, and ambiguity aggregation evaluates actions under alternative conditional price means. This evidence supplies bounded action biases to DDQN, PPO, and GRPO, while a held-out calibration diagnostic audits overstatement. Full-model training uses a verified terminal paired-saving reward, and final evaluation uses transaction cost. On held-out test dates within each year of the fixed 52-stock Shenzhen replay panel (2019–2021), all nine DIHAS configurations achieve lower three-year mean test-set transaction cost than locally retrained DQN and PPO baselines and TWAP. Diagnostics on Dev, an internal training diagnostic split, show lower mean trace variance with PICV and lower mean pinball loss with the residual across three algorithms. These retraining comparisons use each variant’s own probes; ambiguity and policy-cost effects remain mixed.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.