LEARNING HOW TO BE RISK-SENSITIVE: STATE-DEPENDENT RISK FUNCTIONALS FOR REINFORCEMENT LEARNING
Abstract
Adapting risk preferences to the state expands the decisions a risk-sensitive agent can implement. We prove a constructive separation within a specified conditional value at risk (CVaR) family: a stationary state-dependent risk map strictly im- plements an outer-optimal policy beyond the implementation capacity of every time-dependent, state-independent map. We develop a learning procedure that selects risk maps through empirical planning and separate outer-utility evaluation before fitting a neural critic. This separates the value of a risk preference from the fidelity of its learned execution. Controlled experiments recover optimal util- ity and verify policy fidelity separately. In a complementary sequential model, adaptive risk exceeds the specified shared-risk baseline under both design and shifted transition laws, matching direct optimization with equal total data at the largest evaluated budget. Exact certificates verify that learned state-only maps implement the target policy under both laws. In financial risk management, a simu- lated currency-liquidation benchmark with hold-or-sell-all policies shows higher mean revenue than selected shared and time-dependent risk controls. Historical Bit- coin quote replay adds evidence of temporal transfer: a frozen risk map improves lower-tail proceeds over those controls in two later periods on Coinbase’s exchange (GDAX in the dataset). Together, these results connect the expressive power of state-dependent risk to a practical procedure for learning explicit valuation rules and verifying their execution.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.