acceptodds
Under review as a conference paper at ICLR 2027

Rational Agents Resist Entropy: Bounded Rationality as the Principled Objective for RL under Finite Samples

Abstract

Reinforcement learning optimises expected return without explicitly constraining representation complexity, which can make learning statistically difficult under finite samples. We propose bounded-rational RL: augment the return with a penalty on the differential entropy of the agent's noisy representations to control statistical complexity. We justify this objective by proving, under explicit conditions, an entropy-dependent bound on the samples needed to evaluate nearby policy changes. Analysing the policy gradient of this objective provides a behavioral characterisation: the complexity penalty supplies an intrinsic reward for low-surprisal representations and can drive the agent to maintain recurring, predictable situations—entropy resistance. This behavioral characterisation distinguishes RL from passive supervised learning and connects bounded-rational RL to active inference and related frameworks that also produce entropy-resisting behavior from different theoretical foundations. Experiments across partially observable control tasks show both: bounded rationality improves mean returns over unconstrained learning, and agents resist entropy, with larger gains in the more demanding task variants studied.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.