Behavioral Tokenization via Maximum Entropy Mixture Policies with Minimum Entropy Components
Abstract
Standard continuous-control learning in embodied agents entangles movement with reward, producing task-specific policies but no explicit action vocabulary that another controller can directly reuse. We consider the problem of behavioral tokenization: how to learn a fixed, compact bank of state-dependent motor behaviors that becomes a shared discrete interface for future tasks. Previous work has studied this problem using action quantization, sub-policies, options and skills, but no solution provides an unsupervised, task-agnostic method to discover short-lived, action-spanning behavioral tokens. We introduce an online unsupervised framework that discovers behavioral tokens via a joint entropy objective—maximizing the cumulative entropy of an overall mixture policy to ensure rich, exploratory behavior under the maximum occupancy principle, while minimizing the entropy of each component policy to enforce diversity and high specialization. We prove convergence of the value function under our policy-iteration algorithm in the tabular setting, and then extend it to continuous control by fixing the discovered components and deploying them deterministically as quantized actions within an online optimizer to maximize reward. Experiments demonstrate that our maxi-mix-mini-com entropy-based tokenization and action quantization provide reusable behavioral tokens that can be sequenced to achieve competitive performance against continuous-action controllers, largely outperform random action quantization and skill discovery methods, and can be fine-tuned to solve downstream tasks showing superior performance to state-of-art off-policy controllers.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.