Double Identity: Decoupling Control State and Task Belief in Reinforcement Learning
Abstract
In partially observable control problems with persistent hidden task properties, immediate control and task inference may benefit from different temporal contexts. Conditions relevant to the next action can change at every transition, while evidence identifying an unobserved goal or dynamics regime accumulates over longer horizons. Encoding both in a single representation creates conflicting demands: the representation must be highly responsive to local changes while preserving as much task-relevant information from prior interactions as possible. We introduce State-Task Adaptive Control (STAC), an actor-critic architecture that assigns these inference roles to separate causal encoders with different temporal contexts and learning objectives. A control encoder summarizes recent interactions and is trained through temporal-difference learning. At the same time, a task encoder attends causally to the full interaction history and is trained with goal supervision, reward prediction, and contrastive objectives. A trainable adapter maps the task encoder's output to belief features used by the actor and critics. Under a hidden-goal protocol using the ten Meta-World ML10 training tasks, STAC achieves mean evaluation success rates of 71% within the training region and 62% in held-out regions, exceeding the strongest evaluated baseline on each split by 20% and 18%, respectively. Architectural ablations further support limiting the control encoder to recent interactions while using a separate encoder to aggregate available task evidence. Together, these findings show that tailoring both temporal context and learning objectives to the distinct demands of control-state estimation and task inference can improve learning and generalization in the evaluated manipulation setting.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.