acceptodds
Under review as a conference paper at ICLR 2027

HETEROGENEOUS RECURRENCE FOR PARTIALLY OBSERVED CONTROL

Abstract

Agents for partially observed control usually rely on a single kind of recurrence. Yet the best memory depends on the hidden part of the state. When the hidden quantity is a derivative of the observation, as a velocity is of observed positions, the masked system is observable. Each new observation then carries a correction signal, and a recurrence can recover the hidden quantity by reading the current observation beside its memory. When the hidden quantity is an integral, no observation corrects its initial error, and a plain gated recurrence does as well. One agent may face both cases. In this paper, we therefore propose DuoCell, a heterogeneous recurrence for partially observed control. It splits one fixed state budget between a structured prior-innovation path for the first case and a GRU for the second. On masked MuJoCo with velocities hidden, DuoCell leads a width-matched GRU by +0.52 normalized IQM with fewer parameters. It also beats eleven other baselines, including an LSTM and a tuned TransformerXL, and its lead over the GRU baselines holds at five times the training budget. Because the regime follows from the sensor layout, it can be read off before training. We fixed this prediction before running eight unseen environments, and it held in both regimes. Where the criterion expected a benefit, the structured recurrence gained +0.33 normalized IQM over the GRU.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.