Auxiliary supervision shapes optimization geometry for recurrent sequence models
Abstract
Recurrent sequence-to-vector models are widely used for tasks such as time series forecasting and sequence classification or regression, where supervision is often applied only after the complete input sequence has been processed. This can create a temporal parameter-information bottleneck, leaving important parameter directions weakly informed and thereby causing poor conditioning, potentially slow and numerically sensitive training, and undesirable rollout behavior on unseen sequences. We show that auxiliary supervision can alleviate this bottleneck by exposing parameter information at intermediate positions before it becomes invisible or is attenuated by the recurrent dynamics. Our sensitivity-based analysis establishes that suitably chosen auxiliary observations can improve local parameter observability and prevent deterioration of generalized Gauss–Newton conditioning with increasing sequence length, thereby supporting faster and more numerically robust training. At the same time, auxiliary supervision acts as an implicit regularizer that can promote desired intermediate or rollout behavior. Moreover, we characterize the resulting primary-task performance trade-off and show that compatible auxiliary objectives need not sacrifice globally optimal primary-task performance. Analytical results and extensive experiments with modern Mamba models on the standard ListOps sequence-classification benchmark corroborate the identified mechanisms and demonstrate the practical benefits of appropriately chosen auxiliary supervision for training efficiency and primary-task prediction accuracy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.