acceptodds
Under review as a conference paper at ICLR 2027

HiCA-RNN: Multi-Scale Credit Assignment for Long-Term Dependency Learning

Abstract

Windowing a recurrent computation changes how credit is transported; adding intermediate targets changes where credit enters. Neither is sufficient alone: reinitializing the state at each window boundary does not shorten the backward horizon unless the boundary is also detached, and intermediate targets supply local training signals without preserving long-range credit. In this work we introduce HiCA-RNN that combines both within a single fast–slow recurrence. A shared ungated cell processes each window, a gated slow state connects the windows, and a feedback recurrence supplies detached targets to intermediate fast states. The states remain connected across the full sequence: the task loss is trained by full-graph BPTT, while the target loss introduces credit directly at selected states. We derive the resulting adjoint and identify the contribution of each target to the shared-parameter gradient. The derivation exposes a constraint on the method: local credit is useful only when its magnitude, its transport through the slow dynamics, and its direction relative to the task gradient are favorable. Neither short windows nor orthogonal feedback ensures these properties. HiCA-RNN achieves on sequential MNIST and on permuted MNIST, exceeding tuned LSTMs on both, obtains the lowest Copy loss on all horizons, and matches or improves on burn-in LSTMs in eight of twelve system-identification settings. On sequential CIFAR-10 HiCA-RNN uses approximately less peak training memory than a parameter-matched LSTM, which achieves slightly better accuracy; LSTM also remains stronger on Adding.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.