acceptodds
Under review as a conference paper at ICLR 2027

What the Loss Never Asks For: Which Group a Recurrent Network Represents, and Which Words Break It

Abstract

Which finite-group state-tracking tasks a recurrent architecture can express is well characterised; which of several correct implementations training actually selects is not. We study this selection problem for LSTMs and input-switched linear RNNs trained on word problems over nine groups. In the linear models the state is a function of the running product if and only if the learned operators satisfy the group's defining relations on the states they reach; cross-entropy enforces them only as far as the margin requires at training lengths, and in ordinarily trained models they fail at order one. Identities that do hold, imposed by a penalty or by the architecture, present a group whose elements the state follows; training does not select the target group, which satisfies them all. Penalised on identities of the Klein four-group that also present an infinite group, models realise the infinite group: its word breaks them within a thousand tokens while its finite-order words survive, and breaks 1.96 times later (predicted 2). Penalised on identities of , they realise its double cover, whose sign pattern appears in 44 of 48 models and in no control; one reflection per token, the delta rule's transition at , breaks first on , the travelling word of the infinite dihedral group. For isometric recurrences we bound how a relator defect accumulates and separate eigenplanes whose phase error drifts from those too small to reach the margin; snapping the two drifting planes to their ideal angles removes the break on the targeted word. A squared-norm variant of MatrixNet's relational penalty lengthens periodic-word horizons by one to three orders of magnitude without making any model exact; where it stalls, its floors are local minima of the penalty, derived in closed form for . Instruments are calibrated on planted representations; we report all 19 pre-registered predictions, including the one that failed.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.