acceptodds
Under review as a conference paper at ICLR 2027

Design and Training Shape Component Reliance in Small Transformers

Abstract

Architectural choices can change which components a Transformer learns to rely on. We compare dense and mixture-of-experts (MoE) feedforward networks (FFNs) on three algorithmic tasks and character-level language modeling. MoE models retain higher accuracy after FFN removal on carry-based addition and TinyStories. This advantage persists in deeper models on both tasks and in wider TinyStories models. Controlled one-layer addition experiments show that increasing expert count at fixed total FFN width improves no-FFN performance and increases relative attention attribution, while intact accuracy remains near perfect. We identify restricted active capacity and evolving expert assignments as contributing factors. Critically, allowing assignments to evolve increases no-FFN performance even when router weights remain frozen. Dropout and shuffled assignments further show that training conditions can change reliance. The effect varies across tasks and training stages: the gap grows with continued training in one-layer modular addition, while histogram counting shows little separation. Fitted-readout comparisons distinguish recoverable answer information from its use by the trained model, helping interpret these differences. Together, these findings suggest ways to steer component reliance through FFN design and training.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.