acceptodds
Under review as a conference paper at ICLR 2027

Investigating FFN Parameter Allocation via Cross-Layer Sharing in Transformer LMs

Abstract

This paper studies feed-forward network (FFN) parameter sharing in Transformer-based language models. We share FFN parameters across a contiguous middle block and reallocate the saved budget to widen FFNs, retaining FFN computation in every layer. A broad search using models with 570M non-embedding parameters identifies many configurations with lower average downstream task loss than a standard Transformer baseline. Improvements concentrate in late-heavy configurations that retain more layer-specific FFNs in the late layers. Mirrored comparisons with identical parameter counts and FLOPs within each pair also favor late-heavy allocation on average task loss in most cases. The selected configuration improves average downstream accuracy at the 2B scale without retuning the sharing configuration. These results highlight the importance of where independently parameterized FFNs are retained and suggest that sharing and reallocation can improve downstream performance at a nearly matched parameter budget.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.