acceptodds
Under review as a conference paper at ICLR 2027

Spec-CoFrGeNet: Reparametrizing Continued Fraction Networks for Stable and Accurate Training

Abstract

Continued fractions have been shown to be a function class capable of constructing expressive yet parameter-efficient neural networks. These neural networks are not only effective for supervised learning, but also for training large language models (LLMs). However, the rational form of this function class introduces poles when the denominator approaches zero, which prior work has avoided by enforcing a small positive (non-smooth) lower bound. We show that performance and training stability can be highly sensitive to the choice of this lower bound. In this work, we introduce a novel transformation of the partial denominators, inspired by its connection to the Green's function in physics, that mitigates the above issue elegantly by constraining the spectra of the constructed tridiagonal Hamiltonian. Our approach further eliminates the need for incremental training used by previous works for stability, thus avoiding involved modifications to standard pre-training procedures. We demonstrate the effectiveness of our reparametrization by replacing FFNs with a previously proposed continued fraction architecture for both language and image generative modeling. Across diverse models, such as GPT-2 XL, (nano) Llama-MoE, and Granite, we show that our proposed method achieves performance comparable to the original (FFN-based) architectures while achieving %-% parameter savings, and in most cases even outperforms the more complicated incremental training baseline. We further demonstrate, for the first time for continued fraction networks, that their effectiveness also carries over to image generative modeling, that utilize diffusion objectives. On a state-of-the-art Diffusion Transformer, we show that our approach again achieves performance comparable to the original FFN-based models with % parameter savings and significantly outperforms the previous incremental training approach.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.