acceptodds
Under review as a conference paper at ICLR 2027

Deep Delta Learning

Abstract

Transformer residual streams are updated by addition. A sufficiently expressive residual block can represent content replacement, but the residual update itself has no operation that reads, compares, and replaces content. We introduce Deep Delta Learning (DDL), which applies the delta rule over network depth. Each layer reads the residual state along a learned direction, compares the readout with a learned target, and writes a gated rank-1 correction back along the same direction. A closed gate gives the identity map, and a unit gate overwrites the selected readout with the target. DDL works with the usual vector state or with an expanded state that stores several value channels, while attention and MLP blocks keep the original model width. We pretrain decoder-only models with approximately GPT-2 small and medium sizes on the FineWeb-Edu dataset. At both scales, every DDL variant has lower validation loss and higher average one-shot accuracy than the additive baseline, and the best expanded variants raise that average by 0.91 and 1.18 points.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.