acceptodds
Under review as a conference paper at ICLR 2027

Ember: A Kilobyte-State Optimizer for Token Interfaces

Abstract

Language models learn continuous programs over discrete symbols, with the embedding table and LM-head acting as the read/write interface between them. We show that this interface has gradient geometry distinct from dense hidden weights which can be exploited to improve the Pareto frontier across supervised finetuning, RL, and pretraining, while only utilizing kilobytes of optimizer state. We introduce Ember, a lightweight optimizer for embedding and LM-head matrices that utilizes optimizer VRAM, instead of Adam's , and forgoes the need to shard both token table optimizer states. On 50k-step pretraining across two datasets and three paired seeds, Ember lowers val loss by – nats relative to tuned AdamW; we also provide empirical evidence that Ember scales across batch size and parameter count. Finally, we provide a distributed Ember implementation that merges cleanly with existing ZeRO/FSDP setups (code to be released).

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.