acceptodds
Under review as a conference paper at ICLR 2027

TIME-UNIFORM GUARANTEES FOR TRANSFORMER EARLY EXIT VIA MARTINGALE STOPPING

Abstract

Can a transformer exit early and still guarantee that its prediction stays close to the full-depth answer? Existing early exit methods (entropy thresholds, learned confidence scores, patience counters) offer no such guarantee, and in practice collapse unpredictably: CALM loses 14–19 MCC points on two of six GLUE tasks under identical training conditions. We introduce *Martingale Transformers*, which connect early exit to martingale theory. The key observation is that a transformer's layer-wise predictions form a stochastic process amenable to concentration analysis. We first show this process is non-martingale in pretrained BERT (, all 12 layers, six GLUE tasks), then add a zero-parameter regularizer that enforces the martingale property; stopping is simple, exiting when the quadratic variation of the belief sequence drops below a threshold . Our main theoretical result is a time-uniform deviation bound via Ville's inequality: the early exit prediction satisfies simultaneously at all layers , with probability , which we verify empirically at on all datasets. On BERT-large (24 layers, 5 seeds), QV stopping matches or improves baseline MCC on 5 of 6 GLUE tasks while saving 5–9% of FLOPs; on TinyLlama-1.1B it degrades gracefully to full-depth inference (98% MCC retention on MNLI) rather than producing incorrect early exits.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.