acceptodds
Under review as a conference paper at ICLR 2027

Learning Decision-Stump Thresholds in Context: Dynamics of Softmax Attention

Abstract

Estimating a decision threshold requires locating observations near an unknown boundary. We study how gradient-based pretraining learns this statistical rule in a two-parameter softmax-attention model with a fixed feature and inequality direction. Pretraining uses labeled contexts and their true thresholds; a fresh threshold must be inferred from context alone. Under a large-resolution initialization, constant-step gradient descent on tasks with examples each produces a frozen estimator with error for each fixed interior threshold and every fresh-context size . The two terms separate finite-pretraining accuracy from fresh-context localization. The mechanism is coordinated parameter divergence: population training calibrates the relative label and feature scores, then increases the attention scale as , giving population threshold error . To transfer this mechanism to a fixed finite corpus, we control gradient errors relative to the shrinking directions of progress at successive parameter scales. This certifies a growing training interval without requiring long-time tracking of the population trajectory. We also identify the boundary limitation of the one-head model and explain statistically what a reflected symmetrization could achieve.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.