acceptodds
Under review as a conference paper at ICLR 2027

A DEPTH–WIDTH SEPARATION FOR QUADRATIC IN-CONTEXT LEARNING WITHPROTECTED-LABEL BILINEAR TRANSFORMERS

Abstract

Quadratic in-context regression has task coefficients for inputs in , although each prediction is a scalar. Can depth reduce the token width needed to predict these tasks? We study quadratic regression with Gaussian inputs using a bilinear Transformer whose attention can access labels but whose feed-forward maps cannot read or modify them. For one or two blocks at width , the minimum achievable relative risk converges to one, uniformly over context length. Three blocks suffice at width : an explicit network reaches relative risk at most with context examples. Thus the minimum width for fixed relative risk below one falls from to between two and three blocks. The construction evaluates the quadratic projection-kernel estimator without storing all quadratic features. Attention applies a context-dependent matrix to the query, and a later block completes the prediction using -dimensional states. This shows that sequential computation through context can replace explicit storage of a quadratic feature representation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.