A Depth-Indexed Sample-Complexity Exchange Rate Between In-Context Learning and Low-Rank Fine-Tuning
Abstract
In-context learning (ICL) and low-rank (LoRA) fine-tuning are rarely compared on equal footing. We place both on one statistical model, a rank- linear map estimated from demonstrations or fine-tuning examples, and ask how many demonstrations match an -example fine-tune at depth . The answer depends on what attention can store. Any in-context estimator that is linear in the labels, including gradient-descent linear attention and any head bottleneck, needs asymptotically at least demonstrations per fine-tuning example, at every depth. With register tokens, the role played by system prompts, soft prompts and memory tokens, a fixed weight-tied block implements factored gradient descent exactly, so ICL and LoRA become the same algorithm. The exchange rate then becomes depth-indexed: it diverges below a critical depth and reaches above it, and we prove a two-sided escape-time bound whose constants are independent of the initialization scale , which the registers set. Task families with effective parameters lower the achievable rate to , and pretrained registers shift by roughly without removing the transition. Trained linear transformers confirm the separation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.