acceptodds
Under review as a conference paper at ICLR 2027

When Can Behavioral Tests Distinguish In-Context Computations?

Abstract

In-context learning in transformers is often studied by comparing predictions with familiar algorithms. We ask when a benchmark can distinguish these computa- tional explanations. For linear regression, our test compares a model with the best possible output of a specified family of gradient-style updates, allowing all scalar settings to be chosen separately for each prompt. Lower model error excludes the whole family, rather than one optimizer configuration. On a released 12-layer transformer and one independently trained replication, shorter contexts make pre- dictions less accurate but provide stronger risk-based exclusions. With 22 ex- amples in 20 dimensions, both models outperform the best 12-update predictors, including in an exploratory analysis removing the hardest tenth of prompts. With forty examples, the same families are already more accurate than the models. An exact Gaussian residual law explains the comparator’s improvement: longer con- texts lower its minimum achievable error. A separate symmetry analysis distin- guishes useful departures from errors that depend on example mixing and feature coordinates; at forty examples, these dependencies account for 96–97% of error in the recovered regression coefficients. Prediction accuracy and the strength of behavioral evidence are different objectives: an easier prediction task can make a computational explanation harder to exclude.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.