acceptodds
Under review as a conference paper at ICLR 2027

When Do Structured Prompts Help In-Context Learning? A Relevance–Redundancy Theory of In-Context Prediction

Abstract

In-context learning is commonly analyzed under independent demonstrations, yet real prompts are often structured: demonstrations can be correlated with one another and with the query. When does this structure improve in-context prediction, and when does it merely make the context redundant? We develop a relevance–redundancy account in the solvable setting of in-context linear regression, allowing an arbitrary positive-semidefinite, unit-diagonal sample-similarity kernel over the demonstrations and query. For a linear-attention predictor trained on independent prompts and deployed on a structured one, we derive an exact finite-dimensional excess-risk law. Structure decomposes into a redundancy cost among demonstrations, a relevance gain between demonstrations and the query, and a higher-order mediated interaction. In the weak-structure regime, prediction improves precisely when , where is the context-to-dimension ratio and is the label-noise variance. We then characterize which relevance–redundancy combinations arise from coherent structure. Under graph-induced covariance, the two terms become internal context connectivity and query connectivity. In a sparse stochastic block model, a single latent context-composition variable changes relevance linearly but redundancy quadratically, yielding explicit homophily- and heterophily-dependent help–hurt phase transitions. Finally, trained full linear attention closely follows the exact law, while nonlinear softmax attention often depends on the same coordinates but with architecture-dependent coefficients. Thus, relevance and redundancy describe a shared prompt geometry, while their exchange rate is model-specific.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.