In-Context Task Learning: Reading Out a Kernel Classifier
Abstract
Task learning in context requires using demonstrated mappings when pretrained knowledge and other task cues do not determine them. We investigate whether small sets of attention heads can support direct prediction and reuse of demonstration evidence. Starting from Llama-3.1-8B on TREC-fine with 180 demonstrations, we identify mapping-sensitive carrier heads through paired intact/randomized-label contrasts. Their writes support a kernel classifier based on query-dependent retrieval of contextual class votes, characterized through query-weight controls and value interventions. This account motivates Selective TL Memory, which caches demonstration-side states while preserving live retrieval and direct readout. Across models, natural-language tasks, and demonstration budgets, thirty-head full-access readouts match or exceed final-model accuracy in 12 of 18 settings. Selective access to additional demonstrations improves a fixed carrier readout in all 18 settings at both tested head budgets; thirty-head selective readout exceeds every adapted vector baseline in 14. Contextual label states often support compact reuse, while access and receiver controls show that the selected readouts depend on states constructed by the surrounding network. In a prefix-cached benchmark, shared memory supports repeated queries with latency close to the smaller prompt and reduced resident storage relative to full access.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.