Instant Context Optimization: Compiling Few-Shot In-Context Demonstrations into Pre-Trained Channel Magnitudes via Circuit-Bounded Sparse Optimization
Abstract
In-Context Learning (ICL) enables Large Language Models (LLMs) to perform tasks from a small set of demonstrations (e.g., ), but requires persistent Key-Value (KV) cache expansion and quadratic prefill latency on every forward pass. While compiling demonstrations directly into model weights would eliminate this runtime prompt tax, standard few-shot parameter optimization () is fundamentally ill-posed, causing severe overfitting and representation collapse. In this work, we propose Instant Context Optimization (ICO), a framework that compiles demonstrations into task-specific static weights once per task offline, recovering the performance benefits of ICL without serving-time prompt costs. Formally framed as episodic prompt-to-weight internalization via circuit-bounded test-time distillation, ICO leverages three structural constraints exposed by the model's own in-context activation routing: (1) Where to update: Gradient-based parameter saliency localizes updates to sparse, task-relevant subcircuits (); (2) How to update: Updates are restricted to channel-wise gain modulation along fixed directional bases (), provably preserving the layerwise kernel nullspace (); and (3) How to score updates: Reasoning trajectories are supervised via Demonstration-Conditioned Teacher Reference Scoring with Group Minimum-Anchored Relative Advantage (). Across 38 benchmarks on models up to 27B parameters, ICO recovers (up to on GSM8K) few-shot headroom with a universal grand macro served gain in pure zero-shot setting, safeguarded by an empirical Risk-Monotonic AR validation gate.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.