Does the Label Space Need to Evolve? Bounding the Gains of Self-Improving Context Engineering for In-Context Classification
Abstract
Can a classifier improve by rewriting its label space during inference? We study this question with EvoLabel, which separates a prediction architecture from an optional evolution layer. The predictor compares two retrieval views over a fixed labeled pool, returns immediately when their canonical predictions agree, and invokes a configured arbiter otherwise. A coverage gate avoids the second view when leave-one-out retrieval is unreliable. The evolution layer aggregates prediction-pair feedback and may propose structured Merge Describe updates to the presented label state, without changing model weights or adding labeled examples. In a 24-method, 12-dataset, three-seed comparison, the adaptive retrieval–arbitration configuration reaches mean accuracy, versus for the strongest external baseline; on four many-class intent datasets the corresponding result is versus . The same configuration exceeds the frozen dual-arm baseline on three additional models by – percentage points. On the intent tasks, agreement covers of queries and identifies a substantially more accurate subset ( versus on disagreements), although this contrast is descriptive rather than a calibrated risk guarantee. Crucially, the evolution-enabled runs in the replication commit no label-space edits: the observed gains support the prediction architecture, not successful online evolution. A conditional decomposition expresses rewrite gain as affected-query mass times conditional accuracy change under an explicit outside-subset invariance assumption. Together with a frequency-matched operator audit, this explains what the null does and does not establish. The study provides an audited, trace-backed comparison and a bounded claim: complementary views and arbitration help on this suite, while unseen-dataset validation and more effective rewrite operators remain open.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.