EMO: A Closed-Loop Test-Time Computation Framework via Alternating Context Optimization and Logits Steering
Abstract
Test-time context optimization adapts frozen language models through textual feedback, but ordinary conditional generation provides no separate control over how strongly to reinforce the distributional changes induced by a context update. We introduce EMO, a training-free framework that couples question-specific context optimization with logits steering. The E-step revises a system context using textual gradients from an external judge without access to gold answers. The M-step contrasts next-token distributions under the revised and preceding contexts at a shared decoding prefix, transforms this contrast into a steering vector, and applies a confidence-constrained logit correction. Feedback on the steered response drives the next context update, coupling adaptation across rounds with intervention within each generation. A logit-space analysis connects the steering vector to the context-induced logit displacement and identifies sufficient conditions for one-round target-KL improvement. Under greedy decoding across eight models and six benchmarks, EMO improves the zero-shot baseline in 39 of 40 aggregated settings (47 of 48 before merging the two AIME editions), averaging percentage points over the 40 model–task-group pairs. With the same judge, it adds points on average over context-only optimization, winning in 25 settings and tying in four, with MMLU-Pro gains across all eight models. At matched target-model generation budgets, accuracy advantages over best-of- are strongest on Llama, while oracle coverage is higher in four of six comparisons despite fewer candidates. The coverage–selection gap motivates allocating test-time computation jointly to candidate generation and verification.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.