acceptodds
Under review as a conference paper at ICLR 2027

Between Tokens and Trajectories: Event-Level Test-Time Optimization for LLM Reasoning

Abstract

Test-time scaling has emerged as an effective way to improve the reasoning ability of large language models. Recent work has begun to complement sampling and search with instance-level test-time optimization, using reward signals to refine a reasoning trajectory. However, representative approaches differ in optimization granularity: some revise one token at a time, while others optimize token decisions independently. As a result, it is difficult for them to use a reward computed over the full response to coordinate updates across dependent reasoning decisions. We instead treat a short span of discrete token decisions that are autoregressively dependent as a single reasoning event, allowing one reward signal to guide their coordinated revision. Based on this formulation, we introduce Adaptive Event Policy Gradient (AEPG). At its core is a temporary event policy, adapted with reward-guided policy gradients to coordinate the revision of the entire event. During this process, each revised token is fed back into the frozen language model, so subsequent decisions condition on the updated context. Across three instruction-tuned LLMs and five mathematical reasoning benchmarks, AEPG achieves the highest average accuracy among the evaluated test-time methods on all three backbones. Ablation studies further show that multi-token event revision is an important contributor to these gains, highlighting event-level optimization as a promising direction for more effective test-time reasoning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.