acceptodds
Under review as a conference paper at ICLR 2027

PACE: POLICY-NATIVE ADAPTIVE DECISION TIMING FOR LONG-HORIZON REASONING

Abstract

Long-horizon reasoning requires deciding not only what actions to take, but how many to execute open-loop before replanning. This number—the execution depth—balances replanning cost against compounding execution errors. Most current systems either fix execution depth as a hand-tuned scalar or adjust it at inference time using heuristic rules decoupled from the policy; both can be suboptimal. We treat execution depth as a learnable, history-conditioned variable of the policy itself and propose PACE, a model-native vision–language policy that jointly predicts what to execute and for how long under a hard decision budget. We evaluate PACE on four long-horizon environments spanning full and partial observability and visual and textual modalities: Sliding Puzzle, Sokoban, ALFWorld, and ScienceWorld. Across all seven evaluation settings, PACE improves success rates by 3.1–15.6 percentage points while using fewer decisions on average, achieving empirical success–decision Pareto dominance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.