PACE: POLICY-NATIVE ADAPTIVE DECISION TIMING FOR LONG-HORIZON REASONING
Abstract
Long-horizon reasoning requires deciding not only what actions to take, but how many to execute open-loop before replanning. This number—the execution depth—balances replanning cost against compounding execution errors. Most current systems either fix execution depth as a hand-tuned scalar or adjust it at inference time using heuristic rules decoupled from the policy; both can be suboptimal. We treat execution depth as a learnable, history-conditioned variable of the policy itself and propose PACE, a model-native vision–language policy that jointly predicts what to execute and for how long under a hard decision budget. We evaluate PACE on four long-horizon environments spanning full and partial observability and visual and textual modalities: Sliding Puzzle, Sokoban, ALFWorld, and ScienceWorld. Across all seven evaluation settings, PACE improves success rates by 3.1–15.6 percentage points while using fewer decisions on average, achieving empirical success–decision Pareto dominance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.