acceptodds
Under review as a conference paper at ICLR 2027

Gated Successor Calibration for Solver-Guided LLM Agents

Abstract

Game solvers provide turn-level feedback for training LLM agents when task outcomes are sparse. However, a reduction in solver cost does not reveal whether the chosen action is better than the alternatives available in the same state. Solver feedback can also continue to shape updates when every rollout in a sampled task group succeeds. We propose Gated Successor Calibration (GSC), a method that calibrates solver feedback against local alternatives and gates its use by group success. Given a state, GSC subtracts the mean solver progress of distinct effective successors from the chosen action's progress. Intuitively, the same progress deserves more credit when alternative actions make less progress than when they make comparable progress. For all-success groups, GSC removes the solver-derived process term while leaving outcome credit and the policy objective unchanged. Across Sokoban, Rush Hour, and Minesweeper, GSC outperforms a strong solver-guided reinforcement learning baseline and also exhibits zero-shot transfer to ALFWorld and WebShop when trained only on game tasks. Specifically, with Qwen3-4B-Instruct-2507, GSC achieves average in-domain success and overall zero-shot success 6.8 and 1.1 percentage points higher than this baseline, respectively.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.