acceptodds
Under review as a conference paper at ICLR 2027

ACRE: Action-level Credit Assignment with Rubric Evaluation for Long-Horizon Agents

Abstract

Long-horizon agentic tasks require reinforcement learning (RL) methods to optimize both local decisions and final task completion. While group-based RL methods provide an effective framework for learning from sparse outcome rewards, it remains difficult to determine which intermediate actions contribute to success or failure. Recent stepwise policy optimization methods improve credit assignment by introducing intermediate signals, but these signals are still derived from downstream consequences, leaving local decision quality entangled with future outcomes. We introduce Action-level Credit Assignment with Rubric Evaluation for Long-Horizon Agents (ACRE). Rather than estimating process advantages from outcome-derived returns, ACRE uses reusable rubrics for different functional action families to turn semantic evaluations of intermediate actions into process rewards. These rewards are mean-centered within matched environment-state groups to form relative process advantages, which are combined with the episode-level outcome advantage for policy optimization. In this way, ACRE provides a direct step-level learning signal while preserving episode-level supervision for task completion. Across ALFWorld and WebShop with Qwen2.5-1.5B/3B/7B-Instruct, ACRE consistently improves over return-based process advantage estimation. We further show that rubric-based outcome evaluation can approach verifiable environment rewards, suggesting a practical direction for agentic RL in more open-ended settings where deterministic outcome verification is unavailable.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.