acceptodds
Under review as a conference paper at ICLR 2027

HiRUB: Hierarchical Rubric Rewards for Long-Horizon Large Language Model Agents

Abstract

Long-horizon agent reinforcement learning (RL) relies on informative trajectory-level rewards that indicate how a complete interaction advances the requested outcome. Success in these tasks can depend on a chain of interdependent decisions, with different actions contributing unequally to the eventual result. Consequently, terminal feedback alone is often too coarse to distinguish useful progress from ineffective behavior. LLM judges can provide richer trajectory assessments, yet placing them inside the RL loop requires repeated inference for every rollout and becomes costly as trajectories and rollout groups grow. To make rubric-style feedback practical without retaining this online expense, we present , a hierarchical rubric reward that uses an LLM to generate task-aware criteria before training and applies them through executable checks during RL. These criteria are related to a task schema, allowing each complete rollout to receive a task-aware score that combines partial progress, semantic constraints, successful completion, and penalties. This design keeps the reward computation lightweight and compatible with different RL backbones while avoiding online LLM judging. Experiments across ALFWorld and WebShop with 1.5B and 7B parameters show consistent gains over the evaluated RL baselines. In a matched ALFWorld comparison against an online LLM-judge reward, HiRUB improves validation success from 90.1% to 92.2% and reduces total training time from 31.1 to 19.6 hours. Together, these results show that task-aware trajectory scoring can improve agent learning without placing judge inference in the training loop.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.