acceptodds
Under review as a conference paper at ICLR 2027

PROSE: Learning Progress from Cross-Trajectory Outcomes for Long-Horizon Embodied Tasks

Abstract

Long-horizon embodied tasks are difficult to learn from sparse terminal rewards, because an agent may need to act for hundreds of steps before receiving any feedback about its intermediate behavior. Existing approaches provide intermediate guidance by scoring individual observations based on their relevance to the task goal or by relying on predefined subgoals or dense progress supervision. However, goal relevance alone may be weakly predictive of eventual task success, and the effectiveness of subgoal-based guidance depends on whether they remain appropriate under the current environment dynamics and execution capabilities. Instead of requiring explicit intermediate targets, we identify recurring behavior patterns associated with successful outcomes across trajectories and use them to supervise promising behaviors before terminal feedback arrives. We introduce PROSE (Progress Recognition from Outcome Statistics across Experience), which abstracts long trajectories into visually coherent periods, aligns similar periods across episodes into recurring situations and transitions, estimates their progress from cross-trajectory outcome statistics, and refines it along individual trajectories by penalizing subsequently lost progress and accounting for how each transition is visually and behaviorally executed. When progress persistently stalls, a targeted exploration controller uses the agent's previous observations to redirect exploration away from unpromising regions. Across five long-horizon MineDojo tasks, integrating PROSE with two world-model agents improves success rate by 9.1 points and reduces steps to first success by 14.4% on average. These results show that PROSE complements different world-model agents by turning sparse outcomes into intermediate supervision for long-horizon embodied tasks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.