Beyond First Success: Expanding Achievement Rewards in Open-World Survival Games
Abstract
Long-horizon tasks often require repeatedly performing intermediate behaviors that have already succeeded. Rewards for first success alone do not directly reflect subsequent progress through repetition, while continually rewarding repeated behavior can favor easily triggered events. We introduce Structured Ordinal Reward (SOR), which expands achievement feedback along two dimensions: event identity and recurrence level. Using first-success labels and structured reference trajectories, SOR learns event rules, identifies a finite set of observationally supported recurrence levels, and allocates a total auxiliary reward budget according to event recurrence support. The constructed reward program remains fixed during downstream training. Given reference data and approximately one million downstream training steps, SOR improves PPO's cumulative training achievement score in Crafter by 19.5% and its independent native evaluation score in Craftax-full by 37.2%. We further conduct payment-timing controls and budget-permutation experiments within SOR to study how feedback along recurrence and reward allocation across events affect learning. Experiments combining SOR with Achievement Distillation demonstrate compatibility with achievement representation learning, supporting its use as an auxiliary reward module for existing learners. Code is available at https://anonymous.4open.science/r/structured-ordinal-reward-DE30/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.