NoiseRub: Noise-Calibrated Dynamic Rubric Rewards for Real-Time Visual Games
Abstract
Training agents in real-time visual games from screen-only interaction requires dense feedback beyond sparse terminal outcomes. Vision-language judges can provide rubric-based rewards, but dynamically weighting these rewards is challenging because across-segment score variance can reflect behavioral differences, judge inconsistency, or limited opportunities for a rubric to apply. We introduce NoiseRub, a noise-calibrated approach to dynamic rubric rewards. Rather than treating observed score variance as evidence of an informative rubric, NoiseRub asks how much of that variation provides reliable evidence for adapting each rubric's weight and how well that evidence is supported by evaluation opportunities. It therefore treats dynamic weighting as a reliability-estimation problem: judge-induced variation should be discounted, while weight adaptation should reflect how often a rubric can be meaningfully evaluated. Each rubric's reliable variation is estimated by subtracting a within-segment noise floor, obtained from repeated audits, from its across-segment variance; the resulting nonnegative residual is then modulated by opportunity coverage. The resulting weights aggregate rubric scores into segment rewards for online policy optimization. In Black Myth: Wukong, NoiseRub achieves 63.5% agreement with human overall preferences, compared with 57.4% for fixed weights and 60.0% for POW3R-Game. Under matched initialization and interaction budgets, it achieves the highest HP reduction on all four evaluated bosses, outperforming POW3R-Game by 4.45–8.04 percentage points. Supplementary videos are available on the https://visual-game-supplement.pages.dev.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.