acceptodds
Under review as a conference paper at ICLR 2027

When Global Grades Fail Locally: Learning Local Supervision from Aggregate Feedback and Hindsight

Abstract

Agentic systems often receive one grade for an entire trajectory, even though learning requires feedback on individual steps. Copying the final grade to every step can therefore create wrong local labels: a failed trajectory may contain correct steps, and a successful trajectory may contain mistakes. We study how to recover local correctness without densely labeling every step. Our method combines aggregate feedback with a small number of local labels to train a local critic. Hindsight then allows the critic to reassess earlier claims using evidence that becomes available later in the trajectory. Our theory explains what aggregate feedback leaves unresolved and when hindsight provides additional information. It also shows how local correctness relates to errors in inherited labels. We evaluate the method in a controlled multimodal forecast-and-recount setting with independent simulator ground truth. On a frozen final set, the retrospective critic reduces local-correctness log loss by 32% relative to a matched static critic, from 0.5781 to 0.3943. A separate frozen validation set shows the same pattern. These results show that aggregate feedback, sparse local supervision, and hindsight can recover useful local supervision without dense step-level labels.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.