When Global Grades Fail Locally: Learning Local Supervision from Aggregate Feedback and Hindsight
Abstract
Agentic systems often receive one grade for an entire trajectory, even though learning requires feedback on individual steps. Copying the final grade to every step can therefore create wrong local labels: a failed trajectory may contain correct steps, and a successful trajectory may contain mistakes. We study how to recover local correctness without densely labeling every step. Our method combines aggregate feedback with a small number of local labels to train a local critic. Hindsight then allows the critic to reassess earlier claims using evidence that becomes available later in the trajectory. Our theory explains what aggregate feedback leaves unresolved and when hindsight provides additional information. It also shows how local correctness relates to errors in inherited labels. We evaluate the method in a controlled multimodal forecast-and-recount setting with independent simulator ground truth. On a frozen final set, the retrospective critic reduces local-correctness log loss by 32% relative to a matched static critic, from 0.5781 to 0.3943. A separate frozen validation set shows the same pattern. These results show that aggregate feedback, sparse local supervision, and hindsight can recover useful local supervision without dense step-level labels.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.