acceptodds
Under review as a conference paper at ICLR 2027

SRAO: Coordinating Auxiliary Guidance and Feedback for Efficient Learning

Abstract

Learning systems often combine inexpensive guidance with feedback that is costly to obtain. Their efficiency depends on coordinating how guidance shapes an update with how feedback corrects it. We introduce Shared Residual Audited Optimization (SRAO), which uses corrective feedback both to improve the current update and to inform the use of guidance. For selective gradient acquisition, one reference corrects the estimate and evaluates all available guidance-acquisition rules. The key insight is that the estimation gain from combining guidance and acquisition choices can pay for learning how to combine them. We establish this connection through a sharp variance bound and a cumulative learning guarantee. Scalar and gradient studies explain the mechanism, including about 40% lower estimation error at matched reference cost. On-policy distillation retains similar observed mean accuracy with 35.5% fewer reference-gradient batches. Agentic question answering extends the coordination principle to reflection guidance and outcome feedback, improving mean task quality. Together, these results connect feedback efficiency to the way guidance enters learning updates.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.