Backward Truncation Changes the Target Geometry of Local Credit Assignment
Abstract
Adaptive states can keep updating in the forward pass while their derivatives are truncated. We show that the local supervision matched to the full task gradient depends on these retained derivatives: the adjustment can be directional rather than a global loss rescaling. We derive a direction-wise matching rule and test its consequences with controlled interventions. Changing only retained state derivatives reverses the average preference between fixed supervision rules. Directional controls separate the correction from scalar weakening; Quotient reduces mean test error by 9.4% relative to Point. With known dynamics, we also construct supervision before training, obtaining mean gains across two tracker variants and two optimizers. In a coherent optical receiver using recordings from a 1125-km link, gradient-matching preferences vary with horizon and trainable space, while training effects depend on configuration and retained derivatives. With dispersion, nonlinear, and residual compensation trainable under truncation, fixed Complex Quotient supervision improves receiver quality factor (Q-factor) over Point at all 20 tested power–channel workpoints (five powers, four channels), with gains in all 40 paired training-capture comparisons and a mean of 0.045 dB. These findings establish backward differentiation policy as an explicit design variable for local supervision in adaptive systems.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.