DAGA: Distributional Actor-Critic with Action-Gradient Alignment and Stability-Triggered Guidance
Abstract
Multi-critic actor–critic methods improve continuous-control learning by aggregating critic values or return distributions, yet deterministic actors are updated through critic action gradients. Critics can therefore agree in value while inducing different policy-update directions, a phenomenon we call value–gradient mismatch. We propose DAGA, a Gaussian multi-critic framework combining Stability-Triggered Guidance (STG), which adapts actor guidance using distribution-level reliability, with Action-Gradient Policy Alignment (AGPA), which selectively regularizes actor-relevant gradient directions when excessive dispersion coexists with a coherent consensus. We characterize value–gradient decoupling and both mechanisms locally. Experiments evaluate mismatch prevalence, component contributions, cross-task performance, and sensitivity to intervention strength and ensemble size.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.