Adaptive Ensemble Aggregation for Actor-Critics
Abstract
Ensembles are standard in off-policy actor-critic learning, but their efficacy depends critically on how they are aggregated. Many methods construct two ensemble-based reference estimators: a critic reference for bootstrapping and an actor reference for policy optimization. Existing approaches typically control these references through prescribed aggregation rules or externally specified hyperparameters. We introduce Adaptive Ensemble Aggregation (AEA), which calibrates both references based on actor-critic training dynamics. Extensive empirical evaluations study AEA across three settings: interactive learning, sample-efficient learning, and learning with external data collected by other behavior policies. Across multiple continuous-control tasks and external datasets of varying behavioral quality and coverage, AEA performs competitively against strong baselines while using the same AEA update throughout.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.