acceptodds
Under review as a conference paper at ICLR 2027

Adaptive Ensemble Aggregation for Actor-Critics

Abstract

Ensembles are standard in off-policy actor-critic learning, but their efficacy depends critically on how they are aggregated. Many methods construct two ensemble-based reference estimators: a critic reference for bootstrapping and an actor reference for policy optimization. Existing approaches typically control these references through prescribed aggregation rules or externally specified hyperparameters. We introduce Adaptive Ensemble Aggregation (AEA), which calibrates both references based on actor-critic training dynamics. Extensive empirical evaluations study AEA across three settings: interactive learning, sample-efficient learning, and learning with external data collected by other behavior policies. Across multiple continuous-control tasks and external datasets of varying behavioral quality and coverage, AEA performs competitively against strong baselines while using the same AEA update throughout.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.