Less Tuning, Better Planning: Simplifying Offline Reinforcement Learning via Planning
Abstract
Offline reinforcement learning (RL) aims to learn effective policies entirely from precollected data, making it attractive for domains where online interaction is expensive or unsafe. Yet, existing offline RL methods are highly sensitive to hyperparameters and design choices, often requiring online evaluation to identify a good configuration, undermining the offline setting. This work investigates whether plug-and-play model-based planning can improve both policy performance and robustness to offline RL hyperparameters. We focus on two stages of the planning pipeline: test-time planning and offline policy extraction. For test-time planning, we find that performance is highly sensitive to the planning horizon, reflecting a trade-off between incorporating long-term information and accumulating model error. To address this, we propose Soft Horizon AggRegation for Planning (SHARP), which evaluates candidate actions across multiple horizons and softly aggregates their advantages using uncertainty estimates from an ensemble of dynamics models. For policy extraction, we find that Q-function exploitation can improve zero-shot policy performance but provides little additional benefit when the policy serves as an action proposer for planning, whereas behavioral regularization remains essential. Consequently, simple behavior cloning is sufficient as an action proposer. Building on these findings, we propose SHARP-BC, a plug-and-play planning method that requires no planning-specific hyperparameter tuning. Across D4RL locomotion and Adroit dexterous manipulation tasks, SHARP-BC achieves competitive or superior performance to existing offline planning methods while substantially reducing sensitivity to both planning and underlying offline RL hyperparameters.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.