When Does Navigation Pretraining Help? The Role of Action Selection
Abstract
Does the way a reinforcement-learning agent chooses actions change the measured benefit of pretraining? We study recurrent PPO agents adapting from local maze grids to directional wall distances and goal-visibility signals. Transfer retains the trained controller behind a fresh input encoder. Two settings compare maze pretraining, starting from scratch and equally long obstacle-field pretraining, with eight training seeds per condition. We evaluate the same saved networks by always taking the most probable action (greedy selection) or drawing from the learned probabilities (sampling). In the original setting, sampling improves scratch more than maze pretraining. The pretrained-minus-scratch difference in average success during adaptation on a zero-to-one scale changes from +0.053 to −0.054: a shift of −0.107 (descriptive 95% bootstrap interval [−0.182, −0.044]). However, neither ordering establishes superiority, and the task inventories overlap. An exploratory follow-up on larger mazes with structurally separate training and test sets does not reproduce this reversal: pretraining stays ahead on average, and the change in its advantage is uncertain. A further control preserves the probability of leaving the greedy action but chooses the alternatives equally. Sampling improves performance beyond this control, without clearly changing the pretraining advantage. Action selection can change a transfer comparison, but that sensitivity does not necessarily carry over to another task setting.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.