Neuron Sharing in Multi-Task Reinforcement Learning Policies Follows the Geometry of the Task Parameterization
Abstract
Multi-task RL policies are usually treated as monolithic, yet prior work shows subnetworks for individual behaviors exist within them. We investigate how these subnetworks are structured within the network to identify neuron reusage and specificity. Using gradient-based importance profiles and causal ablation on policies trained with Proximal Policy Optimization, we measure neuron sharing between tasks. In a continuously parameterized locomotion task (AntDir), neuron sharing decays monotonically with angular distance between task goal directions, consistently across seeds and network widths. Removing one task's top neurons preferentially degrades its angular neighbors. This organization persists when all continuous goal features are replaced by one-hot task IDs, so it is learned from the behavioral geometry of the task family rather than inherited from the input encoding. In HalfCheetahVel, overlap also tracks velocity distance, but ablation damage is nonselective, showing that geometry-aligned overlap does not imply functional modularity. On a fully learned seven-task Meta-World subset, which lacks a continuous task parameterization, the ordering of pairwise overlap barely replicates across seeds and ablation effects are asymmetric. Together, these results indicate that neuron sharing in these policies follows the geometry of the task parameterization when one exists, and that overlap alone does not establish functional modularity.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.