acceptodds
Under review as a conference paper at ICLR 2027

Continuous Expert Representation Geometry for Multi-Task Reinforcement Learning

Abstract

Multi-task reinforcement learning (MTRL) requires representations that are shared enough to support transfer yet specialized enough to limit interference. However, expert-based policies typically fix the degree of representational separation through architectural or optimization choices, leaving a basic question unresolved: how much separation should a multi-task learner impose? We study expert specialization as a continuous representation-geometry problem and introduce Continuous Orthogonality Control (COC), which exposes a controllable sharing–specialization spectrum while holding the expert networks, task conditioning, aggregation mechanism, optimization objective, and model capacity fixed. COC interpolates between an overlap-permitting expert frame and an orthogonal boundary, while realized geometry is measured independently through expert coherence. Across the evaluated settings, the favorable geometry is not universal. On Meta-World, both MT10 and MT50 achieve their highest observed performance within a broad intermediate-separation region across five seeds, whereas the maximally separated boundary achieves the highest observed return among the evaluated MiniGrid settings. Increasing the control coordinate systematically reduces expert coherence, confirming that COC produces the intended geometric intervention, while auxiliary response-structure and expert-usage diagnostics reveal setting-dependent internal responses. On Meta-World, however, performance varies non-monotonically with realized coherence, indicating that coherence is a geometry diagnostic rather than a universal monotonic criterion for selecting the favorable regime. Finally, on MT10, neither the tested open-loop schedules nor end-to-end optimization of learnable separation strengths matches the best fixed regime; the learned active coefficients instead move toward the maximally separated boundary. These results separate two problems that are often conflated: controlling representation geometry and identifying which geometry is favorable for a given multi-task learning setting.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.