acceptodds
Under review as a conference paper at ICLR 2027

Q-Map Fusion: Recovering Shared Dynamics from Separately Learned Values

Abstract

Value functions in reinforcement learning are tied to particular rewards, yet tasks in the same environment share the dynamics that give rise to those values. We study how knowledge acquired separately across tasks can be combined after training to support new decisions. Building on Bellman inversion, Q-Map Fusion jointly fits a shared transition model to the constraints supplied by multiple Q-functions, then plans under new rewards and termination rules. Recovery uses the source values and known task definitions, requiring no training trajectories, no updates to the Q-functions, and no additional environment interaction. The usefulness of an additional task depends on the ambiguities its values resolve and the accuracy with which they are learned. For finite-state models, our analysis characterizes when the joint constraints control specified one-step predictions even when the full transition model remains unidentified. Experiments in FourRooms, MountainCar, Pendulum, and Reacher demonstrate effective model recovery and planning for held-out tasks across the evaluated tabular and continuous-control settings. Separate matched-budget and perturbation studies show that the benefit of additional values depends on both complementary task information and robustness to value errors. We hope this perspective encourages the reuse of learned values as shared environment knowledge for decisions beyond their original rewards.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.