acceptodds
Under review as a conference paper at ICLR 2027

Deep Q-Learning Without a Target Network

Abstract

A target network is the standard device for stabilizing deep Q-learning, at the price of storing a full copy of the online network and tuning its refresh period. We ask what a target actually needs to store and answer by lowering the storage level from a full copy, through partial copies, to none. A first-order analysis shows that a target decouples only the parameter directions it stores, and in deep networks negative feedback from a frozen copy is what keeps partial targets alive. Removing storage entirely exposes spurious solutions at which the loss vanishes while the values are wrong. A bounded multiplicative output scaling, which scales the network's own prediction by a factor slightly above one before regressing it onto the bootstrap, repairs this with a value-error bound. We call the resulting method Tempered Q-learning. In a stress test where the full copy diverges, it survives while storing nothing, matches or exceeds the copy's return once the copy's period is matched in environment time, and on ALE Breakout reaches the full copy's mean return at a fraction of the spread. The replacement is two lines of code, with no second network and no synchronization period to tune.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.