acceptodds
Under review as a conference paper at ICLR 2027

From Bids to Modes: Policy Refinement with Latent Decision Modes for Auto-Bidding

Abstract

Auto-bidding plays a central role in modern computational advertising, yet learning improved policies solely from offline trajectories remains challenging. Existing methods typically perform policy improvement directly in the original continuous bid space, which is only sparsely covered by training data. Moreover, bid-level optimization obscures distinctions among different decision tendencies, making it difficult to switch among them dynamically as market and budget conditions evolve. To address this, we introduce **PRISM** (**P**olicy **R**efinement with latent dec**IS**ion **M**odes), a framework for auto-bidding that optimizes policies over a compact mode-level decision space. To construct this space, PRISM uses return-tilted mode discovery to learn a partition that captures outcome-relevant distinctions between modes. Building on the discovered modes, it selects the best-evaluated mode from contextually supported candidates and realizes it through a continuous-bid policy trained with mode-referenced weighting. Experiments on AuctionNet show that PRISM consistently outperforms competitive auto-bidding baselines, while D4RL evaluations highlight its potential as a general policy-refinement framework for offline data containing diverse decision behaviors.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.