Confidence-Aware Pareto-Evolving Auto-bidding via Uncertainty-Regularized Counterfactual Regret
Abstract
Constrained auto-bidding requires maximizing marketing value while strictly satisfying efficiency constraints such as Target Cost-Per-Action (CPA). Recent Decision-Transformer-based methods introduce a dual-stream constraint context and a counterfactual regret loss to actively push the policy toward the Pareto frontier. However, we identify two fundamental gaps that cap their performance: (i) the Pareto frontier is constructed once from historical data and remains static, anchoring the learning target at the empirical rather than the theoretical optimum; and (ii) the outcome predictor used for counterfactual evaluation is a deterministic point estimator, so out-of-distribution counterfactuals can hallucinate high utility and mislead policy updates. We propose CoPE-Bid, a Confidence-aware Pareto-Evolving auto-bidding framework built on two synergistic mechanisms: (i) Uncertainty-Regularized Counterfactual Regret (UR-CRO) replaces the point-estimate predictor with a lightweight ensemble and evaluates counterfactuals via a lower-confidence-bound utility, so predictor uncertainty automatically discounts hallucinated regret; and (ii) Frontier-Anchored Self-Improvement (FASI) periodically performs confidence-gated imagination rollouts from the current policy and admits only high-confidence, Pareto-improving trajectories into the frontier, letting the sampling anchor co-evolve with the policy. The two modules are coupled through a shared uncertainty signal—UR-CRO tells FASI what to trust, and FASI tells UR-CRO what to aim for. Extensive experiments on AuctionNet, its sparse variant, and online A/B tests demonstrate that CoPE-Bid consistently outperforms state-of-the-art baselines in both value acquisition and constraint satisfaction, with particularly large gains in the noise-heavy sparse regime.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.