acceptodds
Under review as a conference paper at ICLR 2027

Beyond Interaction Capacity: Estimator Scaling with Recursive Models for CTR Prediction

Abstract

Click-Through Rate (CTR) prediction, a core task in recommendation and advertising systems, relies on modeling interactions among sparse categorical features. Explicit cross networks are a central paradigm for CTR prediction, and recent progress has largely come from increasing the interaction capacity of a single predictor through deeper cross networks and more expressive cross operators. We revisit whether continually increasing interaction capacity remains the most effective way to improve predictive performance, and find that its benefits quickly exhibit diminishing returns even as capacity continues to grow. This motivates a complementary scaling direction that we call estimator scaling, where additional resources are used to incorporate multiple related estimators rather than only enlarging a single predictor. Through theoretical analysis and diagnostic experiments, we show that the gains from estimator scaling are governed by the amount of non-shared predictive variation available across estimators. However, exploiting this variation naively can be expensive: independently trained models provide substantial estimator diversity but require deployment cost to grow with ensemble size. This motivates a parameter-efficient realization of estimator scaling that can incorporate diversity from multiple estimator sources without maintaining multiple full models. Building on this view, we introduce RECursive Averaged Predictor (RECAP), a parameter-efficient recursive CTR model that operationalizes estimator scaling at three levels: distillation across independently trained models, exponential moving averaging over training trajectories, and aggregation over inference-time routes within a weight-shared recursive backbone. This design improves predictive performance without requiring parameter growth proportional to the ensemble size. Experiments across multiple CTR benchmarks show that RECAP approaches the accuracy of a five-model ensemble of a strong cross-network baseline at one fifth of its deployed parameters, while doubling the baseline's interaction capacity recovers only a third of that gain. Overall, these results establish new state-of-the-art predictive performance on standard benchmarks, while placing the proposed approach on a favorable performance–parameter Pareto frontier.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.