Scaling the Inference Program: System-Level Test-Time Scaling for Recommendation
Abstract
Test-time scaling improves predictions without updating model weights, but it typically spends additional computation independently for each query. Recommendation exposes a distinct system-level opportunity: various requests repeatedly pass through the same frozen recommendation models, making the reusable, request-conditioned inference program a natural object on which to spend additional compute. We introduce System-Level Test-Time Scaling and instantiate it with the proposed framework ProgramRec. Starting from a blank program interface, a large language model iteratively constructs contract-valid inference programs that jointly specify feature construction , a recursive execution graph , and numerical parameters . Aggregate validation feedback and the history of evaluated programs guide subsequent proposals. The selected program is then frozen and reused deterministically across requests without further LLM calls or model updates. Across three datasets, the strongest ProgramRec variant achieves the best overall performance, while larger discovery budgets yield gains in every setting. These results establish reusable inference-program construction as a new scaling axis for recommendation. The code is available at https://anonymous.4open.science/r/ProgramRec-0675.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.