Test-Time Symmetrization: Scaling Laws for Model-Agnostic Learning under Structure
Abstract
Test-time compute has emerged as a powerful way to improve the accuracy of language models at inference time, with scaling laws quantifying how accuracy improves with additional forward passes. Existing scaling laws, however, focus on repeated sampling from a fixed prompt, even though many tasks admit multiple equivalent representations of the same underlying problem. A graph does not depend on how its nodes are numbered, and a program does not depend on the names of its variables, yet language models can change their answers under such rewritings. In this paper, we study test-time symmetrization, which uses forward passes on rewritings of the input and aggregates the resulting answers over the group of rewritings. We prove that no amount of sampling and repeated decoding from a single representation can replace rewriting, while only random rewritings suffice to recover the information provided by all rewritings. The same set of rewritings provides simultaneous approximation across models, decoding rules, and sampled reasoning traces, and the logarithmic scaling is optimal in general. Finally, exact computations confirm the rate and its logarithmic dependence on , and graph-reasoning experiments with a frozen open-weight model (Qwen2.5-1.5B) show that a single set of rewritings attains it across six tasks and six decoding rules. Our results provide a principled way to enforce symmetries at inference time when they cannot be built into the model, which we call model-agnostic symmetrization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.