Computationally Efficient Gradient Proxies for Large Reasoning Models
Abstract
Gradient-based optimization supports applications throughout the development and deployment of large reasoning models (LRMs), but repeatedly obtaining gradients through explicit reasoning traces can be expensive. Inspired by recent advances in efficient reasoning, we introduce computationally efficient gradient proxies using shorter reasoning traces to obtain useful optimization guidance for the original model. We instantiate three proxies by omitting explicit reasoning, reasoning in continuous latent space, or combining these two approaches. We assess the quality and cost of these gradient proxies and evaluate their performance in four representative downstream tasks across stages of model development and deployment, including data preparation, training, task adaptation, and inference. Extensive experiments demonstrate improved performance with lower computational cost in various tasks, supporting the generality of gradient proxies across these distinct optimization problems.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.