Can Test-Time Memory Evolution Outperform Test-Time Training for Scientific Discovery?
Abstract
Recent test-time training (TTT) frameworks have enabled small open-weight Large Language Models (LLMs) to outperform inference-only systems powered by frontier models on scientific discovery tasks. However, TTT introduces additional computational cost for online weight updates during the discovery process, making computational efficiency an important consideration when scaling to larger models. This motivates us to explore an efficient test-time adaptation method for scientific discovery. Under the same discovery framework, we introduce Test-Time Memory Evolution (TTME), a plug-and-play test-time adaptation method, as an alternative to TTT. TTME continually updates the external memory by distilling the search experience to guide subsequent discovery without updating model weights. By controlling the discovery framework, we have a clean comparison between TTME and TTT on both performance and cost. Our evaluation shows that TTME outperforms TTT in all five mathematical tasks, while reducing computational cost by in FLOPs and in runtime. These results establish TTME as an effective and computationally efficient alternative to TTT in the evaluated discovery settings.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.