Verified Cognitive Schema Transfer for Test-Time Learning
Abstract
Test-time scaling improves the problem-solving capabilities of large language models (LLMs) by allocating additional inference compute to each problem. Existing memory-augmented methods enable problem-solving experience to accumulate and transfer across tasks, but accumulated experience does not automatically yield a systematic understanding of solution principles. A key challenge is to use verification feedback to abstract structures shared across concrete problem-solving experiences and organize them into cognitive schemas with explicit applicability conditions that guide subsequent reasoning. To address this challenge, we introduce Verified Cognitive Schema Transfer (VCST), a framework for accumulating and reusing knowledge during inference without updating model parameters. VCST extracts task-level schemas from verified solutions, abstracts shared solution structures across tasks, and reuses the resulting schemas after checking their applicability conditions. For new tasks, it retrieves and applies relevant knowledge through tag-based filtering and semantic reranking. Under a fixed API budget, VCST adaptively balances cognition construction and task solving. Experiments across four benchmark suites, 13 benchmarks, and nine backbone LLMs show that VCST improves accuracy over strong test-time scaling baselines. On LiveCodeBench, VCST improves the final solve rate over S∗ by 6.1–20.7 percentage points across nine backbone models, and enables a 7B model to outperform a 32B model equipped with S∗. VCST also generalizes broadly, improving over the strongest corresponding baselines by 2.2–2.4 points on multi-task reasoning and 8.7 points on code execution prediction. On multi-language code generation, VCST achieves the highest accuracy in all six evaluated languages.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.