CoT-Meta: Budgeted Metacognitive Control for Test-Time Reasoning
Abstract
Effective test-time reasoning requires deciding which partial trajectories to expand, which errors to repair, and when to conclude a search. We introduce CoT-Meta, a training-free framework that combines strategy-conditioned thought generation, tree-structured search, online process evaluation, and state-dependent control over expansion, pruning, repair, stopping, and abstention. A compact meta-state links intermediate reasoning quality to the allocation of a total-token budget. At , CoT-Meta achieves 92.8% on MATH-500, 90.4% on GPQA-Diamond, 75.8% on BBEH, and 56.7% Pass@1 on LiveCodeBench with Claude-4.5-Opus, and leads the evaluated baselines on all four tasks with Llama-3.1-8B-Instruct. Holding the inference components fixed and changing only action selection raises four-task macro accuracy by 2.0 points over a threshold-only policy on each backbone. A 900-state counterfactual audit associates this gain with higher best-action agreement and lower decision regret. Evaluator interventions, candidate-coverage audits, and token-matched repair comparisons connect the end-to-end gains to intermediate signal quality and targeted trajectory correction. Nested confidence-prefix evaluation further characterizes selective prediction separately from additional-budget fallback. These results establish an empirical case for explicit metacognitive control as a design principle for budgeted test-time reasoning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.