acceptodds
Under review as a conference paper at ICLR 2027

SPOT: Stateful Post-Training Search by LLM Researchers and Actionable Training Insights

Abstract

Post-training requires coupled decisions about data, model updates, and checkpoint initialization; because each experiment changes the evidence and checkpoints available next, these decisions form a sequential joint search. LLM researchers can close this loop, but existing systems either expose a narrow action space or permit pipeline changes that make experiments difficult to compare. We present SPOT (Stateful POst-Training Search), which converts hypotheses into structured joint interventions and executes them through a shared interface. SPOT records both the proposed recipe and the operation actually consumed by training, then returns its telemetry, evaluation outcomes, and resulting checkpoint to guide subsequent pro- posals. Across six search campaigns, SPOT evaluates 141 candidate checkpoints; every LLM researcher improves over the shared base, and the best search-time composite score increases from 0.4092 to 0.5152 on a 0–1 scale. Interface ab- lations further show that restricting proposal scope changes both update choice and checkpoint inheritance, rather than merely reducing the number of editable fields. Beyond endpoint scores, the aligned records reveal a diagnostic hierarchy for choosing the next experiment. Does the sampled data provide usable learning contrast, or do different-quality outputs receive indistinguishable rewards? Does the surviving signal reward the intended behavior, or does the proxy improve while the target behavior deteriorates? Can the current search space express the required repair, or does persistence across explored controls indicate a missing intervention? These questions distinguish when to continue training, change sampling, redesign the objective, or expand the intervention space.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.