acceptodds
Under review as a conference paper at ICLR 2027

TIMELOOP: Turning Training Trajectories into Realized Future Supervision

Abstract

Training trajectories are usually consumed as optimization traces, even though their delayed intervention consequences contain information for later decisions. We study whether those consequences can become reusable empirical supervision rather than predictions of an unobserved future. TIMELOOP stores matched-state branch outcomes as a Future Memory whose records bind a training state, action, future horizon, and realized utility. At a later query, it retrieves comparable realized futures to rank candidate data mixtures before new outcomes are revealed. Across 20 natural decision states, Full TIMELOOP recovered the realized best action in 20/20 cases, whereas removing future supervision or using an Always-Balanced policy recovered 10/20. A natural signal-attribution study further found 22/22 top-1 decisions and mean Spearman correlation 0.9818 for Full TIMELOOP, compared with 54.55% top-1 for state-only, future-only, and randomized controls. In a controlled shift, old memory without aligned new evidence achieved 3/20, while progressively adding shift-aligned experience increased mean top-1 accuracy from 0.15 to 1.00 across five acquisition orders. These results support an experience-conditioned decision mechanism: realized training futures can become reusable supervision, but TIMELOOP is not a zero-shot oracle for unseen utility changes. The evidence is bounded to one 336M model, a controlled four-domain ecosystem, and five hand-defined actions; reliability also depends on memory support, horizon matching, action coverage, anchor isolation, and the cost of acquiring future branches.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.