From Hidden-State Decoding to Selective Temporal Repair
Abstract
When can temporal information recovered from a language model's hidden states improve its answers? We study Qwen2.5-3B, Llama-3.2-3B, and Gemma-2-2B through four interfaces: local decoding, frozen transfer, graph composition, and selective replacement. On 1,082 held-out OmniTemp event pairs, frozen readouts exceed a lexical control fitted on the same labels by 12.07–16.58 Macro-F1 points, with positive document-bootstrap intervals. A separate coupled mechanism combines per-model adapters, a shared event-timeline decoder, source-fitted gates, and cross-model agreement. On 706 TDG edges, it improves all three native outputs by 3.85–11.14 points, with positive gain intervals and more helpful than harmful edits. The broader evaluation identifies the scope of these gains. Local readouts fail the predeclared strict-order criterion; only Gemma retains a positive lexical-control margin under MATRES transfer; and readout-driven graph solving reaches 31.7–33.3% accuracy against a 100% gold-edge oracle. TDG's gains coexist with low overlap recall, while later, separately fitted variants fail joint transfer criteria on MEANTIME and AQUAINT. These results establish benefits from a shared selective mechanism on the TDG population and motivate evaluating coverage, class recall, and harmful edits alongside decoding accuracy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.