How Interaction History Shapes Evidence Use In Language Models
Abstract
LLM agents rely on interaction history to support multi-step tasks, yet this history can change not only what information they access, but also how they use it in decisions. We investigate this process in MemSyco-Bench, where historical preferences conflict with task evidence. Our circuit-level analysis of LLM evidence use, combining input edits and activation interventions, yields three findings. First, preference and correct evidence retain opposing influences in later layer answer competition; preference’s larger state-change amplitude helps it prevail, rather than evidence being entirely suppressed. Second, repeated preference and middle-layer attention mass contribute to this amplitude advantage, linking historical reading to downstream decision effects. Third, early preference-content states weaken later scope-sensitive reading: reducing these states makes option queries more responsive to applicability, but does not strengthen scope’s absolute correction of the answer. Thus, content and scope interact during reading without ensuring appropriately constrained decisions. Matched tests show that early preference-state and scope-related state interventions shift answers in the same broad directions in Qwen3.5-4B and Qwen3.5-27B, without establishing an identical circuit. In a 300-instance behavioral screen, among cases made robustly correct by preference deletion, the fraction reversed by full history is lower at 27B than at 4B (18.3% vs. 24.1%), but remains nonzero. The tested 27B model is therefore not immune to this history-induced failure; the comparison does not isolate model scale as its cause.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.