acceptodds
Under review as a conference paper at ICLR 2027

TS-Edit: A Model Editing Framework for Adversarial Defense in Time-Series Foundation Models

Abstract

Time-series foundation models (TSFMs) report strong zero-shot forecasting performance, but where they store factual associations internally that drive a particular forecast, and whether the association can be corrected after a malicious attack are open questions. While these models are increasingly deployed in industrial settings such as power grids, traffic and finance, they remain vulnerable to careful backdoor injections inducing catastrophic operational hazards. To address this, we introduce TS-Edit, a novel model editing framework for adversarial defense in TSFMs. We extend model interpretability approaches built for language models such as causal tracing, ROME, sequential ROME and MEMIT on decoder-only transformer models (TimesFM 2.5, Toto 2.0) and evaluate the framework on popular time-series benchmarks. We use masked white-box attacks (FGSM, BIM and PGD) to corrupt a single segment of the context window; causal tracing then localizes the damage across various MLP layers and a closed-form rank-one update to the identified feed-forward weight then repairs the corrupted forecast. For instance, a single edit after a PGD attack on the UCI Electricity data reduces the corrupted-forecast NRMSE by while preserving the model's behavior on clean inputs. The same localize-then-edit procedure, applied per attack, also repairs BIM and FGSM by and at negligible drift on clean-input predictions (specificity) on the edited model. Subsequent iterative editing exposes a restoration-versus-specificity trade-off. Sequential ROME gives the most consistent restoration across attacks, whereas MEMIT attains both the lowest error and the lowest specificity under the strongest attack, PGD. In every case the clean-input drift remains negligible. We further verify the robustness of the model using cross quarter attacks with random starts and varying attack types. Together, the results demonstrate that both the key-value memory interpretation of feed-forward layers and their localized editability can extend beyond language models. The experiments also demonstrate that attack salience within TSFMs can be localized, diagnosed, and repaired through surgical edits to individual weights.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.