acceptodds
Under review as a conference paper at ICLR 2027

Watermarking Values, Not Tokens: Distribution-Free Verification for LLM Time-Series Forecasts

Abstract

Large language models can generate probabilistic time-series forecasts as numerical ensembles, allowing downstream users to account for predictive uncertainty rather than rely on a single point forecast. When such forecasts are released and reused, verifying their provenance becomes important. Existing watermarking methods, however, are poorly aligned with this setting in two respects. First, watermark evidence is typically defined over generated tokens, whereas forecast releases are manipulated as numerical values through operations such as rounding, reordering, subsampling, and cropping. Second, watermarking an autoregressive forecast can feed modified values back into later predictions, causing local edits to propagate through the forecast horizon. To address these limitations, we propose V-band, a generation-time value-space watermarking framework for frozen autoregressive language-model forecasts. V-band uses the owner’s secret key to label numerical bands at each forecast step and tilts the model’s conditional digit probabilities so that sampled values are more likely to fall within preferred bands. We further introduce value-feedback-free steering, which preserves unmarked history across forecast steps while conditioning lower-order digits within each value on the higher-order digits that will appear in the released value. The resulting watermark is verified directly from released numerical values by ranking the owner key against independent decoy keys. We establish finite-sample false-positive control under key exchangeability, exact invariance to row reordering, and theoretical connections between numerical displacement, accessible band diversity, and watermark retention. Experiments on 1,455 units across five evaluation sets show evidence retention under the nine registered release operations with small pooled CRPS changes. A separate study of 26 compatible checkpoints across six model families examines transfer under the same band-width rule and verifier.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.