acceptodds
Under review as a conference paper at ICLR 2027

READ-Bench: Benchmarking Historical Instance Retrieval for Time-Series Diagnosis

Abstract

Time-series diagnostic systems rarely rely on retrieving relevant historical cases, and when they do, retrieval is typically evaluated only indirectly through downstream prediction. We introduce **READ-Bench**, a benchmark for historical-case retrieval across 12 diagnostic datasets, centered on multivariate time series, that defines relevance by shared fault or event type rather than signal shape, so that visually different traces of the same fault count as relevant while similar-looking traces of different faults do not. Organizing existing datasets into queries, corpora, and explicit relevance judgments, READ-Bench is, to our knowledge, the broadest testbed to date for historical-case retrieval across diagnostic domains. Analogous to retrieval-augmented generation, we treat retrieval as a base retriever followed by a reranker, evaluating classical distances, symbolic retrievers, self-supervised and foundation-model embedders, and their fusion for search, and label-aware and language-model rerankers for reranking, under one protocol that varies supervision, normal-series pollution, and corpus scale with significance testing across datasets. Under a common channel-independent retrieval interface, pretrained representations offer no statistically detectable advantage over strong classical and symbolic baselines for search alone. The decisive factor is instead a small amount of resolved-case supervision at reranking, namely a Gaussian-process reranker that propagates a few neighbor labels in the embedding space and improves rankings far more than swapping among more sophisticated unsupervised representations or language-model reasoning, a gain that holds under corpus pollution and at full corpus scale. Guided by these findings, we introduce normal-residual scoring, which ranks each window by its departure from normal operation, and build a system that fuses a normal-residual-scored foundation-model embedder with a dynamic time warping leg by reciprocal-rank fusion, then reranks with the label-aware Gaussian-process reranker. This system improves NDCG@10 over its own search stage on all 12 datasets, by from the reranking step alone and by over the strongest single base retriever applied uniformly across datasets.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.