TimeVista: Exploring and Exploiting Vision-Language Models as Judges for Time Series Forecasting
Abstract
High-quality time series forecasting is pivotal for real-world decision-making. However, traditional numerical metrics often fail to reveal complex temporal patterns and are limited in their ability to align with human intuitive preferences. While the “LLM-as-a-Judge” paradigm has revolutionized text evaluation by providing flexible, human-aligned judgment, its application to time series remains largely unexplored. In this paper, we leverage Vision-Language Models (VLMs) as judges for time series forecasting, harnessing their ability to comprehend time series plots grounded in textual constraints. Specifically, we propose a novel framework integrating Micro- and Macro-Level judgments informed by contextual information to evaluate time series forecasting. To this end, we introduce TimeVista, comprising 5,563 time series samples with evaluation rubrics, and Meta-TimeVista, comprising 1,025 forecast instances annotated by five experts. Meta-evaluation shows that VLM judges possess better alignment with human judgments than point-wise and structure-aware numerical baselines. Further analyses examine the effects of visual rendering and rubric design on VLM judgments. Building upon our benchmark, we comprehensively assess recent Time Series Foundation Models (TSFMs) under the VLM-as-a-Judge paradigm. Our results demonstrate that VLMs serve as robust and interpretable judges, providing a comprehensive, human-aligned standard for evaluating time series models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.