VisionTS: Visual Masked Autoencoders Are Free-Lunch Zero-Shot Time Series Forecasters
Mouxiang ChenLefei ShenZhuo LiXiaoyun Joy WangJianling SunChenghao Liu
Demonstrates that off-the-shelf visual masked autoencoders pre-trained on ImageNet can outperform dedicated time series foundation models in zero-shot forecasting by casting 1D sequences into 2D masked image reconstruction tasks without requiring domain-specific pre-training.
Organizations increasingly rely on time series forecasting for core operational decisions such as energy planning, supply chain management, and traffic routing. Historically, building these systems required training individual models for each dataset from scratch. Recent attempts to create universal foundation models have either repurposed text-based large language models or trained new architectures on massive collections of domain-specific time series data. However, text models exhibit fundamental cross-domain incompatibilities with numerical data, while native time series datasets suffer from extreme diversity and inconsistency, creating substantial barriers to reliable transfer learning.
The article demonstrates that standard computer vision foundation models pre-trained exclusively on natural images can serve as effective zero-shot time series forecasters without requiring initial training on numerical data. To evaluate this approach, the authors introduced VISIONTS, a framework that reformulates numerical forecasting into a visual image reconstruction problem. The system transforms one-dimensional time series data into two-dimensional matrices segmented by temporal periodicity, renders these as grayscale images, and frames the future forecasting window as masked visual patches for an image completion model to reconstruct.
Evaluation across extensive benchmarks—including 8 long-term forecasting datasets, 29 Monash archive benchmarks, and the 23-dataset GIFT-Eval benchmark—revealed strong performance advantages. In zero-shot settings without any prior time series exposure, VISIONTS achieved the top ranking on the GIFT-Eval leaderboard, outperforming dedicated time series foundation models. Across long-term forecasting benchmarks, it delivered an average Mean Squared Error reduction of 8% to 84% compared to few-shot text-based models and reduced errors by approximately 6% compared to native time series models. When fine-tuned for just a single training pass on specific downstream datasets, VISIONTS achieved state-of-the-art results in 46 out of 80 test conditions. Furthermore, because it resamples inputs to fixed image dimensions, it maintained near-constant computational inference speed as input context lengths increased, avoiding the computational slowdown typical of standard sequence models.
These findings indicate that continuous visual pixel variations share deep mathematical similarities with real-world numerical trends, seasonality, and physical dynamics. For technology and operations leadership, this approach offers a cost-effective pathway to high-performance forecasting by bypassing the expensive compute cycles and curation efforts required to train massive time series models from scratch. Organizations can rapidly deploy off-the-shelf vision architectures to achieve competitive zero-shot predictions or invest minimal compute in parameter-efficient fine-tuning on proprietary data.
Before executing large-scale production deployments, organizations should validate this method on their specific internal data streams. VISIONTS processes multivariate series using channel independence—forecasting each variable separately—which limits its ability to model complex inter-variable interactions. Additionally, qualitative analyses indicate that the model tends to make aggressive trend projections on unstructured, volatile data where conservative baselines might yield lower variance. However, given its strong empirical validation across diverse domains, exploring pre-trained visual architectures presents a reliable, high-performing alternative to conventional time series foundation models.
- Paper: Image as a Foreign Language: BEIT Pretraining for Vision and Vision-Language Tasks, Wenhui Wang et al. (2023). BEIT-3’s masked-image pretraining provides the visual mask-and-reconstruct foundation that VisionTS repurposes for forecasting numerical sequences.
No sufficiently relevant recommendations were found.
