VisionTS: Visual Masked Autoencoders Are Free-Lunch Zero-Shot Time Series Forecasters

Mouxiang ChenLefei ShenZhuo LiXiaoyun Joy WangJianling SunChenghao Liu

article2025ICML107 citations

Demonstrates that off-the-shelf visual masked autoencoders pre-trained on ImageNet can outperform dedicated time series foundation models in zero-shot forecasting by casting 1D sequences into 2D masked image reconstruction tasks without requiring domain-specific pre-training.

Listen

Organizations increasingly rely on time series forecasting for core operational decisions such as energy planning, supply chain management, and traffic routing. Historically, building these systems required training individual models for each dataset from scratch. Recent attempts to create universal foundation models have either repurposed text-based large language models or trained new architectures on massive collections of domain-specific time series data. However, text models exhibit fundamental cross-domain incompatibilities with numerical data, while native time series datasets suffer from extreme diversity and inconsistency, creating substantial barriers to reliable transfer learning.

The article demonstrates that standard computer vision foundation models pre-trained exclusively on natural images can serve as effective zero-shot time series forecasters without requiring initial training on numerical data. To evaluate this approach, the authors introduced VISIONTS, a framework that reformulates numerical forecasting into a visual image reconstruction problem. The system transforms one-dimensional time series data into two-dimensional matrices segmented by temporal periodicity, renders these as grayscale images, and frames the future forecasting window as masked visual patches for an image completion model to reconstruct.

Evaluation across extensive benchmarks—including 8 long-term forecasting datasets, 29 Monash archive benchmarks, and the 23-dataset GIFT-Eval benchmark—revealed strong performance advantages. In zero-shot settings without any prior time series exposure, VISIONTS achieved the top ranking on the GIFT-Eval leaderboard, outperforming dedicated time series foundation models. Across long-term forecasting benchmarks, it delivered an average Mean Squared Error reduction of 8% to 84% compared to few-shot text-based models and reduced errors by approximately 6% compared to native time series models. When fine-tuned for just a single training pass on specific downstream datasets, VISIONTS achieved state-of-the-art results in 46 out of 80 test conditions. Furthermore, because it resamples inputs to fixed image dimensions, it maintained near-constant computational inference speed as input context lengths increased, avoiding the computational slowdown typical of standard sequence models.

These findings indicate that continuous visual pixel variations share deep mathematical similarities with real-world numerical trends, seasonality, and physical dynamics. For technology and operations leadership, this approach offers a cost-effective pathway to high-performance forecasting by bypassing the expensive compute cycles and curation efforts required to train massive time series models from scratch. Organizations can rapidly deploy off-the-shelf vision architectures to achieve competitive zero-shot predictions or invest minimal compute in parameter-efficient fine-tuning on proprietary data.

Before executing large-scale production deployments, organizations should validate this method on their specific internal data streams. VISIONTS processes multivariate series using channel independence—forecasting each variable separately—which limits its ability to model complex inter-variable interactions. Additionally, qualitative analyses indicate that the model tends to make aggressive trend projections on unstructured, volatile data where conservative baselines might yield lower variance. However, given its strong empirical validation across diverse domains, exploring pre-trained visual architectures presents a reliable, high-performing alternative to conventional time series foundation models.

No sufficiently relevant recommendations were found.

Cover for VisionTS: Visual Masked Autoencoders Are Free-Lunch Zero-Shot Time Series Forecasters

Abstract

Foundation models have emerged as a promising approach in time series forecasting (TSF). Existing approaches either repurpose large language models (LLMs) or build large-scale time series datasets to develop TSF foundation models for universal forecasting. However, these methods face challenges due to the severe cross-domain gap or in-domain heterogeneity. This paper explores a new road to building a TSF foundation model from rich, high-quality natural images. Our key insight is that a visual masked autoencoder, pre-trained on the ImageNet dataset, can naturally be a numeric series forecaster. By reformulating TSF as an image reconstruction task, we bridge the gap between image pre-training and TSF downstream tasks. Surprisingly, without further adaptation in the time series domain, the proposed VisionTS could achieve better zero-shot forecast performance than existing TSF foundation models. With fine-tuning for one epoch, VisionTS could further improve the forecasting and achieve state-of-the-art performance in most cases. Extensive experiments reveal intrinsic similarities between images and real-world time series, suggesting that visual models may offer a “free lunch” for TSF and highlight the potential for future cross-modality research. Our code is publicly available at https://github.com/Keytoyze/VisionTS.

Table of Contents

  • 1. Introduction
  • 2. Preliminaries
  • 3. Methodology
  • 4. Experiments
  • 4.1. Zero-Shot Time Series Forecasting
  • 4.2. Further Analysis of VISIONTS
  • 4.3. Full-Shot Long-Term Time Series Forecasting
  • 5. Related Work
  • 6. Conclusion
  • Acknowledgments
  • Impact Statement
  • References
  • Appendix
  • A. Details of Experiments
  • A.1. Benchmark and baselines
  • A.2. Periodicity selection
  • B. Zero-Shot Forecasting
  • B.1. Hyperparameters
  • B.2. Full forecasting results of the long-term TSF benchmark
  • B.3. Comparison of TimesFM and LLMTime
  • B.4. Comparison of traditional methods
  • B.5. Comparison of concurrent works
  • B.6. Full forecasting results of the Monash TSF benchmark
  • B.7. Impact of backbones
  • B.8. Impact of the different image encoding strategies
  • B.9. Hyperparameter analysis
  • C. Full-Shot Forecasting
  • C.1. Training details
  • C.2. Full results and standard deviations
  • C.3. Ablation study and fine-tuning strategy comparison
  • D. Visualization

Knowls

  1. Knowl 1 — VISIONTS turns forecasting into masked image completion

    model/method

    VISIONTS uses a visual masked autoencoder (MAE) pre-trained on ImageNet as a time-series forecaster, without first training it on time-series data. For a univariate series, it converts the observed context into a two-dimensional image, presents the image region corresponding to the context as visible patches, and treats the region corresponding to the forecast horizon as masked patches. The MAE reconstructs the masked region; VISIONTS converts that reconstruction back into a numeric forecast. Thus the pre-trained model’s image-completion task supplies the forecasting operation. For multivariate inputs, VISIONTS follows channel independence: it forecasts each variable separately rather than modeling interactions between variables.

  2. Knowl 2 — Period-based segmentation supplies the time-series image structure

    model/method

    Given a univariate context of length LL and a selected period length PP, VISIONTS divides the context into ⌊L/P⌋\lfloor L/P\rfloor non-overlapping subsequences of length PP and stacks them as columns of a matrix with PP rows. This arrangement places within-period variation along one axis and variation across periods at corresponding phases along the other. The period is selected from candidates suggested by the data’s sampling frequency, with validation performance used to choose among them; when there is no clear periodicity, P=1P=1 is allowed. On ETTh1, ETTh2, ETTm1, and ETTm2, using the frequency-based period selection rather than always setting P=1P=1 reduced the average MSE from 0.5590.559 to 0.3440.344 and the average MAE from 0.4920.492 to 0.3700.370.

  3. Knowl 3 — Normalization and patch alignment adapt the series image to the MAE

    equation

    VISIONTS normalizes the segmented matrix IrawI_{\mathrm{raw}} using its instance mean and standard deviation, then scales it by a dimensionless factor rr:

    Inorm=r Iraw−Mean⁡(Iraw)StandardDeviation⁡(Iraw).I_{\mathrm{norm}}=r\,\frac{I_{\mathrm{raw}}-\operatorname{Mean}(I_{\mathrm{raw}})}{\operatorname{StandardDeviation}(I_{\mathrm{raw}})}.

    The normalized matrix is rendered as a grayscale image by copying its values into all three color channels. To match the MAE’s fixed-size input, let NN be the number of patches along each side of its square image, SS the pixel width of each patch, LL the context length, and HH the forecast horizon. VISIONTS resizes the image to dimensions (NS,nS)(NS,nS), where the number of visible patch columns is

    n=⌊cNLL+H⌋,n=\left\lfloor cN\frac{L}{L+H}\right\rfloor,

    and c∈[0,1]c\in[0,1] is an alignment factor. The NnNn patches on the left are visible; the N(N−n)N(N-n) patches on the right are masked for reconstruction. The paper uses bilinear interpolation for resizing and reports r=c=0.4r=c=0.4 as effective defaults across most zero-shot settings. After reconstruction, it resizes the image back, averages the three channels, reverses the normalization, and flattens the result to recover the forecast.

  4. Knowl 4 — VISIONTS is competitive with zero-shot and few-shot long-term forecasters

    empirical result

    On six long-term forecasting datasets—ETTh1, ETTh2, ETTm1, ETTm2, Electricity, and Weather—VISIONTS was evaluated at forecast horizons 9696, 192192, 336336, and 720720 without time-series training. Its aggregate MSE was 0.3090.309 and aggregate MAE was 0.3450.345. The corresponding values for MOIRAI Small, Base, and Large were respectively 0.327/0.3570.327/0.357, 0.310/0.3440.310/0.344, and 0.329/0.3500.329/0.350 (MSE/MAE). VISIONTS therefore improved on MOIRAI Small and Large in average MSE and was comparable to MOIRAI Base; it recorded the best result in most of the dataset-level comparisons reported by the paper. The traditional and text-based baselines in this comparison were fine-tuned using 10% of each downstream dataset, whereas VISIONTS used no downstream time-series training.

  5. Knowl 5 — Cross-domain benchmarks show strong zero-shot transfer

    empirical result

    On the 29-dataset Monash benchmark, VISIONTS achieved normalized MAE 0.7290.729 and ranked second overall; the leading MOIRAI Small result was 0.6570.657, while the cross-domain LLMTime result was 1.0411.041. The paper reports that VISIONTS outperformed the models trained individually on each dataset in the aggregate comparison. On GIFT-Eval, a benchmark covering 23 datasets, the paper reports that VISIONTS ranked first in normalized MASE on the leaderboard as of November 2024. These evaluations extend the zero-shot results beyond the long-term benchmark to heterogeneous datasets and forecasting settings.

  6. Knowl 6 — One-epoch layer-normalization tuning strengthens long-term forecasting

    empirical result

    For full-shot evaluation on eight long-term datasets, VISIONTS fine-tuned only the MAE’s layer-normalization parameters for one epoch per dataset, except on Illness, where it trained for 100 epochs with early stopping because of the limited training data. Across 80 reported dataset-and-metric comparisons, VISIONTS achieved 46 first-place results; PatchTST achieved 19, GPT4TS 12, and Time-LLM 4. Training treated each variable as an independent sample. In the reported comparisons, this small amount of downstream adaptation was sufficient for VISIONTS to outperform the other baselines in most cases.

  7. Knowl 7 — Ablations indicate that pretrained visual weights and layer normalization matter

    empirical result

    On the four ETT datasets in the full-shot ablation, the standard pretrained VISIONTS obtained average MSE 0.3330.333 and MAE 0.3690.369. Removing the visual model raised these to 0.5650.565 and 0.5200.520; replacing it with a randomly initialized MAE gave 0.4170.417 and 0.4140.414. The tested attention-only and Transformer replacements also performed worse than the pretrained model. Among fine-tuning strategies, updating only layer-normalization parameters gave the best reported average MSE and MAE (0.3330.333 and 0.3690.369); updating all parameters gave 0.4170.417 and 0.4140.414, and freezing all parameters gave 0.3600.360 and 0.3750.375. The results support the paper’s conclusion that pretrained visual representations contribute to forecasting and that layer-normalization tuning is preferable among the tested strategies.

  8. Knowl 8 — Encoder embeddings show overlap and heterogeneity across modalities

    empirical result

    To examine transfer between images and time series, the authors encoded 1,000 ImageNet images and 300 samples from each time-series dataset with the same MAE and mask, then visualized the embeddings in two dimensions using t-SNE. The ImageNet embeddings formed a relatively concentrated distribution, while time-series embeddings were more scattered and could occupy separated regions—for example, ETTm1 clustered apart from Monash. Some series, including ETTm1 and Electricity, fell within the ImageNet embedding distribution. The authors interpret this overlap as evidence consistent with a smaller representation gap for those series, and the separation among time-series datasets as evidence of cross-domain heterogeneity; the visualization is suggestive rather than a demonstration of causality.

  9. Knowl 9 — Forecasting quality changes little with larger MAE backbones

    empirical result

    In zero-shot evaluation across six long-term datasets, three ImageNet-pretrained MAE sizes—Base (112M parameters), Large (330M), and Huge (657M)—had average MSE values of 0.3090.309, 0.3110.311, and 0.3150.315, respectively. The larger backbones therefore did not improve aggregate forecasting accuracy in this experiment. The paper suggests that larger visual models may overfit image-specific features and reduce transferability, but presents this as a possible explanation rather than an established mechanism. A separate test using the visual inpainting model LaMa produced average MSE 0.3740.374 and MAE 0.3920.392 on the four ETT datasets, compared with 0.3440.344 and 0.3700.370 for MAE-based VISIONTS.

  10. Knowl 10 — VISIONTS does not model cross-variable dependence or forecast distributions

    limitation

    VISIONTS forecasts multivariate series using channel independence, so it does not capture interactions among variables. The paper also identifies distribution forecasting as unsupported by the visual model used. Qualitative zero-shot examples show a further failure mode: on less-structured input, VISIONTS can make aggressive trend predictions that sometimes increase error, while MOIRAI may make more conservative predictions. In one illustrated ETTh1 example, VISIONTS had MAE 0.3270.327 and MOIRAI Large had MAE 0.1720.172. The authors identify architectures capable of handling variable interactions and distributional forecasts as directions for future work.

Coverage note — Detailed runtime profiling and per-horizon result tables are omitted as secondary or redundant to the benchmark findings above.

References

  1. 1.Aksu, T., Woo, G., Liu, J., Liu, X., Liu, C., Savarese, S., Xiong, C., and Sahoo, D. Gift-eval: A benchmark for general time series forecasting model evaluation, 2024. URL https://arxiv.org/abs/2410.10393.
  2. 2.Alexandrov, A., Benidis, K., Bohlke-Schneider, M., Flunkert, V., Gasthaus, J., Januschowski, T., Maddix, D. C., Rangapuram, S., Salinas, D., Schulz, J., Stella, L., Türkmen, A. C., and Wang, Y. GluonTS: Probabilistic and Neural Time Series Modeling in Python. Journal of Machine Learning Research, 21(116):1–6, 2020. URL http://jmlr.org/papers/v21/19-820.html.
  3. 3.Ansari, A. F., Stella, L., Turkmen, C., Zhang, X., Mercado, P., Shen, H., Shchur, O., Rangapuram, S. S., Arango, S. P., Kapoor, S., et al. Chronos: Learning the language of time series. arXiv preprint arXiv:2403.07815, 2024.
  4. 4.Bao, H., Dong, L., Piao, S., and Wei, F. BEit: BERT pre-training of image transformers. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=p-BhZSz59o4.
  5. 5.Bian, Y., Ju, X., Li, J., Xu, Z., Cheng, D., and Xu, Q. Multi-patch prediction: Adapting llms for time series representation learning. arXiv preprint arXiv:2402.04852, 2024.
  6. 6.Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258, 2021.
  7. 7.Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. Language models are few-shot learners, 2020. URL https://arxiv.org/abs/2005.14165.
  8. 8.Chen, M., Shen, L., Fu, H., Li, Z., Sun, J., and Liu, C. Calibration of time-series forecasting: Detecting and adapting context-driven distribution shift. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, KDD ’24, pp. 341–352, New York, NY, USA, 2024. Association for Computing Machinery. ISBN 9798400704901. doi: 10.1145/3637528.3671926. URL https://doi.org/10.1145/3637528.3671926.
  9. 9.Das, A., Kong, W., Sen, R., and Zhou, Y. A decoder-only foundation model for time-series forecasting. In Forty-first International Conference on Machine Learning, 2024.
  10. 10.Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. Imagenet: A large-scale hierarchical image database. In 2009 IEEE Conference on Computer Vision and Pattern Recognition, pp. 248–255, 2009. doi: 10.1109/CVPR.2009.5206848.
  11. 11.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 4171–4186, Minneapolis, Minnesota, June 2019. Association for Computational Linguistics. doi: 10.18653/v1/N19-1423. URL https://aclanthology.org/N19-1423.
  12. 12.Dong, J., Wu, H., Wang, Y., Qiu, Y.-Z., Zhang, L., Wang, J., and Long, M. Timesiam: A pre-training framework for siamese time-series modeling. In Forty-first International Conference on Machine Learning, 2024.
  13. 13.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=YicbFdNTTy.
  14. 14.Ekambaram, V., Jati, A., Dayama, P., Mukherjee, S., Nguyen, N., Gifford, W. M., Reddy, C., and Kalagnanam, J. Tiny time mixers (ttms): Fast pre-trained models for enhanced zero/few-shot forecasting of multivariate time series. Advances in Neural Information Processing Systems, 37:74147–74181, 2024.
  15. 15.Feng, C., Huang, L., and Krompass, D. Only the curve shape matters: Training foundation models for zero-shot multivariate time series forecasting through next curve shape prediction. arXiv preprint arXiv:2402.07570, 2024.
  16. 16.Fu, F., Chen, J., Zhang, J., Yang, C., Ma, L., and Yang, Y. Are synthetic time-series data really not as good as real data?, 2024. URL https://arxiv.org/abs/2402.00607.
  17. 17.Godahewa, R. W., Bergmeir, C., Webb, G. I., Hyndman, R., and Montero-Manso, P. Monash time series forecasting archive. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2), 2021. URL https://openreview.net/forum?id=wEc1mgAjU-.
  18. 18.Goswami, M., Szafer, K., Choudhry, A., Cai, Y., Li, S., and Dubrawski, A. Moment: A family of open time-series foundation models. In Forty-first International Conference on Machine Learning, 2024.
  19. 19.Gruver, N., Finzi, M., Qiu, S., and Wilson, A. G. Large language models are zero-shot time series forecasters. Advances in Neural Information Processing Systems, 36, 2023.
  20. 20.Han, L., Ye, H.-J., and Zhan, D.-C. The capacity and robustness trade-off: Revisiting the channel independent strategy for multivariate time series forecasting. IEEE Transactions on Knowledge & Data Engineering, (01):1–14, 2024.
  21. 21.Hatami, N., Gavet, Y., and Debayle, J. Classification of time-series images using deep convolutional neural networks. In Tenth international conference on machine vision (ICMV 2017), volume 10696, pp. 242–249. SPIE, 2018.
  22. 22.He, K., Chen, X., Xie, S., Li, Y., Dollár, P., and Girshick, R. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 16000–16009, 2022.
  23. 23.Hsu, W.-N., Bolte, B., Tsai, Y.-H. H., Lakhotia, K., Salakhutdinov, R., and Mohamed, A. HuBERT: Self-supervised speech representation learning by masked prediction of hidden units. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 29:3451–3460, 2021.
  24. 24.Jin, M., Wang, S., Ma, L., Chu, Z., Zhang, J. Y., Shi, X., Chen, P.-Y., Liang, Y., Li, Y.-F., Pan, S., et al. Time-llm: Time series forecasting by reprogramming large language models. In The Twelfth International Conference on Learning Representations, 2024.
  25. 25.Kim, T., Kim, J., Tae, Y., Park, C., Choi, J.-H., and Choo, J. Reversible instance normalization for accurate time-series forecasting against distribution shift. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=cGDAkQo1C0p.
  26. 26.Krizhevsky, A., Sutskever, I., and Hinton, G. E. Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems, 25, 2012.
  27. 27.Li, X., Kang, Y., and Li, F. Forecasting with time series imaging. Expert Systems with Applications, 160:113680, 2020.
  28. 28.Li, Z., Li, S., and Yan, X. Time series as images: Vision transformer for irregularly sampled time series. Advances in Neural Information Processing Systems, 36, 2024.
  29. 29.Lin, S., Lin, W., Wu, W., Chen, H., and Yang, J. Sparsetsf: Modeling long-term time series forecasting with 1k parameters. In Forty-first International Conference on Machine Learning, 2024.
  30. 30.Liu, P., Guo, H., Dai, T., Li, N., Bao, J., Ren, X., Jiang, Y., and Xia, S.-T. Calf: Aligning llms for time series forecasting via cross-modal fine-tuning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 39, pp. 18915–18923, 2025.
  31. 31.Liu, Y., Wu, H., Wang, J., and Long, M. Non-stationary transformers: Exploring the stationarity in time series forecasting, 2022.
  32. 32.Liu, Y., Zhang, H., Li, C., Huang, X., Wang, J., and Long, M. Timer: Generative pre-trained transformers are large time series models. In Forty-first International Conference on Machine Learning, 2024.
  33. 33.Ma, Q., Liu, Z., Zheng, Z., Huang, Z., Zhu, S., Yu, Z., and Kwok, J. T. A survey on time-series pre-trained models. arXiv preprint arXiv:2305.10716, 2023.
  34. 34.Nie, Y., Nguyen, N. H., Sinthong, P., and Kalagnanam, J. A time series is worth 64 words: Long-term forecasting with transformers. In The Eleventh International Conference on Learning Representations, 2022.
  35. 35.Peebles, W. and Xie, S. Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 4195–4205, 2023.
  36. 36.Qiu, X., Hu, J., Zhou, L., Wu, X., Du, J., Zhang, B., Guo, C., Zhou, A., Jensen, C. S., Sheng, Z., and Yang, B. TFB: towards comprehensive and fair benchmarking of time series forecasting methods. Proc. VLDB Endow., 17(9):2363–2377, 2024.
  37. 37.Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019.
  38. 38.Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 10684–10695, 2022.
  39. 39.Schick, T. and Schütze, H. Exploiting cloze-questions for few-shot text classification and natural language inference. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pp. 255–269, 2021.
  40. 40.Semenoglou, A.-A., Spiliotis, E., and Assimakopoulos, V. Image-based time series forecasting: A deep convolutional neural network approach. Neural Networks, 157:39–53, 2023.
  41. 41.Shi, X., Wang, S., Nie, Y., Li, D., Ye, Z., Wen, Q., and Jin, M. Time-moe: Billion-scale time series foundation models with mixture of experts. arXiv preprint arXiv:2409.16040, 2024.
  42. 42.Sood, S., Zeng, Z., Cohen, N., Balch, T., and Veloso, M. Visual time series forecasting: an image-driven approach. In Proceedings of the Second ACM International Conference on AI in Finance, pp. 1–9, 2021.
  43. 43.Suvorov, R., Logacheva, E., Mashikhin, A., Remizova, A., Ashukha, A., Silvestrov, A., Kong, N., Goka, H., Park, K., and Lempitsky, V. Resolution-robust large mask inpainting with fourier convolutions. In Proceedings of the IEEE/CVF winter conference on applications of computer vision, pp. 2149–2159, 2022.
  44. 44.Tan, M., Merrill, M. A., Gupta, V., Althoff, T., and Hartvigsen, T. Are language models actually useful for time series forecasting? arXiv preprint arXiv:2406.16964, 2024.
  45. 45.Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023.
  46. 46.Wang, J., Zhao, S., Luo, Z., Zhou, Y., Jiang, H., Li, S., Li, T., and Pan, G. CBramod: A criss-cross brain foundation model for EEG decoding. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net/forum?id=NPNUHgHF2w.
  47. 47.Wang, Z. and Oates, T. Imaging time-series to improve classification and imputation. In Proceedings of the 24th International Conference on Artificial Intelligence, pp. 3939–3945, 2015a.
  48. 48.Wang, Z. and Oates, T. Spatially encoding temporal correlations to classify temporal data using convolutional neural networks. arXiv preprint arXiv:1509.07481, 2015b.
  49. 49.Wimmer, C. and Rekabsaz, N. Leveraging vision-language models for granular market change prediction. arXiv preprint arXiv:2301.10166, 2023.
  50. 50.Woo, G., Liu, C., Sahoo, D., Kumar, A., and Hoi, S. CoST: Contrastive learning of disentangled seasonal-trend representations for time series forecasting. In International Conference on Learning Representations, 2022a. URL https://openreview.net/forum?id=PilZY3omXV2.
  51. 51.Woo, G., Liu, C., Sahoo, D., Kumar, A., and Hoi, S. C. H. Etsformer: Exponential smoothing transformers for time-series forecasting. CoRR, abs/2202.01381, 2022b. URL https://arxiv.org/abs/2202.01381.
  52. 52.Woo, G., Liu, C., Kumar, A., Xiong, C., Savarese, S., and Sahoo, D. Unified training of universal time series forecasting transformers. In Forty-first International Conference on Machine Learning, 2024.
  53. 53.Wu, H., Xu, J., Wang, J., and Long, M. Autoformer: Decomposition transformers with auto-correlation for long-term series forecasting. Advances in Neural Information Processing Systems, 34:22419–22430, 2021.
  54. 54.Wu, H., Hu, T., Liu, Y., Zhou, H., Wang, J., and Long, M. Timesnet: Temporal 2d-variation modeling for general time series analysis. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=ju_Uqw384Oq.
  55. 55.Xue, H. and Salim, F. D. Promptcast: A new prompt-based learning paradigm for time series forecasting. IEEE Transactions on Knowledge and Data Engineering, 2023.
  56. 56.Yang, L., Wang, Y., Fan, X., Cohen, I., Zhao, Y., and Zhang, Z. Vitime: A visual intelligence-based foundation model for time series forecasting. arXiv preprint arXiv:2407.07311, 2024.
  57. 57.Yue, Z., Wang, Y., Duan, J., Yang, T., Huang, C., Tong, Y., and Xu, B. Ts2vec: Towards universal representation of time series. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, pp. 8980–8987, 2022.
  58. 58.Zaken, E. B., Goldberg, Y., and Ravfogel, S. Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pp. 1–9, 2022.
  59. 59.Zeng, A., Chen, M., Zhang, L., and Xu, Q. Are transformers effective for time series forecasting? In Proceedings of the AAAI conference on artificial intelligence, volume 37, pp. 11121–11128, 2023.
  60. 60.Zerveas, G., Jayaraman, S., Patel, D., Bhamidipaty, A., and Eickhoff, C. A transformer-based framework for multivariate time series representation learning. In Proceedings of the 27th ACM SIGKDD conference on knowledge discovery & data mining, pp. 2114–2124, 2021.
  61. 61.Zhang, K., Wen, Q., Zhang, C., Cai, R., Jin, M., Liu, Y., Zhang, J. Y., Liang, Y., Pang, G., Song, D., et al. Self-supervised learning for time series analysis: Taxonomy, progress, and prospects. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024.
  62. 62.Zhang, Y., Zhang, Y., Zheng, M., Chen, K., Gao, C., Ge, R., Teng, S., Jelloul, A., Rao, J., Guo, X., et al. Insight miner: A time series analysis dataset for cross-domain alignment with natural language. In NeurIPS 2023 AI for Science Workshop, 2023.
  63. 63.Zhou, H., Zhang, S., Peng, J., Zhang, S., Li, J., Xiong, H., and Zhang, W. Informer: Beyond efficient transformer for long sequence time-series forecasting. In Proceedings of the AAAI conference on artificial intelligence, volume 35, pp. 11106–11115, 2021.
  64. 64.Zhou, T., Ma, Z., Wen, Q., Wang, X., Sun, L., and Jin, R. Fedformer: Frequency enhanced decomposed transformer for long-term series forecasting. In International Conference on Machine Learning, pp. 27268–27286. PMLR, 2022.
  65. 65.Zhou, T., Niu, P., Sun, L., Jin, R., et al. One fits all: Power general time series analysis by pretrained lm. In Advances in neural information processing systems, volume 36, pp. 43322–43355, 2023.

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/