Domain Adaptation for Time Series Forecasting via Attention Sharing

Xiaoyong JinYoungsuk ParkDanielle C. MaddixHao WangYuyang Wang

article2022ICML132 citations

Proposes an end-to-end domain adaptation framework for time series forecasting that aligns shared attention queries and keys across data-rich and data-scarce domains while preserving domain-specific values to improve multi-horizon predictions.

Listen

Modern predictive decision systems—such as retail inventory management, cloud resource provisioning, and vehicle control—increasingly rely on deep neural networks to produce accurate time series forecasts. However, these models require vast amounts of data to capture complex temporal dynamics effectively. In many real-world operational environments, organizations face severe data scarcity, such as limited historical observations (cold-start scenarios) or few available monitoring series (few-shot scenarios). When organizations attempt to transfer models trained on rich external data sources to these scarce target environments, they encounter domain shift—a distributional discrepancy that severely degrades forecasting accuracy.

The article introduces and evaluates the Domain Adaptation Forecaster (DAF), an end-to-end multi-horizon forecasting framework designed to solve data scarcity by transferring statistical strengths from a data-rich source domain to a data-scarce target domain.

To address this challenge without forcing incompatible features across disparate datasets, the authors designed a hybrid neural architecture. The framework maintains private encoders and decoders for each domain to handle domain-specific measurements and scales, while implementing a shared, attention-based module governed by a domain discriminator. Using adversarial training, the model forces structural temporal query and key patterns into a shared, domain-invariant latent space while keeping concrete values domain-specific. The model simultaneously performs input reconstruction and future-step extrapolation to stabilize training. The authors validated DAF across simulated cold-start and few-shot conditions as well as real-world benchmark datasets covering electricity consumption, traffic occupancy, store sales, and web page visits, comparing it against both single-domain models and cross-domain adaptation baselines using Normalized Deviation as the primary error metric.

The evaluation yielded several key findings:

  1. DAF consistently outperformed or matched all single-domain and cross-domain baselines across synthetic and real-world benchmarks, demonstrating the strongest accuracy gains in the most severely data-constrained settings.
  2. Cross-domain models trained jointly and end-to-end achieved markedly superior performance compared to two-stage fine-tuning approaches (such as DATSING) or traditional single-domain forecasters.
  3. In real-world multi-horizon transfer tasks, DAF lowered forecasting error across multiple cross-domain pairs (for example, achieving an error metric of 0.125 on electricity-to-traffic adaptations compared to 0.141 for DeepAR), outperforming recurrent-neural-network-based adaptation variants.
  4. Ablation analyses confirmed that keeping value embeddings domain-specific while enforcing query and key domain-invariance via adversarial training is the most critical design factor driving performance.
  5. DAF exhibited flexibility across distinct sampling frequencies, successfully transferring knowledge between hourly datasets (traffic and electricity) and daily datasets (sales and web traffic).

These findings demonstrate that organizations do not need to rely solely on expensive or slow target-domain data collection to deploy high-performing deep forecasting models. By aligning temporal dynamics across domains rather than raw output values, organizations can significantly mitigate the operational and financial risks of cold starts and small-sample deployments. This enables faster deployment cycles, reduced infrastructure costs for fine-tuning separate large-scale models, and improved decision reliability in newly launched operations.

For engineering and operational teams facing time series data scarcity, adopting an attention-sharing domain adaptation architecture represents a viable alternative to single-domain models or generic pre-training workflows. Organizations evaluating this approach should prioritize end-to-end joint training pipelines over staged fine-tuning. Before deploying in mission-critical environments, practitioners should run pilot validations across domain pairs to select optimal kernel and network hyperparameters.

While the empirical results are robust across tested benchmarks, the authors note that formal theoretical justifications for domain-invariant attention spaces remain an open area of inquiry. Furthermore, the empirical validation in the article is focused on univariate forecasting setups. Decision-makers should exercise appropriate caution when attempting to generalize these findings directly to complex, highly correlated multivariate forecasting systems without prior testing.

arXiv: 2102.06828
Cover for Domain Adaptation for Time Series Forecasting via Attention Sharing

Abstract

Recently, deep neural networks have gained increasing popularity in the field of time series forecasting. A primary reason for their success is their ability to effectively capture complex temporal dynamics across multiple related time series. The advantages of these deep forecasters only start to emerge in the presence of a sufficient amount of data. This poses a challenge for typical forecasting problems in practice, where there is a limited number of time series or observations per time series, or both. To cope with this data scarcity issue, we propose a novel domain adaptation framework, Domain Adaptation Forecaster (DAF). DAF leverages statistical strengths from a relevant domain with abundant data samples (source) to improve the performance on the domain of interest with limited data (target). In particular, we use an attention-based shared module with a domain discriminator across domains and private modules for individual domains. We induce domain-invariant latent features (queries and keys) and retrain domain-specific features (values) simultaneously to enable joint training of forecasters on source and target domains. A main insight is that our design of aligning keys allows the target domain to leverage source time series even with different characteristics. Extensive experiments on various domains demonstrate that our proposed method outperforms state-of-the-art baselines on synthetic and real-world datasets, and ablation studies verify the effectiveness of our design choices.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Domain Adaptation in Forecasting
  • 4. The Domain Adaptation Forecaster (DAF)
  • 4.1. Sequence Generators
  • 4.2. Domain Discriminator
  • 4.3. Adversarial Training
  • 5. Experiments
  • 5.1. Baselines and Evaluation
  • 5.2. Synthetic Datasets
  • 5.4. Additional Experiments
  • 5.5. Ablation Studies
  • 6. Conclusions
  • References
  • A. Dataset Details
  • A.1. Synthetic Datasets
  • A.2. Real-World Datasets
  • B. Implementation Details
  • B.1. Baselines
  • B.2. Hyperparameters
  • C. Detailed Experiment Results

Knowls

  1. Knowl 1 — DAF shares attention matching while keeping forecasts domain-specific

    model/method

    The Domain Adaptation Forecaster (DAF) transfers information from a data-rich source time-series domain to a data-scarce target domain using two sequence generators. Each generator has its own encoder and decoder, while the attention module is shared. The shared module learns queries and keys intended to be domain-invariant; the domain-specific encoders produce values, and the domain-specific decoders turn attention outputs into reconstructions and forecasts. Thus the domains share which historical patterns are matched, but can retain different signal scales and patterns in the values and outputs. Both generators are trained on their own domain’s data, and the target generator is used for target-domain forecasting.

  2. Knowl 2 — DAF constructs multi-scale patterns and shared query-key representations

    model/method

    For an input history X=[zt]t=1TX=[z_t]_{t=1}^{T} from either domain, DAF’s private encoder creates value embeddings vtv_t with a position-wise multilayer perceptron (MLP), and pattern embeddings ptp_t using MM temporal convolutions with different kernel sizes. The convolutional outputs are concatenated to form the multi-scale pattern representation; the pattern and value dimensions are kept equal. A shared position-wise MLP maps each pattern embedding to a query qtq_t and key ktk_t in a common dd-dimensional space. For a query at position tt, DAF normalizes a positive semi-definite kernel score over the eligible neighborhood N(t)\mathcal{N}(t):

    α(qt,kt′)=κ(qt,kt′)∑u∈N(t)κ(qt,ku),t′∈N(t).\alpha(q_t,k_{t'})=\frac{\kappa(q_t,k_{t'})}{\sum_{u\in\mathcal{N}(t)}\kappa(q_t,k_u)},\qquad t'\in\mathcal{N}(t).

    Here κ\kappa is the attention kernel, and t′t' and uu index key positions. The attention output is an MLP applied to the weighted sum of values associated with those keys. The query-key mapping and attention are shared across source and target, whereas value embeddings and the output decoders remain domain-specific.

  3. Knowl 3 — DAF uses attention both to reconstruct history and to extrapolate forecasts

    model/method

    DAF trains its sequence generators to reconstruct each observed input without allowing a position to attend to its own value, and to forecast future values autoregressively. For reconstruction at an observed position t≤Tt\leq T, the eligible keys are N(t)={1,…,T}∖{t}\mathcal{N}(t)=\{1,\ldots,T\}\setminus\{t\}, and key position t′t' is paired with value position μ(t′)=t′\mu(t')=t'. The reconstruction therefore combines other observed values, with weights determined by the query-key similarities.

    For the first forecast step T+1T+1, let ss be the convolution window size and sˉ=⌈(s−1)/2⌉\bar{s}=\lceil(s-1)/2\rceil. DAF reuses the query qT−sˉq_{T-\bar{s}} associated with the final complete historical window, attends to keys at positions N(T+1)={s,…,T−sˉ−1}\mathcal{N}(T+1)=\{s,\ldots,T-\bar{s}-1\}, and pairs each key at t′t' with the value that follows its local window: μ(t′)=t′+sˉ+1\mu(t')=t'+\bar{s}+1. The decoder maps the resulting attention output to a one-step prediction. That prediction is fed back through the encoder and attention module to generate later forecast steps recursively.

  4. Knowl 4 — DAF jointly fits both domains while adversarially aligning queries and keys

    equation

    Let DSD_S and DTD_T be source and target datasets, and let GSG_S and GTG_T be their respective sequence generators. DAF minimizes reconstruction and forecasting error on both domains while adversarially training a binary domain discriminator DD on the latent queries and keys:

    min⁡GS,GTmax⁡D  Lseq(DS;GS)+Lseq(DT;GT)−λLdom(DS,DT;D,GS,GT).\min_{G_S,G_T}\max_D\; L_{\mathrm{seq}}(D_S;G_S)+L_{\mathrm{seq}}(D_T;G_T)-\lambda L_{\mathrm{dom}}(D_S,D_T;D,G_S,G_T).

    The coefficient λ≥0\lambda\geq 0 weights the domain-classification term. For a dataset containing NN histories of length TT and forecast horizon τ\tau, the sequence loss sums each series’ mean input-reconstruction loss and mean future-prediction loss:

    Lseq(D;G)=∑i=1N[1T∑t=1Tℓ(zi,t,z^i,t)+1τ∑t=T+1T+τℓ(zi,t,z^i,t)].L_{\mathrm{seq}}(D;G)=\sum_{i=1}^{N}\left[\frac{1}{T}\sum_{t=1}^{T}\ell(z_{i,t},\hat z_{i,t})+\frac{1}{\tau}\sum_{t=T+1}^{T+\tau}\ell(z_{i,t},\hat z_{i,t})\right].

    Here zi,tz_{i,t} is the observed scalar at time tt for series ii, z^i,t\hat z_{i,t} is its reconstruction or forecast, and ℓ\ell is the pointwise loss; the experiments use mean squared error. The discriminator’s cross-entropy loss LdomL_{\mathrm{dom}} classifies source versus target queries and keys. Maximizing the objective with respect to DD trains it to classify domains, while minimizing with respect to the generators encourages their latent queries and keys to confuse the discriminator. The paper fixes λ=1\lambda=1 in its experiments.

  5. Knowl 5 — Adversarial training alternates generator and discriminator updates

    algorithm

    DAF trains the source and target generators jointly with a shared domain discriminator. Each generator produces an input reconstruction and an autoregressive forecast; the sequence and domain losses are then used for opposing updates. In the experiments, the sequence loss uses mean squared error and λ=1\lambda=1. Training proceeds until the target training data is exhausted in each epoch, for a maximum of 100,000 iterations, with early stopping based on validation normalized deviation.

    Input: Source dataset DSD_S, target dataset DTD_T, maximum epochs EE
    Initialize: Source and target generator parameters ΘG\Theta_G and discriminator parameters θD\theta_D
    for epoch = 1 to EE do
        repeat
            Sample source and target history-future batches from DSD_S and DTD_T
            Produce input reconstructions and autoregressive forecasts with GSG_S and GTG_T
            Compute the two-domain sequence loss and query-key domain-classification loss
            Form L=Lseq(DS;GS)+Lseq(DT;GT)−LdomL=L_{\mathrm{seq}}(D_S;G_S)+L_{\mathrm{seq}}(D_T;G_T)-L_{\mathrm{dom}}
            Update ΘG\Theta_G by gradient descent on LL
            Update θD\theta_D by gradient ascent on LL
        until the target training data is exhausted
    end for
    Output: Trained source and target generators and shared discriminator
  6. Knowl 6 — Synthetic experiments model cold-start and few-shot target forecasting

    experimental setup

    The synthetic source and target datasets contain sinusoidal time series with randomly sampled amplitudes, offsets, frequencies, and phases, plus Gaussian white noise with variance 0.20.2. The experiments use 5,000 source series and a forecast horizon of 18 steps. In the cold-start setting, source histories have length 144, target histories have length T∈{36,45,54}T\in\{36,45,54\}, and the target sinusoid period is fixed at 36; both domains have 5,000 series. In the few-shot setting, source histories and target histories both have length 144, while the target contains N∈{20,50,100}N\in\{20,50,100\} series and the source contains 5,000. Performance is evaluated by normalized deviation (ND), the sum of absolute forecast errors divided by the sum of absolute ground-truth values over all series and forecast steps; lower ND is better.

  7. Knowl 7 — DAF improves or matches baselines in synthetic cold-start and few-shot tests

    data/table

    The table reports mean ND and standard deviation over five runs for selected single-domain and cross-domain forecasters in the synthetic cold-start and few-shot settings. The cold-start rows vary target history length; the few-shot rows vary target-series count. DAF has the lowest reported mean ND in four of the six settings, and is close to the best result in the other two: at 20 target series it scores 0.057 versus 0.059 for RDA-ADDA, while at 50 target series RDA-ADDA scores 0.054 versus DAF’s 0.055. These results show that source data can help in both limited-history and limited-series regimes, though DAF is not the numerical winner in every individual setting.

    Task Target setting DAR VT AttF DATSING RDA-ADDA DAF
    Cold-start T=36T=36 0.053±0.0030.053\pm0.003 0.040±0.0010.040\pm0.001 0.042±0.0010.042\pm0.001 0.039±0.0040.039\pm0.004 0.035±0.0020.035\pm0.002 0.035±0.0030.035\pm0.003
    Cold-start T=45T=45 0.037±0.0020.037\pm0.002 0.039±0.0010.039\pm0.001 0.041±0.0040.041\pm0.004 0.039±0.0020.039\pm0.002 0.034±0.0010.034\pm0.001 0.030±0.0030.030\pm0.003
    Cold-start T=54T=54 0.031±0.0020.031\pm0.002 0.039±0.0010.039\pm0.001 0.038±0.0050.038\pm0.005 0.037±0.0010.037\pm0.001 0.034±0.0010.034\pm0.001 0.029±0.0030.029\pm0.003
    Few-shot N=20N=20 0.062±0.0030.062\pm0.003 0.089±0.0010.089\pm0.001 0.095±0.0030.095\pm0.003 0.078±0.0050.078\pm0.005 0.059±0.0030.059\pm0.003 0.057±0.0040.057\pm0.004
    Few-shot N=50N=50 0.059±0.0040.059\pm0.004 0.085±0.0010.085\pm0.001 0.074±0.0050.074\pm0.005 0.076±0.0060.076\pm0.006 0.054±0.0030.054\pm0.003 0.055±0.0010.055\pm0.001
    Few-shot N=100N=100 0.059±0.0030.059\pm0.003 0.079±0.0020.079\pm0.002 0.071±0.0020.071\pm0.002 0.058±0.0050.058\pm0.005 0.053±0.0070.053\pm0.007 0.051±0.0010.051\pm0.001
  8. Knowl 8 — Cross-dataset tests show lower target ND for DAF across eight transfers

    empirical result

    The real-world evaluation transfers between the hourly electricity-consumption dataset elec, hourly highway-occupancy dataset traf, daily store-sales dataset sales, and daily Wikipedia-visits dataset wiki. The full source dataset is used for adaptation, while the target is restricted to its final 30 days for hourly series or 60 days for daily series; those target portions are split equally into training, validation, and test sets. Forecast histories and horizons are 168 and 24 steps for hourly data and 28 and 7 steps for daily data. The table gives mean ND ±\pm standard deviation over five runs. It compares DAF with the target-only attention forecaster AttF and the best non-DAF cross-domain result for each transfer. DAF has lower mean ND than AttF in all eight transfers and is numerically best in seven; on wiki from sales, DAF ties the strongest comparator, RDA-ADDA.

    Target Source Horizon AttF Best non-DAF cross-domain result DAF
    traf elec 24 0.182±0.0070.182\pm0.007 RDA-ADDA: 0.174±0.0050.174\pm0.005 0.169±0.0020.169\pm0.002
    traf wiki 24 0.182±0.0070.182\pm0.007 RDA-MMD: 0.179±0.0040.179\pm0.004 0.176±0.0040.176\pm0.004
    elec traf 24 0.137±0.0050.137\pm0.005 RDA-DANN: 0.133±0.0050.133\pm0.005 0.125±0.0080.125\pm0.008
    elec sales 24 0.137±0.0050.137\pm0.005 RDA-DANN: 0.135±0.0070.135\pm0.007 0.123±0.0050.123\pm0.005
    wiki traf 7 0.050±0.0030.050\pm0.003 RDA-ADDA or RDA-MMD: 0.045±0.0030.045\pm0.003 0.042±0.0040.042\pm0.004
    wiki sales 7 0.050±0.0030.050\pm0.003 RDA-ADDA: 0.049±0.0030.049\pm0.003 0.049±0.0030.049\pm0.003
    sales elec 7 0.308±0.0020.308\pm0.002 RDA-ADDA: 0.281±0.0010.281\pm0.001 0.277±0.0050.277\pm0.005
    sales wiki 7 0.308±0.0020.308\pm0.002 RDA-ADDA: 0.287±0.0020.287\pm0.002 0.280±0.0070.280\pm0.007
  9. Knowl 9 — Ablations support adversarial query-key alignment and domain-specific values

    data/table

    Four real-world transfer tasks compare DAF with variants that remove adversarial training, stop sharing queries, stop sharing keys, or share values across domains. Each entry is target-domain ND, for which lower is better. Full DAF has the lowest value on all four tasks. The value-sharing variant has the largest degradation relative to DAF in each task, supporting the design choice to keep values domain-specific. Removing adversarial training or either query/key sharing component also worsens the reported result compared with full DAF.

    Target Source No adversarial training No query sharing No key sharing Shared values DAF
    traf elec 0.172 0.171 0.172 0.176 0.168
    elec traf 0.121 0.122 0.120 0.127 0.119
    wiki sales 0.042 0.042 0.044 0.049 0.041
    sales wiki 0.294 0.283 0.282 0.291 0.280
  10. Knowl 10 — The paper leaves theoretical alignment guarantees and multivariate evaluation open

    limitation

    The paper does not provide a theoretical justification for why domain-invariant queries and keys within attention should improve forecasting; it identifies this justification as an open problem. Its experiments evaluate univariate time-series forecasting, and multivariate forecasting is identified as future work. Consequently, the reported results do not establish theoretical guarantees or performance for multivariate settings.

Coverage note — The paper’s additional baseline comparisons and forecast visualizations were omitted because the main quantitative comparisons and the load-bearing attention-sharing mechanism are already captured.

References

  1. 1.Alam, F., Joty, S., and Imran, M. Domain Adaptation with Adversarial Training and Graph Embeddings. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 1077–1087, Melbourne, Australia, 2018. Association for Computational Linguistics. doi: 10.18653/v1/P18-1099. URL http://aclweb.org/anthology/P18-1099.
  2. 2.Bartunov, S. and Vetrov, D. P. Few-shot Generative Modelling with Generative Matching Networks. International Conference on Artificial Intelligence and Statistics, pp. 670–678, 2018.
  3. 3.Ben-David, S., Blitzer, J., Crammer, K., Kulesza, A., Pereira, F., and Vaughan, J. W. A theory of learning from different domains. Machine Learning, 79(1-2):151–175, May 2010. ISSN 0885-6125, 1573-0565. doi: 10.1007/s10994-009-5152-4. URL http://link.springer.com/10.1007/s10994-009-5152-4.
  4. 4.Borovykh, A., Bohte, S., and Oosterlee, C. W. Conditional time series forecasting with convolutional neural networks. arXiv preprint arXiv:1703.04691, 2017.
  5. 5.Bose, J.-H., Flunkert, V., Gasthaus, J., Januschkowski, T., Lange, D., Salinas, D., Schelter, S., Seeger, M., and Wang, Y. Probabilistic demand forecasting at scale. Proceedings of the VLDB Endowment, 10(12):1694–1705, 2017.
  6. 6.Bousmalis, K., Trigeorgis, G., Silberman, N., Krishnan, D., and Erhan, D. Domain Separation Networks. volume 29, pp. 345–351, 2016.
  7. 7.Cai, R., Chen, J., Li, Z., Chen, W., Zhang, K., Ye, J., Li, Z., Yang, X., and Zhang, Z. Time Series Domain Adaptation via Sparse Associative Structure Alignment. Proceedings of the AAAI Conference on Artificial Intelligence, 35:6859–6867, 2021.
  8. 8.Chen, W.-Y., Liu, Y.-C., Kira, Z., Wang, Y.-C. F., and Huang, J.-B. A Closer Look at Few-shot Classification. arXiv:1904.04232 [cs], January 2020. URL http://arxiv.org/abs/1904.04232. arXiv: 1904.04232.
  9. 9.Cortes, C. and Mohri, M. Domain Adaptation in Regression. In Kivinen, J., Szepesvari, C., Ukkonen, E., and Zeugmann, T. (eds.), Algorithmic Learning Theory, volume 6925, pp. 308–323. Springer Berlin Heidelberg, Berlin, Heidelberg, 2011. ISBN 978-3-642-24411-7 978-3-642-24412-4. doi: 10.1007/978-3-642-24412-4 25. URL http://link.springer.com/10.1007/978-3-642-24412-4_25. Series Title: Lecture Notes in Computer Science.
  10. 10.Cuturi, M. and Blondel, M. Soft-DTW: a Differentiable Loss Function for Time-Series. arXiv:1703.01541 [stat], March 2017. URL http://arxiv.org/abs/1703.01541. arXiv: 1703.01541.
  11. 11.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 4171–4186, 2019. URL http://arxiv.org/abs/1810.04805. arXiv: 1810.04805.
  12. 12.Dua, D. and Graff, C. UCI Machine Learning Repository. University of California, Irvine, School of Information and Computer Sciences, 2017. URL http://archive.ics.uci.edu/ml.
  13. 13.Flunkert, V., Salinas, D., and Gasthaus, J. DeepAR: Probabilistic Forecasting with Autoregressive Recurrent Networks. International Journal of Forecasting, 36:1181–1191, 2020. arXiv: 1704.04110.
  14. 14.Ganin, Y. and Lempitsky, V. Unsupervised Domain Adaptation by Backpropagation. International conference on machine learning, pp. 1180–1189, 2015.
  15. 15.Ganin, Y., Ustinova, E., Ajakan, H., Germain, P., Larochelle, H., Laviolette, F., Marchand, M., and Lempitsky, V. Domain-Adversarial Training of Neural Networks. The journal of machine learning research, 17:2096–2030, 2016.
  16. 16.Ghifary, M., Kleijn, W. B., Zhang, M., Balduzzi, D., and Li, W. Deep Reconstruction-Classification Networks for Unsupervised Domain Adaptation. In European conference on computer vision, pp. 597–613, 2016.
  17. 17.Guo, H., Pasunuru, R., and Bansal, M. Multi-Source Domain Adaptation for Text Classification via DistanceNet-Bandits. Proceedings of the AAAI Conference on Artificial Intelligence, 34(05):7830–7838, April 2020. ISSN 2374-3468, 2159-5399. doi: 10.1609/aaai.v34i05.6288. URL https://aaai.org/ojs/index.php/AAAI/article/view/6288.
  18. 18.Gururangan, S., Marasovic, A., Swayamdipta, S., Lo, K., Beltagy, I., Downey, D., and Smith, N. A. Don’t Stop Pretraining: Adapt Language Models to Domains and Tasks. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 8342–8360, Online, 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.acl-main.740. URL https://www.aclweb.org/anthology/2020.acl-main.740.
  19. 19.Han, X. and Eisenstein, J. Unsupervised Domain Adaptation of Contextualized Embeddings for Sequence Labeling. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pp. 4237–4247, Hong Kong, China, 2019. Association for Computational Linguistics. doi: 10.18653/v1/D19-1433. URL https://www.aclweb.org/anthology/D19-1433.
  20. 20.Hoffman, J., Tzeng, E., Park, T., Zhu, J.-Y., Isola, P., Saenko, K., Efros, A. A., and Darrell, T. CyCADA: Cycle-Consistent Adversarial Domain Adaptation. International conference on machine learning, pp. 1989–1998, 2018.
  21. 21.Hu, H., Tang, M., and Bai, C. DATSING: Data Augmented Time Series Forecasting with Adversarial Domain Adaptation. In Proceedings of the 29th ACM International Conference on Information & Knowledge Management, pp. 2061–2064, Virtual Event Ireland, October 2020. ACM. ISBN 978-1-4503-6859-9. doi: 10.1145/3340531.3412155. URL https://dl.acm.org/doi/10.1145/3340531.3412155.
  22. 22.Kan, K., Aubet, F.-X., Januschowski, T., Park, Y., Benidis, K., Ruthotto, L., and Gasthaus, J. Multivariate quantile function forecaster. In International Conference on Artificial Intelligence and Statistics, pp. 10603–10621. PMLR, 2022.
  23. 23.Kar, P. Dataset of Kaggle Competition Rossmann Store Sales, version 2, January 2019. URL https://www.kaggle.com/pratyushakar/rossmann-store-sales.
  24. 24.Kim, J., Park, Y., Fox, J. D., Boyd, S. P., and Dally, W. Optimal operation of a plug-in hybrid vehicle with battery thermal and degradation model. In 2020 American Control Conference (ACC), pp. 3083–3090. IEEE, 2020.
  25. 25.Lai. Dataset of Kaggle Competition Web Traffic Time Series Forecasting, Version 3, August 2017. URL https://www.kaggle.com/ymlai87416/wiktraffictimeseriesforecast/metadata.
  26. 26.Li, C.-L., Chang, W.-C., Cheng, Y., Yang, Y., and Póczos, B. MMD GAN: towards deeper understanding of moment matching network. In Proceedings of the 31st International Conference on Neural Information Processing Systems, pp. 2200–2210, 2017.
  27. 27.Li, S., Jin, X., Xuan, Y., Zhou, X., Chen, W., Wang, Y.-X., and Yan, X. Enhancing the Locality and Breaking the Memory Bottleneck of Transformer on Time Series Forecasting. Advances in Neural Information Processing Systems, pp. 5243–5253, 2019.
  28. 28.Liberty, E., Karnin, Z., Xiang, B., Rouesnel, L., Coskun, B., Nallapati, R., Delgado, J., Sadoughi, A., Astashonok, Y., Das, P., Balioglu, C., Chakravarty, S., Jha, M., Gautier, P., Arpin, D., Januschowski, T., Flunkert, V., Wang, Y., Gasthaus, J., Stella, L., Rangapuram, S., Salinas, D., Schelter, S., and Smola, A. Elastic Machine Learning Algorithms in Amazon SageMaker. In Proceedings of the 2020 ACM SIGMOD International Conference on Management of Data, pp. 731–737, Portland OR USA, June 2020. ACM. ISBN 978-1-4503-6735-6. doi: 10.1145/3318464.3386126. URL https://dl.acm.org/doi/10.1145/3318464.3386126.
  29. 29.Lim, B., Arik, S. O., Loeff, N., and Pfister, T. Temporal Fusion Transformers for Interpretable Multi-horizon Time Series Forecasting. arXiv:1912.09363 [cs, stat], December 2019. URL http://arxiv.org/abs/1912.09363. arXiv: 1912.09363.
  30. 30.Long, M., Cao, Y., Wang, J., and Jordan, M. I. Learning Transferable Features with Deep Adaptation Networks. pp. 97–105, 2015. URL http://arxiv.org/abs/1502.02791. arXiv: 1502.02791.
  31. 31.Motiian, S., Jones, Q., Iranmanesh, S. M., and Doretto, G. Few-Shot Adversarial Domain Adaptation. Advances in Neural Information Processing Systems, pp. 6670–6680, 2017. URL http://arxiv.org/abs/1711.02536. arXiv: 1711.02536.
  32. 32.Oreshkin, B. N., Carpov, D., Chapados, N., and Bengio, Y. Meta-learning framework with applications to zero-shot time-series forecasting. arXiv:2002.02887 [cs, stat], February 2020a. URL http://arxiv.org/abs/2002.02887. arXiv: 2002.02887.
  33. 33.Oreshkin, B. N., Chapados, N., Carpov, D., and Bengio, Y. N-BEATS: Neural basis expansion analysis for interpretable time series forecasting. International Conference on Learning Representations, pp. 31, 2020b.
  34. 34.Park, Y., Mahadik, K., Rossi, R. A., Wu, G., and Zhao, H. Linear quadratic regulator for resource-efficient cloud services. In Proceedings of the ACM Symposium on Cloud Computing, pp. 488–489, 2019.
  35. 35.Park, Y., Maddix, D., Aubet, F.-X., Kan, K., Gasthaus, J., and Wang, Y. Learning quantile functions without quantile crossing for distribution-free time series forecasting. In International Conference on Artificial Intelligence and Statistics, pp. 8127–8150. PMLR, 2022.
  36. 36.Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S. PyTorch: An Imperative Style, High-Performance Deep Learning Library. Advances in neural information processing systems, pp. 8026–8037, 2019.
  37. 37.Purushotham, S., Carvalho, W., Nilanon, T., and Liu, Y. Variational Recurrent Adversarial Deep Domain Adaptation. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings. OpenReview.net, 2017. URL https://openreview.net/forum?id=rk9eAFcxg.
  38. 38.Ramponi, A. and Plank, B. Neural Unsupervised Domain Adaptation in NLP—A Survey. arXiv:2006.00632 [cs], May 2020. URL http://arxiv.org/abs/2006.00632. arXiv: 2006.00632.
  39. 39.Rangapuram, S. S., Seeger, M., Gasthaus, J., Stella, L., Wang, Y., and Januschowski, T. Deep State Space Models for Time Series Forecasting. In Advances in neural information processing systems, pp. 7785–7794, 2018.
  40. 40.Rietzler, A., Stabinger, S., Opitz, P., and Engl, S. Adapt or Get Left Behind: Domain Adaptation through BERT Language Model Finetuning for Aspect-Target Sentiment Classification. Proceedings of The 12th Language Resources and Evaluation Conference, pp. 4933–4941, 2020.
  41. 41.Sen, R., Yu, H.-F., and Dhillon, I. Think Globally, Act Locally: A Deep Neural Network Approach to High-Dimensional Time Series Forecasting. Advances in Neural Information Processing Systems, 32, 2019.
  42. 42.Shi, G., Feng, C., Huang, L., Zhang, B., Ji, H., Liao, L., and Huang, H. Genre Separation Network with Adversarial Training for Cross-genre Relation Extraction. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pp. 1018–1023, Brussels, Belgium, 2018. Association for Computational Linguistics. doi: 10.18653/v1/D18-1125. URL http://aclweb.org/anthology/D18-1125.
  43. 43.Tzeng, E., Hoffman, J., Saenko, K., and Darrell, T. Adversarial Discriminative Domain Adaptation. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 2962–2971, Honolulu, HI, July 2017. IEEE. ISBN 978-1-5386-0457-1. doi: 10.1109/CVPR.2017.316. URL http://ieeexplore.ieee.org/document/8099799/.
  44. 44.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I. Attention Is All You Need. In Advances in neural information processing systems, pp. 5998–6008, 2017.
  45. 45.Wang, H., He, H., and Katabi, D. Continuously indexed domain adaptation. In ICML, 2020.
  46. 46.Wang, R., Maddix, D., Faloutsos, C., Wang, Y., and Yu, R. Bridging Physics-based and Data-driven modeling for Learning Dynamical Systems. In Learning for Dynamics and Control, pp. 385–398, 2021.
  47. 47.Wang, W., Van Gelder, P. H., and Vrijling, J. Some issues about the generalization of neural networks for time series prediction. In International Conference on Artificial Neural Networks, pp. 559–564. Springer, 2005.
  48. 48.Wang, Y., Smola, A., Maddix, D. C., Gasthaus, J., Foster, D., and Januschowski, T. Deep Factors for Forecasting. In International conference on machine learning, pp. 6607–6617, 2019.
  49. 49.Wen, R., Torkkola, K., Narayanaswamy, B., and Madeka, D. A Multi-Horizon Quantile Recurrent Forecaster. arXiv:1711.11053 [stat], November 2017. URL http://arxiv.org/abs/1711.11053. arXiv: 1711.11053.
  50. 50.Wilson, G. and Cook, D. J. A Survey of Unsupervised Deep Domain Adaptation. arXiv:1812.02849 [cs, stat], February 2020. URL http://arxiv.org/abs/1812.02849. arXiv: 1812.02849.
  51. 51.Wilson, G., Doppa, J. R., and Cook, D. J. Multi-Source Deep Domain Adaptation with Weak Supervision for Time-Series Sensor Data. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pp. 1768–1778, Virtual Event CA USA, August 2020. ACM. ISBN 978-1-4503-7998-4. doi: 10.1145/3394486.3403228. URL https://dl.acm.org/doi/10.1145/3394486.3403228.
  52. 52.Wright, D. and Augenstein, I. Transformer Based Multi-Source Domain Adaptation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 7963–7974, Online, 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.emnlp-main.639. URL https://www.aclweb.org/anthology/2020.emnlp-main.639.
  53. 53.Wu, S., Xiao, X., Ding, Q., Zhao, P., Wei, Y., and Huang, J. Adversarial Sparse Transformer for Time Series Forecasting. Advances in Neural Information Processing Systems, 33:11, 2020.
  54. 54.Xu, J., Wang, J., Long, M., and others. Autoformer: Decomposition transformers with auto-correlation for long-term series forecasting. Advances in Neural Information Processing Systems, 34, 2021.
  55. 55.Xu, Z., Lee, G.-H., Wang, Y., Wang, H., et al. Graph-relational domain adaptation. ICLR, 2022.
  56. 56.Yao, Z., Wang, Y., Long, M., and Wang, J. Unsupervised transfer learning for spatiotemporal predictive networks. In International Conference on Machine Learning, pp. 10778–10788. PMLR, 2020.
  57. 57.Yoon, T., Park, Y., Ryu, E. K., and Wang, Y. Robust probabilistic time series forecasting. arXiv preprint arXiv:2202.11910, 2022.
  58. 58.Yu, H.-F., Rao, N., and Dhillon, I. S. Temporal Regularized Matrix Factorization for High-dimensional Time Series Prediction. NIPS, pp. 847–855, 2016.
  59. 59.Zhao, H., Zhang, S., Wu, G., Moura, J. M. F., Costeira, J. P., and Gordon, G. J. Adversarial Multiple Source Domain Adaptation. Advances in neural information processing systems, pp. 8559–8570, 2018.
  60. 60.Zhou, H., Zhang, S., Peng, J., Zhang, S., Li, J., Xiong, H., and Zhang, W. Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting. In Proceedings of the AAAI Conference on Artificial Intelligence, 2021. URL http://arxiv.org/abs/2012.07436. arXiv: 2012.07436.

Citation

MLA
Jin, X., et al. “Domain Adaptation for Time Series Forecasting via Attention Sharing”. International Conference on Machine Learning, vol. 162, 2022, pp. 10280–97, https://proceedings.mlr.press/v162/jin22d.html.
APA
Jin, X., Park, Y., Maddix, D., Wang, H., & Wang, Y. (2022). Domain Adaptation for Time Series Forecasting via Attention Sharing. International Conference on Machine Learning, 162, 10280–10297. https://proceedings.mlr.press/v162/jin22d.html
Chicago
Jin, X., Y. Park, D. Maddix, H. Wang, and Y. Wang. 2022. “Domain Adaptation for Time Series Forecasting via Attention Sharing”. International Conference on Machine Learning 162: 10280–97. https://proceedings.mlr.press/v162/jin22d.html.
Harvard
Jin, X. et al. (2022) “Domain Adaptation for Time Series Forecasting via Attention Sharing”, International Conference on Machine Learning. PMLR, pp. 10280–10297. Available at: https://proceedings.mlr.press/v162/jin22d.html.
Vancouver
1. Jin X, Park Y, Maddix D, Wang H, Wang Y (2022) Domain Adaptation for Time Series Forecasting via Attention Sharing. In: International Conference on Machine Learning. PMLR, pp 10280–10297

BibTeX

@InProceedings{pmlr-v162-jin22d,
  title = 	 {Domain Adaptation for Time Series Forecasting via Attention Sharing},
  author =       {Jin, Xiaoyong and Park, Youngsuk and Maddix, Danielle and Wang, Hao and Wang, Yuyang},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {10280--10297},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/jin22d/jin22d.pdf},
  url = 	 {https://proceedings.mlr.press/v162/jin22d.html},
  abstract = 	 {Recently, deep neural networks have gained increasing popularity in the field of time series forecasting. A primary reason for their success is their ability to effectively capture complex temporal dynamics across multiple related time series. The advantages of these deep forecasters only start to emerge in the presence of a sufficient amount of data. This poses a challenge for typical forecasting problems in practice, where there is a limited number of time series or observations per time series, or both. To cope with this data scarcity issue, we propose a novel domain adaptation framework, Domain Adaptation Forecaster (DAF). DAF leverages statistical strengths from a relevant domain with abundant data samples (source) to improve the performance on the domain of interest with limited data (target). In particular, we use an attention-based shared module with a domain discriminator across domains and private modules for individual domains. We induce domain-invariant latent features (queries and keys) and retrain domain-specific features (values) simultaneously to enable joint training of forecasters on source and target domains. A main insight is that our design of aligning keys allows the target domain to leverage source time series even with different characteristics. Extensive experiments on various domains demonstrate that our proposed method outperforms state-of-the-art baselines on synthetic and real-world datasets, and ablation studies verify the effectiveness of our design choices.}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/