Domain Adaptation for Time Series Forecasting via Attention Sharing
Xiaoyong JinYoungsuk ParkDanielle C. MaddixHao WangYuyang Wang
Proposes an end-to-end domain adaptation framework for time series forecasting that aligns shared attention queries and keys across data-rich and data-scarce domains while preserving domain-specific values to improve multi-horizon predictions.
Modern predictive decision systems—such as retail inventory management, cloud resource provisioning, and vehicle control—increasingly rely on deep neural networks to produce accurate time series forecasts. However, these models require vast amounts of data to capture complex temporal dynamics effectively. In many real-world operational environments, organizations face severe data scarcity, such as limited historical observations (cold-start scenarios) or few available monitoring series (few-shot scenarios). When organizations attempt to transfer models trained on rich external data sources to these scarce target environments, they encounter domain shift—a distributional discrepancy that severely degrades forecasting accuracy.
The article introduces and evaluates the Domain Adaptation Forecaster (DAF), an end-to-end multi-horizon forecasting framework designed to solve data scarcity by transferring statistical strengths from a data-rich source domain to a data-scarce target domain.
To address this challenge without forcing incompatible features across disparate datasets, the authors designed a hybrid neural architecture. The framework maintains private encoders and decoders for each domain to handle domain-specific measurements and scales, while implementing a shared, attention-based module governed by a domain discriminator. Using adversarial training, the model forces structural temporal query and key patterns into a shared, domain-invariant latent space while keeping concrete values domain-specific. The model simultaneously performs input reconstruction and future-step extrapolation to stabilize training. The authors validated DAF across simulated cold-start and few-shot conditions as well as real-world benchmark datasets covering electricity consumption, traffic occupancy, store sales, and web page visits, comparing it against both single-domain models and cross-domain adaptation baselines using Normalized Deviation as the primary error metric.
The evaluation yielded several key findings:
- DAF consistently outperformed or matched all single-domain and cross-domain baselines across synthetic and real-world benchmarks, demonstrating the strongest accuracy gains in the most severely data-constrained settings.
- Cross-domain models trained jointly and end-to-end achieved markedly superior performance compared to two-stage fine-tuning approaches (such as DATSING) or traditional single-domain forecasters.
- In real-world multi-horizon transfer tasks, DAF lowered forecasting error across multiple cross-domain pairs (for example, achieving an error metric of 0.125 on electricity-to-traffic adaptations compared to 0.141 for DeepAR), outperforming recurrent-neural-network-based adaptation variants.
- Ablation analyses confirmed that keeping value embeddings domain-specific while enforcing query and key domain-invariance via adversarial training is the most critical design factor driving performance.
- DAF exhibited flexibility across distinct sampling frequencies, successfully transferring knowledge between hourly datasets (traffic and electricity) and daily datasets (sales and web traffic).
These findings demonstrate that organizations do not need to rely solely on expensive or slow target-domain data collection to deploy high-performing deep forecasting models. By aligning temporal dynamics across domains rather than raw output values, organizations can significantly mitigate the operational and financial risks of cold starts and small-sample deployments. This enables faster deployment cycles, reduced infrastructure costs for fine-tuning separate large-scale models, and improved decision reliability in newly launched operations.
For engineering and operational teams facing time series data scarcity, adopting an attention-sharing domain adaptation architecture represents a viable alternative to single-domain models or generic pre-training workflows. Organizations evaluating this approach should prioritize end-to-end joint training pipelines over staged fine-tuning. Before deploying in mission-critical environments, practitioners should run pilot validations across domain pairs to select optimal kernel and network hyperparameters.
While the empirical results are robust across tested benchmarks, the authors note that formal theoretical justifications for domain-invariant attention spaces remain an open area of inquiry. Furthermore, the empirical validation in the article is focused on univariate forecasting setups. Decision-makers should exercise appropriate caution when attempting to generalize these findings directly to complex, highly correlated multivariate forecasting systems without prior testing.
- Book: Domain-Adversarial Training of Neural Networks, Yaroslav Ganin et al. (2016). Its domain-adversarial training framework explains the domain discriminator and invariant-feature alignment that underpin DAF’s shared attention module.
- Paper: Domain Separation Networks, Konstantinos Bousmalis et al. (2016). Its separation of shared and domain-private representations provides a direct conceptual precursor to DAF’s shared and private modules.
- Paper: A Dual-Stage Attention-Based Recurrent Neural Network for Time Series Prediction, Yao Qin et al. (2017). Its dual-stage attention model establishes how attention can select relevant input features and temporal context for time-series forecasting.
- Paper: A theory of learning from different domains, Shai Ben-David et al. (2010). Its bounds formalize how source–target divergence constrains adaptation, grounding DAF’s effort to transfer useful information across domains.
- Paper: Domain Adaptation for Time Series Under Feature and Label Shifts, Huan He et al. (2023). It carries time-series domain adaptation into feature- and label-shift settings, extending the problem beyond DAF’s attention-based transfer across domains.
