TACTiS: Transformer-Attentional Copulas for Time Series

Alexandre DrouinÉtienne MarcotteNicolas Chapados

article2022ICML54 citations

Introduces TACTiS, a transformer architecture that models non-parametric copulas through attention mechanisms to deliver state-of-the-art multivariate probabilistic forecasting and interpolation across irregular, unaligned, and missing time series data.

Listen

Real-world decision-making in domains such as finance, healthcare, and energy management relies heavily on estimating time-varying quantities and understanding their associated predictive uncertainties. However, observational data in these settings rarely match the rigid assumptions required by classical statistical methods. Industrial time series frequently feature irregular sampling frequencies, missing values, misaligned timestamps, and complex, non-standard probability distributions across hundreds of related variables. Developing a robust, unified framework capable of handling these real-world data imperfections without requiring bespoke domain engineering remains a critical challenge.

The article evaluates and demonstrates Transformer-Attentional Copulas for Time Series (TACTiS), a deep learning model designed for large-scale multivariate probabilistic prediction. The core objective is to accurately estimate the joint predictive distribution over arbitrary missing values—enabling unified forecasting and interpolation—while flexibly capturing complex dependencies across diverse, high-dimensional time series.

To achieve this, the authors develop an architecture combining a transformer encoder with an attention-based decoder that estimates flexible, non-parametric copulas (statistical functions that separate joint dependency structures from individual variable distributions) without relying on restrictive parametric assumptions like Gaussian distributions. The framework uses normalizing flows to capture complex individual distributions alongside an autoregressive attention mechanism to model cross-variable dependencies. The authors conduct mathematical proofs of validity, simulation studies on synthetic processes, and an empirical backtesting evaluation using five diverse, high-dimensional real-world benchmarks (electricity, economic indicators, air quality, solar energy, and traffic) ranging from 107 to 862 variables.

The findings show that TACTiS achieves state-of-the-art predictive accuracy across real-world datasets, obtaining the lowest overall average rank (1.6) against competitive deep learning and classical baselines. The model successfully performs both forecasting and interpolation tasks within a single architecture simply by adjusting data masking patterns, accurately reconstructing hidden trajectory gaps in stochastic processes. In addition, synthetic tests confirm that TACTiS natively processes unaligned, irregularly sampled series. Computational efficiency experiments demonstrate that the model can be trained on high-dimensional data using random subsets of series (bagging) without sacrificing accuracy, while an ablation study confirms that the self-attention encoder and attentional copula decoder are both essential drivers of its strong predictive performance.

These results indicate that organizations can replace fragmented, problem-specific time series pipelines with a single, general-purpose model. This versatility reduces operational maintenance, lowers software development costs, and improves risk mitigation by delivering high-fidelity uncertainty estimates directly from raw, imperfect data streams. Unlike traditional models that require complete, aligned arrays, TACTiS accommodates real-world data irregularities without extensive manual preprocessing.

Organizations seeking to implement large-scale multivariate forecasting should consider adopting attention-based copula frameworks, particularly when managing multi-sensor systems or unaligned data streams. Practitioners should implement subset bagging during training to reduce computational resource usage while maintaining a sufficiently large bag size (at least 5 to 10 series) to preserve cross-series correlation structures. Future development should focus on testing time-specific positional encodings, accelerating sampling inference speeds, and evaluating the architecture as a cross-domain foundation model for time series forecasting with minimal historical observations.

Readers should note certain operational boundaries. Training dynamics can exhibit plateaus during intermediate learning phases, and marginal normalizing flows can struggle when fitting strictly discrete distributions. Furthermore, when sampling forecasts across hundreds of variables, the computational burden remains noticeable because bagging is applied only during training. Nonetheless, backed by theoretical convergence proofs and rigorous backtesting, confidence in the model's core predictive capabilities is high.

Cover for TACTiS: Transformer-Attentional Copulas for Time Series

Abstract

The estimation of time-varying quantities is a fundamental component of decision making in fields such as healthcare and finance. However, the practical utility of such estimates is limited by how accurately they quantify predictive uncertainty. In this work, we address the problem of estimating the joint predictive distribution of high-dimensional multivariate time series. We propose a versatile method, based on the transformer architecture, that estimates joint distributions using an attention-based decoder that provably learns to mimic the properties of non-parametric copulas. The resulting model has several desirable properties: it can scale to hundreds of time series, supports both forecasting and interpolation, can handle unaligned and non-uniformly sampled data, and can seamlessly adapt to missing data during training. We demonstrate these properties empirically and show that our model produces state-of-the-art predictions on multiple real-world datasets.

Table of Contents

  • 1. Introduction
  • 2. Background
  • 2.1. Problem Setting
  • 2.2. Transformers
  • 2.3. Copulas
  • 3. Related Work
  • 4. The TACTiS Model
  • 4.1. Encoder
  • 4.2. Decoder
  • 4.3. Training Procedure
  • 5. Experiments
  • 5.1. Empirical Validation of Attentional Copulas
  • 5.2. Forecasting: Comparison to the State of the Art
  • 5.3. Model Flexibility
  • 6. Discussion
  • Acknowledgements
  • References
  • A. Theory: Proof of Theorem 1
  • B. Implementation Details
  • B.1. Libraries Used
  • B.2. Inverting the Marginal Flows
  • B.3. Bagging: Efficient Training in High Dimensions
  • B.4. Data Normalization
  • C. Forecasting Benchmark
  • C.1. Datasets
  • C.2. Training Procedure
  • C.2.1. DEEP LEARNING MODELS
  • C.2.2. CLASSICAL MODELS
  • C.3. Hyperparameter Search Protocol
  • C.4. Hyperparameter Ranges
  • C.5. Backtesting Protocol
  • C.5.1. EXCEPTIONS FOR TACTIS-TT
  • C.6. Metrics and Additional Results
  • D. Additional Experiments
  • D.1. Can Attentional Copulas Recover a Ground-Truth Copula?
  • D.2. Can TACTiS Learn to Interpolate?
  • D.3. Can TACTiS Learn from Unaligned/Non-Uniformly Sampled Time Series?
  • D.4. Ablation Study
  • D.4.1. TACTIS-GC: USING ECDFS AND A GAUSSIAN COPULA IN THE DECODER
  • D.4.2. TACTIS-IC: USING A TRIVIAL COPULA INSTEAD OF AN ATTENTIONAL COPULA
  • D.4.3. RELATIVE IMPORTANCE OF TACTIS-TT KEY COMPONENTS
  • E. A Deeper Dive into TACTiS Models
  • E.1. Some Good and Bad Forecasts
  • E.2. Looking into Learned Marginal Distributions
  • E.3. Learning Dependencies Between Variables
  • E.3.1. IMPACT OF BAGGING ON INTER-SERIES CORRELATIONS

Knowls

  1. Knowl 1 — Multivariate Time Series Missing Value Estimation Formulation

    definition

    The general problem of multivariate time series prediction is formulated as estimating the joint predictive distribution of missing values across arbitrary time points and variables. Formally, let S={X1,…,Xm}\mathcal{S} = \{X_1, \dots, X_m\} denote a dataset of multivariate time series. For a given realization X={xi∈Rli}i=1nX = \{x_i \in \mathbb{R}^{l_i}\}_{i=1}^n comprising nn univariate time series of arbitrary lengths li∈N+l_i \in \mathbb{N}^+:

    1. Each series xi=[xi1,…,xili]⊤x_i = [x_{i1}, \dots, x_{il_i}]^\top is paired with a Boolean observation mask mi∈{0,1}lim_i \in \{0, 1\}^{l_i}, where mij=1m_{ij} = 1 if xijx_{ij} is observed and mij=0m_{ij} = 0 if missing.
    2. A matrix of time-varying covariates Ci=[ci1,…,cili]∈Rp×liC_i = [c_{i1}, \dots, c_{il_i}] \in \mathbb{R}^{p \times l_i} provides pp-dimensional auxiliary attributes for each observation.
    3. A strictly increasing vector of timestamps ti=[ti1,…,tili]⊤∈Rlit_i = [t_{i1}, \dots, t_{il_i}]^\top \in \mathbb{R}^{l_i} (tij<ti,j+1t_{ij} < t_{i,j+1}) specifies measurement times, natively supporting irregularly sampled and mutually unaligned series.

    The target objective is to estimate the joint conditional probability distribution of all missing values X(m)={xi(m)}i=1nX^{(m)} = \{x_i^{(m)}\}_{i=1}^n given all observed values X(o)={xi(o)}i=1nX^{(o)} = \{x_i^{(o)}\}_{i=1}^n, covariates C={Ci}i=1n\mathcal{C} = \{C_i\}_{i=1}^n, and timestamps T={ti}i=1n\mathcal{T} = \{t_i\}_{i=1}^n:

    P({xi(m)}i=1n  |  {xi(o),Ci,ti}i=1n)P\left(\{x_i^{(m)}\}_{i=1}^n \;\middle|\; \{x_i^{(o)}, C_i, t_i\}_{i=1}^n\right)

    Standard downstream tasks are specified via the mask structure: setting the final tt elements of each mim_i to 0 defines a tt-step probabilistic forecasting task, whereas setting interior elements to 0 defines an interpolation task.

  2. Knowl 2 — TACTiS Joint Density Decomposition and Marginal Flow Parameterization

    model/method

    Transformer-Attentional Copulas for Time Series (TACTiS) estimates the joint predictive density of nmn_m missing values x(m)=(x1(m),…,xnm(m))x^{(m)} = (x_1^{(m)}, \dots, x_{n_m}^{(m)}) conditioned on representations produced by a transformer encoder from observed tokens Z(o)={z1(o),…,zno(o)}Z^{(o)} = \{z_1^{(o)}, \dots, z_{n_o}^{(o)}\} and missing tokens Z(m)={z1(m),…,znm(m)}Z^{(m)} = \{z_1^{(m)}, \dots, z_{n_m}^{(m)}\}.

    For each token (i,j)(i, j), the encoder computes an input embedding combining the masked value, covariates, and mask indicator:

    eij=Embedθemb(xij⋅mij,cij,mij)e_{ij} = \text{Embed}_{\theta_{\text{emb}}}(x_{ij} \cdot m_{ij}, c_{ij}, m_{ij})

    eij′=eijdemb+pije'_{ij} = e_{ij}\sqrt{d_{\text{emb}}} + p_{ij}

    where pij∈Rdembp_{ij} \in \mathbb{R}^{d_{\text{emb}}} is a sinusoidal positional encoding of timestamp tijt_{ij}. Embeddings are processed through multi-head self-attention and normalization layers to yield token encodings zijz_{ij}.

    The decoder factorizes the joint probability density gϕg_\phi via Sklar's theorem:

    ight) = c_{\phi_c}\left(F_{\phi_1}\left(x_1^{(m)} ight), \dots, F_{\phi_{n_m}}\left(x_{n_m}^{(m)} ight)\right) \prod_{k=1}^{n_m} f_{\phi_k}\left(x_k^{(m)} ight)$$ where $\phi = \{\phi_c, \phi_1, \dots, \phi_{n_m}\}$ are distributional parameters generated by neural modules conditioned on $Z^{(o)}$ and $Z^{(m)}$. Each marginal cumulative distribution function (CDF) $F_{\phi_k}: \mathbb{R} \to [0, 1]$ is modeled using a Deep Sigmoidal Flow (DSF) with the terminal logit layer removed to ensure strict monotonicity, differentiability, and output in $[0, 1]$. Marginal parameters are output by a shared network $\phi_k = \text{MarginalParams}_{\theta_F}(z_k^{(m)})$, and marginal densities $f_{\phi_k}(x_k^{(m)}) = \frac{\partial}{\partial x_k^{(m)}} F_{\phi_k}(x_k^{(m)})$ are computed analytically.
  3. Knowl 3 — Attentional Copula and Attention-Based Conditioner

    model/method

    The attentional copula parametrizes a flexible, non-parametric joint copula density cϕcc_{\phi_c} over uniform random variables uk=Fϕk(xk(m))∈[0,1]u_k = F_{\phi_k}(x_k^{(m)}) \in [0, 1] for k∈{1,…,nm}k \in \{1, \dots, n_m\}. Given an arbitrary permutation π=[π1,…,πnm]\pi = [\pi_1, \dots, \pi_{n_m}] of the indices {1,…,nm}\{1, \dots, n_m\}, the copula density is autoregressively factorized as:

    cϕcπ(u1,…,unm)=cϕc1π(uπ1)∏k=2nmcϕckπ(uπk  |  uπ1,…,uπk−1)c_{\phi_c^\pi}(u_1, \dots, u_{n_m}) = c_{\phi_{c1}^\pi}(u_{\pi_1}) \prod_{k=2}^{n_m} c_{\phi_{ck}^\pi}\left(u_{\pi_k} \;\middle|\; u_{\pi_1}, \dots, u_{\pi_{k-1}}\right)

    The first conditional density is defined as the uniform density: cϕc1π(uπ1)=1c_{\phi_{c1}^\pi}(u_{\pi_1}) = 1 for uπ1∈[0,1]u_{\pi_1} \in [0, 1]. For subsequent steps k>1k > 1, an attention-based conditioner computes the conditional distribution parameters ϕckπ\phi_{ck}^\pi by querying a memory containing observed tokens and previous missing tokens in the permutation:

    Mk={(z1(o),u1(o)),…,(zno(o),uno(o)),(zπ1(m),uπ1(m)),…,(zπk−1(m),uπk−1(m))}\mathcal{M}_k = \left\{ \left(z_1^{(o)}, u_1^{(o)}\right), \dots, \left(z_{n_o}^{(o)}, u_{n_o}^{(o)}\right), \left(z_{\pi_1}^{(m)}, u_{\pi_1}^{(m)}\right), \dots, \left(z_{\pi_{k-1}}^{(m)}, u_{\pi_{k-1}}^{(m)}\right) \right\}

    where ui(o)u_i^{(o)} are the flow-transformed observed token values. For each (z,u)∈Mk(z, u) \in \mathcal{M}_k, key and value projections k=Keyθk(z,u)∈Rdattk = \text{Key}_{\theta_k}(z, u) \in \mathbb{R}^{d_{\text{att}}} and v=Valueθv(z,u)∈Rdattv = \text{Value}_{\theta_v}(z, u) \in \mathbb{R}^{d_{\text{att}}} are generated. A query q=Queryθq(zπk(m))∈Rdattq = \text{Query}_{\theta_q}(z_{\pi_k}^{(m)}) \in \mathbb{R}^{d_{\text{att}}} is constructed from the target missing token. Multi-head cross-attention and feed-forward residual layers produce an updated state z′′z'', from which parameters are predicted: ϕckπ=DistParamsθdist(z′′)\phi_{ck}^\pi = \text{DistParams}_{\theta_{\text{dist}}}(z'').

    Each conditional distribution cϕckπc_{\phi_{ck}^\pi} is modeled as a piecewise constant distribution over BB equal-width bins on [0,1][0, 1], enabling the representation of multimodal dependencies without parametric constraints.

  4. Knowl 4 — Validity of Attentional Copulas via Permutation-Averaged Optimization

    theoretical result

    Let gϕπ(x1(m),…,xnm(m))g_{\phi^\pi}(x_1^{(m)}, \dots, x_{n_m}^{(m)}) be the joint density estimator factorized according to permutation π∈Π\pi \in \Pi, where Π\Pi is the set of all permutations of {1,…,nm}\{1, \dots, n_m\}.

    Theorem. If model parameters Θ={θenc,θdec}\Theta = \{\theta_{\text{enc}}, \theta_{\text{dec}}\} minimize the expected negative log-likelihood across permutations sampled uniformly at random:

    ight) \right]$$ then the attentional copula $c_{\phi_c^\pi}$ embedded within $g_{\phi^\pi}$ is a valid copula density. **Scope and Conditions of Validity:** 1. By construction, each conditional component of the attentional copula has support restricted to the unit interval $[0, 1]$, ensuring the joint support is $[0, 1]^{n_m}$. 2. The expected loss over $\pi \sim \Pi$ equals the negative log of a geometric mean of densities across permutations. Minimizing this objective achieves equality with the arithmetic mean if and only if $g_{\phi^\pi}$ (and thus $c_{\phi_c^\pi}$) is permutation invariant for all $\pi \in \Pi$. 3. Because the leading element in any permutation is fixed to $\mathcal{U}[0, 1]$ by definition ($c_{\phi_{c1}^\pi}(u) = 1$), permutation invariance guarantees that the marginal distribution of every variable $u_k$ is identically $\mathcal{U}[0, 1]$.
  5. Knowl 5 — TACTiS Permutation-Averaged Training Objective

    equation

    The parameter set Θ={θenc,θdec}\Theta = \{\theta_{\text{enc}}, \theta_{\text{dec}}\} of the encoder and decoder is optimized by minimizing the expected negative log-likelihood of the factorized density estimator over uniformly distributed random permutations π∼Π\pi \sim \Pi and training time series samples X∼SX \sim \mathcal{S}:

    Θ∗=arg⁡min⁡ΘEπ∼Π,X∼S[−log⁡gϕπ(x1(m),…,xnm(m))]\Theta^* = \arg\min_\Theta \mathbb{E}_{\pi \sim \Pi, X \sim \mathcal{S}} \left[ -\log g_{\phi^\pi}\left(x_1^{(m)}, \dots, x_{n_m}^{(m)}\right) \right]

    where Π\Pi is the set of all permutations of missing token indices {1,…,nm}\{1, \dots, n_m\}, and gϕπg_{\phi^\pi} is the joint density factorized along permutation π\pi:

    ight) = -\log c_{\phi_c^\pi}\left(F_{\phi_1}\left(x_1^{(m)} ight), \dots, F_{\phi_{n_m}}\left(x_{n_m}^{(m)} ight)\right) - \sum_{k=1}^{n_m} \log f_{\phi_k}\left(x_k^{(m)} ight)$$ In stochastic gradient descent implementations, a single permutation $\pi$ is sampled uniformly at random for each mini-batch without differentiating through permutation sampling.
  6. Knowl 6 — TACTiS Autoregressive Sampling Algorithm

    algorithm

    Joint trajectories of missing tokens are generated by autoregressively sampling uniform latent variables from the attentional copula along a permutation π\pi, inverting the Deep Sigmoidal Flows (DSF), and destandardizing the outputs.

    Input: Observed token encodings Z(o)Z^{(o)}, missing token encodings Z(m)Z^{(m)}, flow-transformed observed values u(o)u^{(o)}, permutation π=[π1,…,πnm]\pi = [\pi_1, \dots, \pi_{n_m}], marginal flow models Fϕ1,…,FϕnmF_{\phi_1}, \dots, F_{\phi_{n_m}}, series standardization means meani\text{mean}_i and variances variancei\text{variance}_i
    Output: Predicted missing values X(m)=(x1(m),…,xnm(m))X^{(m)} = (x_1^{(m)}, \dots, x_{n_m}^{(m)})
    Sample uπ1(m)∼U[0,1]u_{\pi_1}^{(m)} \sim \mathcal{U}[0, 1]
    Initialize conditioner memory M=Z(o)∪{u(o)}∪{(zπ1(m),uπ1(m))}\mathcal{M} = Z^{(o)} \cup \{u^{(o)}\} \cup \{(z_{\pi_1}^{(m)}, u_{\pi_1}^{(m)})\}
    for k=2k = 2 to nmn_m do
        Compute keys K=Keyθk(M)K = \text{Key}_{\theta_k}(\mathcal{M}) and values V=Valueθv(M)V = \text{Value}_{\theta_v}(\mathcal{M})
        Compute query q=Queryθq(zπk(m))q = \text{Query}_{\theta_q}(z_{\pi_k}^{(m)})
        Compute updated representation z′′z'' via multi-head cross-attention over K,VK, V with query qq
        Compute conditional distribution parameters ϕckπ=DistParamsθdist(z′′)\phi_{ck}^\pi = \text{DistParams}_{\theta_{\text{dist}}}(z'')
        Sample uπk(m)∼cϕckπ(⋅)u_{\pi_k}^{(m)} \sim c_{\phi_{ck}^\pi}(\cdot) from the piecewise constant distribution on [0,1][0, 1]
        Update memory M=M∪{(zπk(m),uπk(m))}\mathcal{M} = \mathcal{M} \cup \{(z_{\pi_k}^{(m)}, u_{\pi_k}^{(m)})\}
    end for
    for k=1k = 1 to nmn_m do
        Rescale sample to avoid unregularized flow tails: u~k=0.05+0.9⋅uk(m)\tilde{u}_k = 0.05 + 0.9 \cdot u_k^{(m)}
        Invert marginal CDF via binary search on monotonic DSF: x~k(m)=Fϕk−1(u~k)\tilde{x}_k^{(m)} = F_{\phi_k}^{-1}(\tilde{u}_k)
        Destandardize: xk(m)=variancek⋅x~k(m)+meankx_k^{(m)} = \sqrt{\text{variance}_k} \cdot \tilde{x}_k^{(m)} + \text{mean}_k
    end for
    return X(m)X^{(m)}

    Binary search is used for flow inversion due to the strict monotonicity of DSF. The [0.05,0.95][0.05, 0.95] rescaling prevents numerical instability and artifacts caused by unregularized extreme tails in normalizing flows.

  7. Knowl 7 — Variable Bagging for Scalable High-Dimensional Training

    model/method

    To train attention-based copula models on high-dimensional datasets with nn time series without incurring prohibitive memory and compute footprints, TACTiS utilizes variable bagging during training:

    1. Mini-Batch Subsampling: Each training mini-batch samples a random subset ("bag") of b≪nb \ll n univariate series (e.g., b=20b = 20). Encoder self-attention and decoder attentional copulas operate strictly over the tokens belonging to the selected bb series.
    2. Epoch Batch Scaling: To compensate for reduced data volume per batch, the number of mini-batches per epoch is adjusted to ⌊(1600/batch size)×(n/b)⌋\lfloor (1600 / \text{batch size}) \times (n / b) \rfloor.
    3. Full-Dimensional Inference: Because attention mechanisms operate on arbitrary token set sizes without modifying network weights, the model trained with bag size bb is deployed at inference time on all nn time series simultaneously.

    Empirical validation shows that training with bag sizes b≥5b \ge 5 converges to the same inter-series correlation structure and predictive accuracy as larger bags, whereas training with b=1b = 1 prevents the attention conditioner from learning cross-series dependency masks.

  8. Knowl 8 — Factorized Temporal Transformer Layers (TACTiS-TT)

    model/method

    For aligned multivariate time series with nn variables and maximum sequence length lmax⁡=max⁡ilil_{\max} = \max_i l_i, full self-attention across all n⋅lmax⁡n \cdot l_{\max} tokens requires O((n⋅lmax⁡)2)O((n \cdot l_{\max})^2) time and memory.

    The TACTiS-TT variant replaces full self-attention in the encoder with factorized two-dimensional attention:

    1. Intra-Series Temporal Attention: Multi-head self-attention is computed independently across the temporal tokens of each univariate series (xi1,…,xi,li)(x_{i1}, \dots, x_{i,l_i}).
    2. Inter-Series Spatial Attention: Multi-head self-attention is computed across the nn series tokens at each fixed time step jj, (x1j,…,xnj)(x_{1j}, \dots, x_{nj}).

    This factorization reduces computational complexity to O(n2lmax⁡+nlmax⁡2)O(n^2 l_{\max} + n l_{\max}^2), enabling TACTiS-TT to scale to hundreds of time series and time steps while requiring the series to be temporally aligned.

  9. Knowl 9 — Probabilistic Forecasting Benchmark Results

    data/table

    The multivariate probabilistic forecasting performance of TACTiS-TT was benchmarked against deep learning methods (GPVar, TempFlow, TimeGrad) and classical baselines (Auto-ARIMA, ETS) across five datasets from the Monash Time Series Forecasting Repository. Evaluation was conducted using sequential backtesting with nB=6n_B = 6 folds, reporting CRPS-Sum means with Newey-West autocorrelation-corrected standard errors (3 lags) and average ranks across datasets.

    Model electricity fred-md kdd-cup solar-10min traffic Avg. Rank
    Auto-ARIMA 0.077 0.016 0.043 0.005 0.625 0.066 0.994 0.216 0.222 0.005 4.7 0.3
    ETS 0.059 0.011 0.037 0.010 0.408 0.030 0.678 0.097 0.353 0.011 4.4 0.3
    TempFlow 0.075 0.024 0.095 0.004 0.250 0.010 0.507 0.034 0.242 0.020 3.9 0.2
    TimeGrad 0.067 0.028 0.094 0.030 0.326 0.024 0.540 0.044 0.126 0.019 3.6 0.3
    GPVar 0.035 0.011 0.067 0.008 0.290 0.005 0.254 0.028 0.145 0.010 2.7 0.2
    TACTiS-TT 0.021 0.005 0.042 0.009 0.237 0.013 0.311 0.061 0.071 0.008 1.6 0.2

    TACTiS-TT achieved the lowest CRPS-Sum on 3 of the 5 datasets (electricity, kdd-cup, traffic) and the best overall average rank (1.6), outperforming all deep learning baselines on fred-md and all baselines except GPVar on solar-10min.

  10. Knowl 10 — Ablation Analysis of TACTiS Decoder Components

    empirical result

    To evaluate the specific contributions of the encoder self-attention, Deep Sigmoidal Flow (DSF) marginals, and attentional copulas, TACTiS-TT was compared against two ablations:

    1. TACTiS-IC: Replaces the attentional copula with a trivial independent copula (uniform distribution on [0,1]nm[0, 1]^{n_m}), isolating the effect of copula dependencies while retaining DSF marginals.
    2. TACTiS-GC: Replaces the entire decoder with empirical cumulative distribution functions (ECDF) and a low-rank Gaussian copula (as in GPVar).

    Key empirical findings across datasets:

    • Encoder Impact: Both TACTiS-GC and TACTiS-IC remained competitive with prior state-of-the-art baselines, demonstrating that the self-attention transformer encoder is the primary baseline driver of predictive accuracy.
    • Marginal Flows vs. ECDF: On electricity (CRPS-Sum: 0.022 for TACTiS-IC vs. 0.054 for TACTiS-GC) and fred-md (0.042 vs. 0.050), TACTiS-IC outperformed TACTiS-GC, showing the critical role of DSF in modeling non-Gaussian, flexible marginals.
    • Attentional Copula vs. Independent Copula: While CRPS-Sum (a summation metric) showed modest differences on some datasets, the multivariate Energy Score revealed substantial advantages for the full attentional copula over TACTiS-IC on kdd-cup (2.93×1032.93 \times 10^3 vs. 5.24×1035.24 \times 10^3), solar-10min (2.88×1022.88 \times 10^2 vs. 9.14×1029.14 \times 10^2), and traffic (3.103.10 vs. 10.6910.69).
  11. Knowl 11 — Probabilistic Interpolation and Non-Uniform Sampling Capabilities

    empirical result

    TACTiS demonstrates flexibility beyond standard forecasting by operating directly on masked, irregularly sampled, and unaligned time series:

    1. Probabilistic Interpolation on Stochastic Volatility Data: Univariate sequences of length 125 were generated from a latent stochastic volatility process (yt∣ht∼N(0,eht)y_t \mid h_t \sim \mathcal{N}(0, e^{h_t}), ht∣ht−1∼N(μ+ϕ(ht−1−μ),σ2)h_t \mid h_{t-1} \sim \mathcal{N}(\mu + \phi(h_{t-1} - \mu), \sigma^2) with μ=−9,ϕ=0.99,σ=0.04\mu = -9, \phi = 0.99, \sigma = 0.04, and xt=xt−1+ytx_t = x_{t-1} + y_t). For a missing central gap of 25 steps, TACTiS estimated conditional distributions evaluated against ground-truth MCMC posterior samples (Stan Oracle) via Wasserstein distance (WD) of Energy Scores. TACTiS achieved a mean WD of 0.03910.0391, substantially outperforming a deterministic linear interpolation baseline (mean WD 0.14590.1459).
    2. Unaligned and Non-Uniformly Sampled Series: TACTiS was trained on bivariate noisy sine processes where each series was sampled at independent, non-uniform random time intervals. By treating each observation as a distinct token with an explicit timestamp positional encoding, the general TACTiS architecture produced accurate probabilistic forecasts without requiring interpolation preprocessing or series alignment.
  12. Knowl 12 — Three-Phase Training Dynamics and Copula Initialization Sensitivity

    limitation

    Training TACTiS exhibits a distinct three-phase optimization trajectory that introduces risks during model fitting:

    1. Phase 1 (Marginal Fitting): The model rapidly fits the marginal distributions via the Deep Sigmoidal Flows while driving the attentional copula toward a trivial uniform distribution (independent variables).
    2. Phase 2 (Loss Plateau): The training loss and validation metrics plateau across multiple epochs with negligible visible improvement in prediction quality as the model transitions.
    3. Phase 3 (Non-Trivial Copula Learning): The model begins learning cross-variable dependencies in the attentional copula, leading to renewed loss reduction.

    Implications and Limitations:

    • Conventional early stopping criteria (e.g., stopping after 20 non-improving epochs) can terminate training prematurely during Phase 2, leaving the model stuck with an independent copula.
    • Setting the decoder MLP hidden dimension too small (e.g., <48< 48) significantly prolongs or prevents the transition from Phase 2 to Phase 3, requiring manually enlarged conditioner dimensions.

Coverage note — Deliberately omitted the specific qualitative forecast visualization plots from Appendix E.1 and the empirical copula validation on synthetic bivariate Clayton copula data (Appendix D.1), focusing instead on the full architectural specification, theoretical validity proof, benchmark tables, ablations, and key practical capabilities.

References

  1. 1.Aas, K., Czado, C., Frigessi, A., and Bakken, H. Pair-copula constructions of multiple dependence. Insurance: Mathematics and Economics, 44(2):182–198, 2009. URL https://EconPapers.repec.org/RePEc:eee:insuma:v:44:y:2009:i:2:p:182-198.
  2. 2.Alexandrov, A., Benidis, K., Bohlke-Schneider, M., Flunkert, V., Gasthaus, J., Januschowski, T., Maddix, D. C., Rangapuram, S., Salinas, D., Schulz, J., Stella, L., Turkmen, A. C., and Wang, Y. GluonTS: Probabilistic and Neural Time Series Modeling in Python. Journal of Machine Learning Research, 21(116):1–6, 2020. URL http://jmlr.org/papers/v21/19-820.html.
  3. 3.Bahdanau, D., Cho, K., and Bengio, Y. Neural machine translation by jointly learning to align and translate. In Bengio, Y. and LeCun, Y. (eds.), 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings, 2015. URL http://arxiv.org/abs/1409.0473.
  4. 4.Benidis, K., Rangapuram, S. S., Flunkert, V., Wang, B., Maddix, D., Turkmen, C., Gasthaus, J., Bohlke-Schneider, M., Salinas, D., Stella, L., Callot, L., and Januschowski, T. Neural forecasting: Introduction and literature overview. arXiv.org, 2020. URL http://arxiv.org/abs/2004.10240.
  5. 5.Bohlke-Schneider, M. and Salinas, D. personal communication, 2021.
  6. 6.Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258, 2021.
  7. 7.Box, G. E. P., Jenkins, G. M., Reinsel, G. C., and Ljung, G. M. Time series analysis: forecasting and control. John Wiley & Sons, fifth edition, 2015.
  8. 8.Chapados, N. and Bengio, Y. Augmented functional time series representation and forecasting with gaussian processes. In Platt, J., Koller, D., Singer, Y., and Roweis, S. (eds.), Advances in Neural Information Processing Systems, volume 20. Curran Associates, Inc., 2007. URL https://proceedings.neurips.cc/paper/2007/file/81e74d678581a3bb7a720b019f4f1a93-Paper.pdf.
  9. 9.Chen, Y., Kang, Y., Chen, Y., and Wang, Z. Probabilistic forecasting with temporal convolutional neural network. Neurocomputing, 399:491–501, 2020.
  10. 10.Choromanski, K. M., Likhosherstov, V., Dohan, D., Song, X., Gane, A., Sarlos, T., Hawkins, P., Davis, J. Q., Mohiuddin, A., Kaiser, L., Belanger, D. B., Colwell, L. J., and Weller, A. Rethinking attention with performers. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=Ua6zuk0WRH.
  11. 11.de Bezenac, E., Rangapuram, S. S., Benidis, K., Bohlke-Schneider, M., Kurle, R., Stella, L., Hasson, H., Gallinari, P., and Januschowski, T. Normalizing kalman filters for multivariate time series analysis. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M. F., and Lin, H. (eds.), Advances in Neural Information Processing Systems, volume 33, pp. 2995–3007. Curran Associates, Inc., 2020. URL https://proceedings.neurips.cc/paper/2020/file/1f47cef5e38c952f94c5d61726027439-Paper.pdf.
  12. 12.Gneiting, T. and Raftery, A. E. Strictly proper scoring rules, prediction, and estimation. Journal of the American statistical Association, 102(477):359–378, 2007.
  13. 13.Godahewa, R., Bergmeir, C., Webb, G. I., Hyndman, R. J., and Montero-Manso, P. Monash time series forecasting archive. In Neural Information Processing Systems Track on Datasets and Benchmarks, 2021. forthcoming.
  14. 14.Goodfellow, I., Bengio, Y., and Courville, A. Deep Learning. MIT Press, 2016. http://www.deeplearningbook.org.
  15. 15.Großer, J. and Okhrin, O. Copulae: An overview and recent developments. WIREs Computational Statistics, n/a(n/a): e1557, 2021. doi: https://doi.org/10.1002/wics.1557. URL https://wires.onlinelibrary.wiley.com/doi/abs/10.1002/wics.1557. A good initial read to give a higher level of understanding of what copulas are and what kinds of copulas have been studied.
  16. 16.Ho, J., Jain, A., and Abbeel, P. Denoising diffusion probabilistic models. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M. F., and Lin, H. (eds.), Advances in Neural Information Processing Systems, volume 33, pp. 6840–6851. Curran Associates, Inc., 2020. URL https://proceedings.neurips.cc/paper/2020/file/4c5bcfec8584af0d967f1ab10179ca4b-Paper.pdf.
  17. 17.Hochreiter, S. and Schmidhuber, J. Long short-term memory. Neural computation, 9(8):1735–1780, 1997.
  18. 18.Hoogeboom, E., Gritsenko, A. A., Bastings, J., Poole, B., van den Berg, R., and Salimans, T. Autoregressive diffusion models. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=Lm8T39vLDTE.
  19. 19.Huang, C.-W., Krueger, D., Lacoste, A., and Courville, A. Neural autoregressive flows. In Dy, J. and Krause, A. (eds.), Proceedings of the 35th International Conference on Machine Learning, volume 80 of Proceedings of Machine Learning Research, pp. 2078–2087. PMLR, 10–15 Jul 2018. URL https://proceedings.mlr.press/v80/huang18d.html.
  20. 20.Hyndman, R., Koehler, A. B., Ord, J. K., and Snyder, R. D. Forecasting with exponential smoothing: the state space approach. Springer Science & Business Media, 2008.
  21. 21.Hyndman, R., Athanasopoulos, G., Bergmeir, C., Caceres, G., Chhay, L., O’Hara-Wild, M., Petropoulos, F., Razbash, S., Wang, E., and Yasmeen, F. forecast: Forecasting functions for time series and linear models, 2022. URL https://pkg.robjhyndman.com/forecast/. R package version 8.16.
  22. 22.Hyndman, R. J. and Khandakar, Y. Automatic time series forecasting: the forecast package for R. Journal of Statistical Software, 26(3):1–22, 2008. doi: 10.18637/jss.v027.i03.
  23. 23.Kim, S., Shephard, N., and Chib, S. Stochastic volatility: Likelihood inference and comparison with ARCH models. The Review of Economic Studies, 65(3):361–393, 07 1998. ISSN 0034-6527. URL https://doi.org/10.1111/1467-937X.00050.
  24. 24.Koochali, A., Schichtel, P., Dengel, A., and Ahmed, S. Random noise vs state-of-the-art probabilistic forecasting methods: A case study on crps-sum discrimination ability. arXiv preprint arXiv:2201.08671, 2022.
  25. 25.Krupskii, P. and Joe, H. Flexible copula models with dynamic dependence and application to financial data. Econometrics and Statistics, 16:148–167, 2020. ISSN 2452-3062. doi: https://doi.org/10.1016/j.ecosta.2020.01.005. URL https://www.sciencedirect.com/science/article/pii/S2452306220300216.
  26. 26.Le Guen, V. and Thome, N. Probabilistic time series forecasting with shape and temporal diversity. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M. F., and Lin, H. (eds.), Advances in Neural Information Processing Systems, volume 33, pp. 4427–4440. Curran Associates, Inc., 2020. URL https://proceedings.neurips.cc/paper/2020/file/2f2b265625d76a6704b08093c652fd79-Paper.pdf.
  27. 27.Li, S., Jin, X., Xuan, Y., Zhou, X., Chen, W., Wang, Y.-X., and Yan, X. Enhancing the locality and breaking the memory bottleneck of transformer on time series forecasting. In Wallach, H., Larochelle, H., Beygelzimer, A., d'Alche-Buc, F., Fox, E., and Garnett, R. (eds.), Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc., 2019. URL https://proceedings.neurips.cc/paper/2019/file/6775a0635c302542da2c32aa19d86be0-Paper.pdf.
  28. 28.Lim, B. and Zohren, S. Time-series forecasting with deep learning: a survey. Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences, 379(2194): 20200209, 2021. doi: 10.1098/rsta.2020.0209. URL https://royalsocietypublishing.org/doi/abs/10.1098/rsta.2020.0209.
  29. 29.Lim, B., Arık, S. O., Loeff, N., and Pfister, T. Temporal Fusion Transformers for interpretable multi-horizon time series forecasting. International Journal of Forecasting, 37(4):1748–1764, 2021. ISSN 0169-2070. doi: https://doi.org/10.1016/j.ijforecast.2021.03.012. URL https://www.sciencedirect.com/science/article/pii/S0169207021000637.
  30. 30.Lin, T., Wang, Y., Liu, X., and Qiu, X. A survey of transformers, 2021. URL https://arxiv.org/abs/2106.04554.
  31. 31.Lopez-Paz, D., Hernandez-lobato, J., and Schölkopf, B. Semi-supervised domain adaptation with non-parametric copulas. In Pereira, F., Burges, C. J. C., Bottou, L., and Weinberger, K. Q. (eds.), Advances in Neural Information Processing Systems, volume 25. Curran Associates, Inc., 2012. URL https://proceedings.neurips.cc/paper/2012/file/8e98d81f8217304975ccb23337bb5761-Paper.pdf.
  32. 32.Makridakis, S., Spiliotis, E., and Assimakopoulos, V. Statistical and machine learning forecasting methods: Concerns and ways forward. PLoS ONE, 13(3):e0194889–26, 2018.
  33. 33.Makridakis, S., Spiliotis, E., Assimakopoulos, V., Chen, Z., Gaba, A., Tsetlin, I., and Winkler, R. L. The M5 uncertainty competition: Results, findings and conclusions. International Journal of Forecasting, 2021. ISSN 0169-2070. doi: https://doi.org/10.1016/j.ijforecast.2021.10.009. URL https://www.sciencedirect.com/science/article/pii/S0169207021001722.
  34. 34.Makridakis, S., Spiliotis, E., and Assimakopoulos, V. M5 accuracy competition: Results, findings, and conclusions. International Journal of Forecasting, 2022. ISSN 0169-2070. doi: https://doi.org/10.1016/j.ijforecast.2021.11.013. URL https://www.sciencedirect.com/science/article/pii/S0169207021001874.
  35. 35.Matheson, J. E. and Winkler, R. L. Scoring rules for continuous probability distributions. Management science, 22(10):1087–1096, 1976.
  36. 36.Mayer, A. and Wied, D. Estimation and inference in factor copula models with exogenous covariates, 2021. URL https://arxiv.org/abs/2107.03366.
  37. 37.Montero-Manso, P. and Hyndman, R. J. Principles and algorithms for forecasting groups of time series: Locality and globality. International Journal of Forecasting, 37(4):1632–1653, 2021. ISSN 0169-2070. doi: https://doi.org/10.1016/j.ijforecast.2021.03.004. URL https://www.sciencedirect.com/science/article/pii/S0169207021000558.
  38. 38.Müller, S., Hollmann, N., Arango, S. P., Grabocka, J., and Hutter, F. Transformers can do Bayesian inference. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=KSugKcbNf9.
  39. 39.Nelsen, R. B. An introduction to copulas. Springer Science & Business Media, second edition, 2007.
  40. 40.Newey, W. K. and West, K. D. A simple, positive semi-definite, heteroskedasticity and autocorrelation consistent covariance matrix. Econometrica, 55(3): 703–708, 1987. ISSN 00129682, 14680262. URL http://www.jstor.org/stable/1913610.
  41. 41.Newey, W. K. and West, K. D. Automatic Lag Selection in Covariance Matrix Estimation. The Review of Economic Studies, 61(4):631–653, 10 1994. ISSN 0034-6527. doi: 10.2307/2297912. URL https://doi.org/10.2307/2297912.
  42. 42.Oreshkin, B. N., Carpov, D., Chapados, N., and Bengio, Y. N-BEATS: Neural basis expansion analysis for interpretable time series forecasting. In International Conference on Learning Representations, 2020. URL https://openreview.net/forum?id=r1ecqn4YwB.
  43. 43.Papamakarios, G., Nalisnick, E., Rezende, D. J., Mohamed, S., and Lakshminarayanan, B. Normalizing flows for probabilistic modeling and inference. Journal of Machine Learning Research, 22(57):1–64, 2021. URL http://jmlr.org/papers/v22/19-1028.html.
  44. 44.Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S. Pytorch: An imperative style, high-performance deep learning library. In Wallach, H., Larochelle, H., Beygelzimer, A., d’Alche Buc, F., Fox, E., and Garnett, R. (eds.), Advances in Neural Information Processing Systems 32, pp. 8024–8035. Curran Associates, Inc., 2019. URL http://papers.neurips.cc/paper/9015-pytorch-an-imperative-style-high-performance-deep-learning-library.pdf.
  45. 45.Patton, A. J. A review of copula models for economic time series. Journal of Multivariate Analysis, 110:4–18, 2012.
  46. 46.Peterson, M. An Introduction to Decision Theory. Cambridge Introductions to Philosophy. Cambridge University Press, second edition, 2017. doi: 10.1017/9781316585061.
  47. 47.Pinson, P. and Tastu, J. Discrimination ability of the energy score. DTU Informatics, 2013.
  48. 48.R Core Team. R: A Language and Environment for Statistical Computing. R Foundation for Statistical Computing, Vienna, Austria, 2020. URL https://www.R-project.org/.
  49. 49.Rasul, K. PyTorchTS, 2021a. URL https://github.com/zalandoresearch/pytorch-ts.
  50. 50.Rasul, K. personal communication, 2021b.
  51. 51.Rasul, K., Seward, C., Schuster, I., and Vollgraf, R. Autoregressive denoising diffusion models for multivariate probabilistic time series forecasting. In Meila, M. and Zhang, T. (eds.), Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pp. 8857–8868. PMLR, 18–24 Jul 2021a. URL https://proceedings.mlr.press/v139/rasul21a.html.
  52. 52.Rasul, K., Sheikh, A.-S., Schuster, I., Bergmann, U. M., and Vollgraf, R. Multivariate probabilistic time series forecasting via conditioned normalizing flows. In International Conference on Learning Representations, 2021b. URL https://openreview.net/forum?id=WiGQBFuVRv.
  53. 53.Rémillard, B., Papageorgiou, N., and Soustra, F. Copula-based semiparametric models for multivariate time series. Journal of Multivariate Analysis, 110:30–42, 2012. ISSN 0047-259X. doi: https://doi.org/10.1016/j.jmva.2012.03.001. URL https://www.sciencedirect.com/science/article/pii/S0047259X1200070X. Special Issue on Copula Modeling and Dependence.
  54. 54.Salinas, D., Bohlke-Schneider, M., Callot, L., Medico, R., and Gasthaus, J. High-dimensional multivariate forecasting with low-rank Gaussian copula processes. Advances in Neural Information Processing Systems, 32: 6827–6837, 2019.
  55. 55.Salinas, D., Flunkert, V., Gasthaus, J., and Januschowski, T. DeepAR: Probabilistic forecasting with autoregressive recurrent networks. International Journal of Forecasting, 36(3):1181–1191, 2020.
  56. 56.Seabold, S. and Perktold, J. statsmodels: Econometric and statistical modeling with Python. In 9th Python in Science Conference, 2010.
  57. 57.Shih, S.-Y., Sun, F.-K., and Lee, H.-y. Temporal pattern attention for multivariate time series forecasting. Machine Learning, 108(8-9):1421–1441, 2019.
  58. 58.Shukla, S. N. and Marlin, B. Multi-time attention networks for irregularly sampled time series. In International Conference on Learning Representations, 2021a. URL https://openreview.net/forum?id=4c0J6lwQ4.
  59. 59.Shukla, S. N. and Marlin, B. M. A survey on principles, models and methods for learning from irregularly sampled time series, 2021b. URL https://arxiv.org/abs/2012.00168.
  60. 60.Sklar, A. Fonctions de repartition a` n dimensions et leurs marges. Publications de l’Institut Statistique de l’Universite de Paris, 8:229–231, 1959.
  61. 61.Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S. Deep unsupervised learning using nonequilibrium thermodynamics. In Bach, F. and Blei, D. (eds.), Proceedings of the 32nd International Conference on Machine Learning, volume 37 of Proceedings of Machine Learning Research, pp. 2256–2265, Lille, France, 07–09 Jul 2015. PMLR. URL https://proceedings.mlr.press/v37/sohl-dickstein15.html.
  62. 62.Spadon, G., Hong, S., Brandoli, B., Matwin, S., Rodrigues-Jr, J. F., and Sun, J. Pay attention to evolution: Time series forecasting with deep graph-evolution learning. IEEE Transactions on Pattern Analysis & Machine Intelligence, 2021. ISSN 1939-3539. doi: 10.1109/TPAMI.2021.3076155.
  63. 63.Stan Development Team. Stan modeling language users guide and reference manual, 2022. URL https://mc-stan.org.
  64. 64.Sun, C., Hong, S., Song, M., and Li, H. A review of deep learning methods for irregularly sampled medical time series data, 2020. URL https://arxiv.org/abs/2010.12493.
  65. 65.Tabak, E. G. and Turner, C. V. A family of nonparametric density estimation algorithms. Communications on Pure and Applied Mathematics, 66(2):145–164, 2013.
  66. 66.Tang, B. and Matteson, D. Probabilistic transformer for time series analysis. Advances in Neural Information Processing Systems, 34, 2021.
  67. 67.Tashiro, Y., Song, J., Song, Y., and Ermon, S. CSDI: Conditional score-based diffusion models for probabilistic time series imputation. In Advances in Neural Information Processing Systems, volume 34, 2021.
  68. 68.Tay, Y., Dehghani, M., Bahri, D., and Metzler, D. Efficient transformers: A survey, 2020. URL https://arxiv.org/abs/2009.06732.
  69. 69.Uria, B., Murray, I., and Larochelle, H. A deep and tractable density estimator. In International Conference on Machine Learning, volume 32, pp. 467–475. PMLR, 2014.
  70. 70.van den Oord, A., Dieleman, S., Zen, H., Simonyan, K., Vinyals, O., Graves, A., Kalchbrenner, N., Senior, A., and Kavukcuoglu, K. Wavenet: A generative model for raw audio. In Arxiv, 2016. URL https://arxiv.org/abs/1609.03499.
  71. 71.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is all you need. In Advances in neural information processing systems, pp. 5998–6008, 2017.
  72. 72.Wei, W. W. Multivariate time series analysis and applications. John Wiley & Sons, 2018.
  73. 73.Wiese, M., Knobloch, R., and Korn, R. Copula & marginal flows: Disentangling the marginal from its joint. arXiv preprint arXiv:1907.03361, 2019.
  74. 74.Williams, C. K. and Rasmussen, C. E. Gaussian processes for machine learning, volume 2. MIT press Cambridge, MA, 2006.
  75. 75.Wu, H., Xu, J., Wang, J., and Long, M. Autoformer: Decomposition transformers with auto-correlation for long-term series forecasting. In Ranzato, M., Beygelzimer, A., Dauphin, Y., Liang, P., and Vaughan, J. W. (eds.), Advances in Neural Information Processing Systems, volume 34, pp. 22419–22430. Curran Associates, Inc., 2021. URL https://proceedings.neurips.cc/paper/2021/file/bcc0d400288793e8bdcd7c19a8ac0c2b-Paper.pdf.
  76. 76.Wu, S., Xiao, X., Ding, Q., Zhao, P., Wei, Y., and Huang, J. Adversarial sparse transformer for time series forecasting. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., and Lin, H. (eds.), Advances in Neural Information Processing Systems, volume 33, pp. 17105–17115. Curran Associates, Inc., 2020. URL https://proceedings.neurips.cc/paper/2020/file/c6b8c8d762da15fa8dbbdfb6baf9e260-Paper.pdf.
  77. 77.Yanchenko, A. K. and Mukherjee, S. Stanza: A nonlinear state space model for probabilistic inference in non-stationary time series. arXiv, pp. 2006.06553v1, 2020.
  78. 78.Zhang, G., Eddy Patuwo, B., and Y. Hu, M. Forecasting with artificial neural networks:: The state of the art. International Journal of Forecasting, 14(1):35–62, 1998. ISSN 0169-2070. doi: https://doi.org/10.1016/S0169-2070(97)00044-7. URL https://www.sciencedirect.com/science/article/pii/S0169207097000447.

Citation

MLA
Drouin, A., et al. “TACTiS: Transformer-Attentional Copulas for Time Series”. International Conference on Machine Learning, vol. 162, 2022, pp. 5447–93, https://proceedings.mlr.press/v162/drouin22a.html.
APA
Drouin, A., Marcotte, É., & Chapados, N. (2022). TACTiS: Transformer-Attentional Copulas for Time Series. International Conference on Machine Learning, 162, 5447–5493. https://proceedings.mlr.press/v162/drouin22a.html
Chicago
Drouin, A., É. Marcotte, and N. Chapados. 2022. “TACTiS: Transformer-Attentional Copulas for Time Series”. International Conference on Machine Learning 162: 5447–93. https://proceedings.mlr.press/v162/drouin22a.html.
Harvard
Drouin, A., Marcotte, É. and Chapados, N. (2022) “TACTiS: Transformer-Attentional Copulas for Time Series”, International Conference on Machine Learning. PMLR, pp. 5447–5493. Available at: https://proceedings.mlr.press/v162/drouin22a.html.
Vancouver
1. Drouin A, Marcotte É, Chapados N (2022) TACTiS: Transformer-Attentional Copulas for Time Series. In: International Conference on Machine Learning. PMLR, pp 5447–5493

BibTeX

@InProceedings{pmlr-v162-drouin22a,
  title = 	 {{TACT}i{S}: Transformer-Attentional Copulas for Time Series},
  author =       {Drouin, Alexandre and Marcotte, \'Etienne and Chapados, Nicolas},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {5447--5493},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/drouin22a/drouin22a.pdf},
  url = 	 {https://proceedings.mlr.press/v162/drouin22a.html},
  abstract = 	 {The estimation of time-varying quantities is a fundamental component of decision making in fields such as healthcare and finance. However, the practical utility of such estimates is limited by how accurately they quantify predictive uncertainty. In this work, we address the problem of estimating the joint predictive distribution of high-dimensional multivariate time series. We propose a versatile method, based on the transformer architecture, that estimates joint distributions using an attention-based decoder that provably learns to mimic the properties of non-parametric copulas. The resulting model has several desirable properties: it can scale to hundreds of time series, supports both forecasting and interpolation, can handle unaligned and non-uniformly sampled data, and can seamlessly adapt to missing data during training. We demonstrate these properties empirically and show that our model produces state-of-the-art predictions on multiple real-world datasets.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/